论文精选

LLM代码理解复杂度评估研究

Complexity-Aware Evaluation of LLM Comprehension

精选理由

DeepSeek-Coder-V2和Llama在代码理解上的表现如何?复杂度增加对模型准确率有什么影响?

该研究提出了一种基于代码复杂度的评估框架,使用循环复杂度、嵌套深度、分支因子和Halstead体积四个指标。研究评估了DeepSeek-Coder-V2和Llama模型在300个Python函数上的自动输入输出预测,以及在60个函数上的语义理解。DeepSeek-Coder-V2整体准确率为78.33%,Llama为70.33%。随着复杂度从低到高,两个模型的准确率均显著下降。

原文 · arXiv: DeepSeek

Complexity-Aware Evaluation of LLM Comprehension

Large language models (LLMs) are increasingly used for software engineering tasks that require understanding existing source code, including behavior prediction, function explanation, debugging, and code review. However, aggregate benchmark accuracy can conceal how model reliability changes as source code becomes structurally more complex. This paper presents a complexity-aware framework for evaluating LLM code comprehension using cyclomatic complexity, nesting depth, branching factor, and Halstead volume. We evaluate DeepSeek-Coder-V2 and Llama through two complementary tasks: automatic input-output prediction over 300 Python functions and manually assessed semantic comprehension over a balanced subset of 60 functions. The functions are grouped into Low-, Medium-, and High-complexity bands. DeepSeek-Coder-V2 achieves an overall automatic accuracy of 78.33%, compared with 70.33% for Llama. However, accuracy decreases substantially from Low to High complexity, from 93.52% to 52.78% for DeepSeek-Coder-V2 and from 87.04% to 47.22% for Llama. Incorrect predictions are consistently associated with higher values of all four complexity metrics, and correlation and logistic-regression analyses confirm broadly comparable negative associations between structural complexity and correctness. Manual semantic comprehension shows the same degradation pattern, with accuracy decreasing from 100.00% to 75.00% for DeepSeek-Coder-V2 and from 90.00% to 60.00% for Llama. These findings demonstrate that complexity-aware evaluation provides a more diagnostic assessment of LLM code-comprehension reliability than aggregate accuracy alone.