论文

LLM图重建的谱理论:边界与实证

A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization

精选理由

这篇论文揭示了LLM图重建评估的数学边界,通过135个案例展示了不同模型的编辑策略差异。

该研究针对语言模型图重建的评估问题,证明了Wasserstein距离在拉普拉斯谱之间的界限。研究分析了135个重建案例,涉及三个开源模型在45个合成图上的表现。77个输出为单向编辑,29个为混合编辑,其中包含边数完全保留但同时添加和删除19条边的情况。不同模型展现出不同的编辑策略,从简单复制输入到尝试完成但产生大量幻觉。

原文 · arXiv cs.LG

A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization

Evaluations of graph reconstruction by language models typically report a single aggregate distance between the original and the reconstructed graph. We prove that for the Wasserstein distance between Laplacian spectra such a summary is bracketed by two edge counts, the net change in edge number from below and the symmetric difference from above, each scaled by $2/n$ where $n$ is the number of vertices. The bracket is sharp: its two ends coincide exactly when the reconstruction only adds edges or only deletes them, and on that class the distance is a rescaled edge count that says nothing about which edges changed. When the ends differ, the residual between the distance and the lower end is positive only if the reconstruction both invented and lost edges, which turns it into a certificate of mixed editing computable from the reported summaries alone. We characterize these regimes in 135 reconstructions produced by three open-weight models over 45 synthetic graphs. Seventy-seven outputs are one-sided and 29 mixed outputs have $X > 0$, including cases where edge count is exactly preserved while nineteen edges were simultaneously invented and lost. The three models differ in editing policy, ranging from copying the input to attempting completion at the cost of large hallucination volume, a distinction that aggregate distortion does not reveal.