这篇论文用一个具体案例(葡萄牙语AMALIA)揭露了LLM做数据标注的隐患:表面一致但内在无效,对做标注或评测的人很有启发。
葡萄牙国家语言模型AMALIA(9B参数)在标注道德基础“权威”时,与人类编码者一致性达与8-13倍大小模型相差6个F1点。但通过“恢复间隙”方法发现,分解提示后AMALIA仅恢复约一半性能,错误分析显示它依赖表面相关捷径(如对权威人物的道德愤怒)。多语言开源模型在相同葡萄牙语语料库上能缩小此差距,表明问题不在语料库本身。研究认为主权LLM基准应测试证据路线而非仅一致性。
Validity of LLMs as data annotators: AMALIA on authority
A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicly funded 9B-parameter model for European Portuguese, appears competitive on agreement alone: asked to code the moral foundation of authority, it agrees with trained human coders to within six F1 points of open models eight to thirteen times its size. Yet agreement is reliability, not validity. For theoretical constructs that must be inferred rather than read from surface features, the question is whether the model follows the construct's theory or reaches the right code by correlated shortcuts. We test this with the recovery gap: the loss in performance when a holistic prompt is decomposed into the codebook's atomic clauses and recombined by the theory's explicit rule. If calibration closes that gap, some portability should survive across models and languages; where it does not, the construct-model instrument is the likely locus of failure. We ask whether a calibrated English instrument transfers to AMALIA-9B and to European Portuguese. For one construct and one corpus, it does not. Decomposition recovers only about half of AMALIA's holistic performance, and error analysis suggests reliance on surface correlates, especially moral outrage near authority figures. An open multilingual LLM closes the gap on the same Portuguese corpus under the same instructions, pointing away from the corpus as the main explanation. AMALIA can still screen and pre-code at scale, but it cannot yet measure this construct well enough to stand alone. The study is a single counterexample, not a verdict on national models; it argues that sovereign-LLM benchmark batteries should test not only agreement with human coders, but the evidential route by which that agreement is warranted.