InsufficiencyBench:评估LLM在用户查询不足方面的法律建议

InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries

精选理由

这篇论文介绍了InsufficiencyBench,这是第一个评估LLM在用户查询不足方面的法律建议的基准。它对法律AI系统提出了新的挑战,并提供了评估模型性能的新方法。

AI 摘要

法律AI系统日益用于回答法律问题,但现有基准假设查询完全指定。InsufficiencyBench是第一个针对查询不足的法律基准,评估模型是否识别查询缺乏法律相关信息,识别缺失内容,并避免过早结论。该基准包含八个典型缺失元素类别,涵盖三种结构故障模式,并构建了202个基准项目,涵盖六个法律领域和24个美国司法管辖区。评估了十个前沿模型,发现没有模型在缺失元素识别上超过F2 = 0.46,平均召回率为0.44。模型要么无差别地打掩护,要么在虚构的假设下沉默。没有模型既能识别和限定对不足查询的响应,又能直接针对完整查询。

原文 · arXiv cs.AI

InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries

Legal AI systems are increasingly used to answer legal questions, yet existing benchmarks assume queries arrive fully specified. In practice, users omit facts that materially determine the legal outcome. We introduce InsufficiencyBench, the first legal benchmark targeting query-side insufficiency: whether a model recognizes when a query lacks legally material information, identifies what is missing, and refrains from premature conclusions. We formalize a taxonomy of eight canonical missing-element categories across three structural failure modes---switch, gating, and fatal prerequisite--- and construct 202 benchmark items (58 base queries, 144 deficient variants) spanning six legal domains and 24 US jurisdictions and annotated by practising attorneys. Evaluating ten frontier models, we find that no model exceeds F2 = 0.46 on missing-element identification and that the median recall is 0.44. Models either hedge indiscriminately or answer silently under fabricated presumptions. No model both identifies and qualifies responses to deficient queries while directly addressing complete ones.