8月25日
10:26
10:26官方一手arXiv: DeepSeek@Toghrul Abbasli, Kentaroh Toyoda, Yuan Wang, Li Chen
This paper introduces claim-level confidence calibration for LLMs, addressing the issue of misaligned confidence with factual correctness. It proposes a framework that decomposes responses into claims and assigns calibrated confidence using inference-time signals. The framework is evaluated on six recent models and shows reduced calibration error on factual questions while revealing failure modes on adversarial false-premise questions.
推荐理由:This research provides a new approach to improve the reliability of LLMs in decision-making by calibrating claim-level confidence, which is particularly useful for high-stakes domains where accurate uncertainty estimation is crucial.