LLMPEDIA 让你能直接查看和比较三大模型的知识准确性,发现它们在 MMLU 外的真实表现。
LLMPEDIA 从 GPT-5-mini、DeepSeek-V3.2 和 Llama-3.3-70B 三个模型家族中提取了约 130 万篇文章。研究团队对分层抽样的事实声明进行审核,发现真实准确率为 68.4%,比 MMLU 基准低 21 个百分点。30.5% 的声明无法被维基百科或网络资源验证,属于长尾知识或合理幻觉。该平台提供五种一键视图,包括链接遍历探索、声明级事实核查、跨模型和政治立场比较等。
LLMPEDIA: Browsing, Verifying, and Comparing the Parametric Encyclopedic Knowledge of LLMs
Flagship language models appear saturated on benchmarks like MMLU (Hendrycks et al., 2021), scoring above 90% - yet benchmarks test only what the experimenter thought to ask, the availability bias of fixed question sets. LLMPEDIA makes this bias measurable and browsable. We recursively materialized ~1.3M articles from three model families' parametric memory (GPT-5-mini, DeepSeek-V3.2, Llama-3.3-70B) without retrieval, then audited a stratified sample of atomic claims against Wikipedia and a curated web stack, coloring every claim supported, refuted, or insufficient (Saeed and Razniewski, 2026). On a uniform random sample the true rate is 68.4% - more than 21 pp below MMLU - with 30.5% of claims insufficient: assertions no benchmark probes and the world's largest encyclopedia cannot adjudicate - long-tail knowledge or plausible hallucination, the evidence cannot tell - extending to free text the coverage gap GPTKB established for triples (Hu et al., 2025). The resulting live, open encyclopedia lets visitors inspect this frontier one claim at a time through five one-click views - link-traversal exploration, claim-level factuality, cross-model and political-persona comparison, and a guided topic drill-down - each page, claim, and verdict at a stable URL. LLMPEDIA is live at https://llmpedia.net
- Hugging Face: Blog09-01 21:39原文