Grip on LLMs框架:荷兰政府用大模型评估基准

From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch

精选理由

荷兰政府项目组做的评估框架,测了30多个模型,告诉你哪个模型在事实性、能耗、成本上表现如何,选型时直接参考。

AI 摘要

荷兰研究者提出“Grip on LLMs”评估框架,针对政府场景的荷兰语大模型。该框架涵盖事实性、诚实性、社会偏见、能耗、成本和训练数据透明度六个维度,测试了30多个多语言及荷兰语专用模型。结果显示没有模型在所有维度上表现最佳,高质量模型通常伴随更高能耗和成本,而偏见与两者无关。事实性和诚实性由不同属性决定,高事实性不意味着高诚实性。研究团队发布了面向非技术用户的公开模型概览。

原文 · arXiv cs.AI

From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch

Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the "Grip on LLMs" framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation. Through an advisory board process, user research, and a survey of the users of a civil-servant chatbot, we identify six evaluation dimensions (factuality, honesty, social bias, energy consumption, cost, and training data transparency) and operationalise them into a benchmark suite covering more than 30 multilingual and Dutch-specific models. Our results reveal that no single model excels across all dimensions, and that trade-offs are unavoidable: higher quality consistently comes at greater environmental impact and financial cost, while bias remains largely independent of both. We further find that factuality (whether a model answers correctly) and honesty (whether a model acknowledges what it does not know) are governed by distinct properties, with high factuality not implying high honesty. To make these findings actionable for non-technical audiences, we release a publicly accessible, user-friendly model overview designed for the full range of stakeholders involved in governmental LLM selection, from engineers to policymakers.