#LLMs
牛津研究团队发表新论文,通过数学论证指出,大模型(LLMs)无法真正“发明”新事物,因为它们本质上是预测下一个词的模型,无法相信现有数据中错误的信息。作者认为,人类能够基于现有数据提出新理论,而大模型则只能反映过去。作者每天使用这些模型,发现新突破通常来自人类判断现有数据错误并做出新理论。
Datasette 发布了两个安全补丁版本,修复了 Claude Fable 5.1、GPT-5.6 和 GPT-6 Astra 等大模型审计发现的漏洞,对在公共网络上运行且混合了公开和私有表格的实例很重要。
This paper introduces a new reporting protocol for scientific reports that could significantly improve reproducibility. It's a must-read for anyone interested in improving the quality of scientific reporting and the use of LLMs in research.
This research provides a new approach to improve the reliability of LLMs in decision-making by calibrating claim-level confidence, which is particularly useful for high-stakes domains where accurate uncertainty estimation is crucial.
SchemaGUI提供了一个新的基准,用于评估可控制GUI生成,特别是对于LLMs在几何空间控制和布局复杂性方面的表现进行了深入分析,对于想要了解GUI生成领域最新进展的人来说是个好资源。
这篇论文评估了LLMs在语法工程中的实用性,特别是针对Cantonese ParGram资源,与GPT-5.4相比,gpt-oss-120b表现稍逊。研究有助于了解LLMs在语法工程中的优势和局限性。
这项研究提出了个性化隐私控制的新方法,通过注意力头干预在LLMs中实现,对于关注隐私保护的用户来说是个好消息,它比现有方法更可靠地保护用户隐私。
Drew Breunig分享了他对Fable的看法,揭示了免费午餐时代的终结,对于关注LLM定价和AI发展的读者来说,这是一篇值得阅读的文章。
Rich Sutton, a Turing Award winner, shares his insights on LLMs, highlighting their limitations and the broader scope of AI intelligence.