有人拿 8 个决策模型打 Tetris,Perplexity Decider 稳居第一,还能自己在 playground 里复现,挺好玩的。
#benchmark
共 12 条 · 7 天 1 条 · 30 天 6 条
信息流里打上「benchmark」标签的资讯、产品与论文,按刊登时间排,新的在上。
10月7日
8 个决策模型玩俄罗斯方块测试,Perplexity Decider 排名第一
9月25日
BRIDGE ASR 2.0 发布:18 种印度语言加西语葡语越语的真实对话语音识别基准
humynlabs 出了个语音识别新基准,专测真实对话和混语言场景,23 个模型已上榜,做语音方向的可以拿来自测。
9月23日
Trains but Doesn't Learn:面向 LLM 智能体交付能力的 Post-Training 基准
这篇论文造了个很较真的基准:不看智能体能不能把指标刷上去,只看它交付的模型到底行不行,连"训练了但没学到东西"这种静默失败都能抓住,四个前沿模型全被拉出来和人类工程师对比。
9月17日
8月21日
论文10:33
Can Agent Memory Systems Track Evolving State?This paper presents a new benchmark and memory method that could significantly improve the performance of AI agents in long interactions. It's a must-read for anyone interested in AI memory systems and benchmarks.
7月30日
6月13日
6月10日
5月11日