8月21日
10:33
10:33官方一手arXiv: DeepSeek@Xinyi Fan, Miri Liu, Ruozhen Yang, Siru Ouyang, Jiawei Han
This paper introduces StateMemBench, a benchmark for memory systems to track evolving states in long interactions. It shows that current memory systems struggle with this task. The paper also proposes StateMem, a new memory method that improves current-state accuracy significantly. It can be applied as a lightweight wrapper to existing systems, enhancing their performance on the benchmark.
推荐理由:This paper presents a new benchmark and memory method that could significantly improve the performance of AI agents in long interactions. It's a must-read for anyone interested in AI memory systems and benchmarks.
7月30日
11:17
11:17官方账号arXiv cs.AI@Jingbo Zhou, Yusai Zhao, Qi Bao, Jingjia Cao, Zhenghai Chen, Chang Gao, Kaiqi Guo, Muxin Guo, Mingxuan Li, Xinjiang Lu, Yanru Ma, Yixiong Xiao, Zenghui Zhang, Le Zhang, Hua Wu
新基准 OmegaUse-OfficeVal 包含 100 个办公套件任务,平均需 2.32 小时人工完成。每个任务提供人工时间与任务价格代理两个经济信号,用于对比人类成本与 LLM 推理成本。评估了 GPT-4、Claude 3.5 等前沿 LLM,结果显示它们虽更便宜更快,但交付质量尚未达到人类水平。该基准代码与数据集已开源。
推荐理由:想做办公自动化?这个新基准用真实成本告诉你:现在LLM虽便宜但还没人类靠谱,值得看看差距在哪。
5月11日
00:20
00:20官方一手OpenAI Blog博客/媒体
75°
OpenAI发布Gym公测版,这是一个用于开发和比较强化学习算法的标准化工具包,包含从模拟机器人到Atari游戏等丰富的环境集合。同时提供结果比较和复现平台,旨在推动RL研究的可复现性和标准化。
事件专题

推荐理由:为AI从业者提供了一个统一的强化学习基准平台,极大降低了算法测试与对比的门槛,是RL研究的必备基础设施。
00:17
00:17官方一手OpenAI Blog博客/媒体
精选80°
OpenAI开源Universe平台,提供一个包含游戏、网站等多样化环境的测试平台,用于衡量和训练AI的通用智能。该平台通过标准化接口,让AI代理能像人类一样与各类应用交互,加速通用人工智能研究。
事件专题

推荐理由:Universe为AI研究者提供了首个大规模、标准化的通用智能评估环境,直接推动AGI训练与基准测试发展。