FHIR-AgentEval:临床LLM智能体基准测试沙盒
FHIR-AgentEval: a modular sandbox for benchmarking clinical LLM agents with memory-augmented configurations
做医疗智能体的可以看看这个沙盒,基于 FHIR 标准,专门测记忆增强配置对临床任务的影响,比自己搭评测省事。
npj Digital Medicine 于2026年10月10日在线发表 FHIR-AgentEval,一个模块化沙盒,用于在 FHIR 标准下评测临床 LLM 智能体。该框架支持记忆增强配置的对比实验,可系统比较不同记忆机制对智能体表现的影响。论文给出了可复用的评测流程与配置接口,面向医疗场景的智能体研究。
FHIR-AgentEval: a modular sandbox for benchmarking clinical LLM agents with memory-augmented configurations
npj Digital Medicine, Published online: 10 October 2026; doi:10.1038/s41746-026-03348-0 FHIR-AgentEval: a modular sandbox for benchmarking clinical LLM agents with memory-augmented configurations