Wyvern:用于生成有据多模态报告的多智能体框架

Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports

精选理由

写技术综述的可以看看,Wyvern自动出图出表还带引用,比基线靠谱多了。

AI 摘要

Wyvern是一个多智能体框架,用于自动生成有据可查的多模态技术报告,能统一集成图像、表格、文本及参考文献。其核心包含声明自动修订阶段,用于强化内容的信息锚定。人类评估显示,Wyvern生成的图表信息量在87%的案例中优于近期基线。Wyvern报告的有用性在63%至100%的实例中高于三种对比方法。自动评估中,Wyvern的引用召回率提升至基线2.3倍,引用精确率达1.6倍。

原文 · arXiv cs.AI

Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports

In the current artificial intelligence-driven innovation era, the pace of knowledge growth is accelerating, and is hard to keep up with. While generative models are increasingly used to synthesize content, they often lack in information grounding. To address these peculiarities of our time, we propose Wyvern, a multi-agent framework for the automated generation of grounded, multimodal technical reports. Wyvern allows for the generation of multimodal outputs, integrating images, tables, and text with supporting references in a unified report. Additionally, a particular focus is placed on the grounding of the content, with the implementation of a claims auto-revision stage. We conduct a human evaluation study to assess the quality of our proposed framework. The results show that the figures' informativeness is perceived as superior to that of a recent baseline in 87% of cases. Furthermore, Wyvern's reports are rated as more useful than those produced by three alternative methods in 63% to 100% of instances. We also carry out automatic evaluations showing that Wyvern gains up to 2.3$\times$ in citation recall and 1.6$\times$ in citation precision with respect to the baselines.