AI产品精选73°

ATIBA:研究论文的完整性与质量检查工具

ATIBA: Grounded Integrity and Quality Checking for Research Papers

精选理由

ATIBA 自动检查论文引用完整性和会议合规性,用 GPT-5.4 做多模式审查,比人工检查更一致可靠。

AI 摘要

ATIBA是一款研究论文质量检查工具,提供五项 grounded integrity 和 quality 检查。该工具通过引用完整性检查验证每篇引文,标记撤回或无法找到的参考文献。它还提供/track 合规性检查,直接从会议的 call-for-papers 页面提取提交标准并评估论文。此外,ATIBA 还包含针对 ACM SIGSOFT 经验标准的合规性检查,使用 GPT-5.4 通过 Azure OpenAI 进行多模式 AI 审查,并提供引用建议功能。研究显示,13 名参与者在六项调查中的平均同意率为 85%。

原文 · arXiv: OpenAI

ATIBA: Grounded Integrity and Quality Checking for Research Papers

Checking a manuscript's reference integrity, its compliance with a target venue's specific submission rules, and its adherence to community reporting standards is manual, repetitive, and different for every venue so in practice it is done inconsistently or skipped. We present ATIBA, a tool that runs five grounded integrity and quality checks on a manuscript: a reference-integrity check that verifies each citation against bibliographic sources and flags retracted or unfindable references; a venue/track compliance check that derives submission criteria directly from a venue's own call-for-papers page and evaluates the manuscript against them, each verdict anchored to a verbatim quote from that page; an empirical-standards compliance check against the ACM SIGSOFT Empirical Standards, with a hallucination defence that discards any evidence quote it cannot locate verbatim in the manuscript; a multi-mode AI review (venue-specific, formal, and page-anchored annotation) powered by GPT-5.4 through Azure OpenAI; and a citation-suggestion feature that proposes candidate references for a manuscript and verifies each against bibliographic sources before it is shown to the user. All five checks are designed around the same principle: an LLM is only trusted to judge, never to invent the evidence it judges against. We evaluated ATIBA through a moderated user study with 13 non-author participants. Agreement across the six survey items ranged from 69% to 92%, with a mean of 85%, providing initial evidence of positive perceived usefulness across the evaluated workflows. These findings establish perceived usefulness; objective accuracy remains to be measured.