求是引擎在AstaBench E2E-Bench-Hard基准测试中表现
Qiushi Engine on AstaBench E2E-Bench-Hard
求是引擎用DeepSeek-v4pro-preview模型在AstaBench测试中表现优异,完成率是官方最佳记录的3.3倍。
求是引擎v0.8在AstaBench E2E-Bench-Hard基准测试中得分0.816,每任务平均成本15.209美元。该测试包含40个任务,要求智能体完成从实验设计到报告交付的全流程。求是引擎在4个任务中完全达标,完成率达10%,比官方最佳记录高出7个百分点。在507个评估项目中,416项达标,达标率82.1%。
Qiushi Engine on AstaBench E2E-Bench-Hard
This report analyzes Qiushi Engine v0.8 across all 40 test tasks in AstaBench E2E-Bench-Hard, a benchmark that requires autonomous agents to carry a research question through experimental design, code implementation, actual execution, result analysis, and report delivery. Qiushi Engine is model-configurable; this evaluation selected DeepSeek deepseek-v4pro-preview as the model backend. The official AstaBench leaderboard records a score of 0.816 and an average benchmark cost of USD 15.209 per task, while the full-precision local recomputation is $81.59 \pm 1.87$. Four tasks satisfied every rubric item, yielding a full-task completion rate of 4/40 = 10% -- 7 percentage points above, and about 3.3 times, the approximately 3% best rate reported for AstaBench's official agents. Across 507 required rubric items, 416 were satisfied (82.1%). Official scoring archives and 40 Meta-Trace records show sustained production and verification of reports, code, and experimental artifacts; the principal gaps lie in repeated runs, external dependencies, specified metrics, and ablation studies. The report explains the benchmark, system workflow, aggregate results, representative cases, and limits of interpretation.