论文

Analog Design Bench 发布:50 个模拟电路任务测试智能体长时程设计能力

Long-Horizon Analog Design Bench: Benchmarking Agents on Hours-Long Analog and Mixed-Signal Circuit Design Tasks

精选理由

模拟电路版 SWE-bench 来了:50 个真实设计任务,15 种智能体跑 2250 次实验,DeepSeek V4 Pro 靠参考拓扑涨了 18.7 个点。

Analog Design Bench 收录 50 个晶体管级模拟与混合信号电路设计任务,由 17 位芯片设计师提供。评测在 2,250 次两小时尝试中测试 15 种智能体配置,全规格通过率从 8.0% 到 78.0% 不等。失败分析显示多数提交未违反合法性规则,而是无法通过电气规格验收。为智能体提供与任务匹配的参考拓扑使 DeepSeek V4 Pro 提升 18.7 个百分点,GPT-5.6 Sol 则主要获得加速。通用技能文档几乎无益,有时反而降低表现。

原文 · arXiv: DeepSeek

Long-Horizon Analog Design Bench: Benchmarking Agents on Hours-Long Analog and Mixed-Signal Circuit Design Tasks

Coding agents now sustain hours-long, tool-driven loops, yet their ability to carry long-horizon analog and mixed-signal circuits to electrical specification remains unmeasured. We introduce Analog Design Bench, a long-horizon agentic benchmark of 50 transistor-level design tasks contributed by 17 chip designers. Agents work with an open-source simulator, while an isolated verifier evaluates the submitted circuit using specification-based electrical tests. We evaluate 15 agent configurations across 2,250 two-hour attempts and observe full-specification pass rates from 8.0% to 78.0%. Coding-benchmark performance correlates with analog results but leaves much of the performance spread unexplained. Our failure analysis shows that most unsuccessful submissions have no recorded legality rejection but fail electrical acceptance, identifying electrical closure as the dominant endpoint challenge. We test time, reasoning effort, agent harness, and supplied design knowledge as interventions. Longer budgets and higher reasoning effort improve performance, while general skill documents provide little benefit and sometimes reduce performance. Supplying a task-matched reference topology, an idealized form of circuit-IP retrieval, raises DeepSeek V4 Pro by 18.7 percentage points and mainly accelerates GPT-5.6 Sol.