理论工具箱:基于智能体的LLM辅助经济理论研究工具

Theorist Toolbox: Tools for Agent Based LLM-assisted economic theory Research

精选理由

这篇论文为你演示了如何用LLM做经济理论研究,重点不是让模型生成答案,而是设计验证流程来确保结果可靠,三种方法对比很清楚。

AI 摘要

该论文提出一种以验证为先的LLM辅助经济理论协议,并实例化为三种方法:单次严谨通道、对抗性验证器对(Claude Opus 4.8提议,OpenAI Codex反驳,作者仲裁)以及带评审门控的结构化多智能体项目。作者在一个开放示例——为Gans-Kominers等级膨胀模型设计Groves/Pigouvian激励相容机制——上评估该协议,三个运行均未产生严格直接揭示VCG/Clarke机制,对抗性通道自身证实了该点。结果揭示三个反复出现的现象:收敛发现、对抗验证的有效性、以及抛光不等于严谨。

原文 · arXiv: OpenAI

Theorist Toolbox: Tools for Agent Based LLM-assisted economic theory Research

Empirical economists inherit a toolbox. Shared packages, replication archives, and circulated guides etc. Theorists largely start from a blank page. By 2026, large language models can produce and check nontrivial mathematics, so the binding constraint on machine-assisted theory is no longer production but trust: a fluent model will prove a false theorem as readily as a true one. I propose a verification-first protocol for doing economic theory with a language model and instantiate it as three reusable methods that differ on a single axis, how the work is checked: a single disciplined pass, an adversarial prover-verifier pair (Claude Opus~4.8 proposing, OpenAI Codex refuting, the author triaging), and a structured multi-agent project with a reviewer gate. I evaluate the protocol on one open worked example: designing a Groves/Pigouvian incentive mechanism for the Gans-Kominers eigengrade model of grade inflation; none of the three runs produced a strict direct-revelation VCG/Clarke mechanism, a point the adversarial pass itself established. The evidence is a single worked example with one model pairing run by one operator, so what follows are demonstrations rather than measured effects. Three phenomena recur. First, convergent discovery: two runs derive the same effective-resistance externality kernel on opposite margins. Second, adversarial verification is load-bearing: the pair caught three of its own false claims and the gate rejected a sub-goal. Third, polish is not rigor: the most finished-looking output was the least verified. The methodological takeaway is that external verification, not model capability, is the design variable.