研究:LLM 推理投入增加并不能稳定提升量化交易收益
The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?
拿 DeepSeek、GPT、Gemini 跑了一年美股、80 万次预测,结论是推理算力花在交易上不划算,做量化的人可以看看实验设计。
arXiv 论文对 DeepSeek、GPT、Gemini 三个模型家族做了控制变量实验,在固定提示词、输出格式和组合构建的条件下调整推理力度。评测覆盖一年美股、数值/可识别新闻/掩码新闻三种输入、超 80 万条资产预测。结果显示三个家族的额外推理都没有带来可靠的净收益提升,其中 DeepSeek 从无推理到最大推理的表现呈非单调变化。重复生成也会导致组合选择不稳定,即使总体评分接近。
The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?
While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs. We conduct a controlled study of representative LLMs from the DeepSeek, GPT, and Gemini families. We vary reasoning effort while holding information available at each formation date, prompts, output formats, and portfolio construction fixed. Our evaluation covers a full year of U.S. equities under three input conditions: numerical, identifiable news, and masked news. It includes more than 800,000 asset predictions and repeated model generations. Across all three model families, additional reasoning does not produce a reliable improvement in net portfolio returns. For DeepSeek, where we examine the full progression from no reasoning to maximum reasoning, performance is nonmonotonic. Repeated generations also produce unstable treatment effects and portfolio selections, even when overall scores remain similar. These findings show that additional reasoning can change financial decisions without reliably improving their economic value, motivating validation for each task before deployment.