RetailAgent:多模态LLM交易代理的结构性逆向时机

RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents

精选理由

这篇论文研究了LLM交易代理是否存在可预测的结构性行为,发现其决策存在稳定的逆向时机模式。

AI 摘要

RetailAgent是一个实验框架,让大型语言模型观察匿名日内股票价格历史和允许状态,然后选择做多或平仓。研究显示,在跨模态、时间跨度、状态和模型家族中,存在持续负向时机。打乱保存的行动序列会显著减弱这一效应,表明行动与后续回报的 alignment 是负向评分的驱动因素。模型使用自我生成的记忆做决策时,策略持续性增强,且在同时使用两种行动的股票日上,时机负向性更明显。

原文 · arXiv cs.AI

RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents

In financial markets, a sequential policy that reacts systematically to price movements may become predictable to other market participants. This paper studies whether large language model (LLM) agents exhibit such directional structure through RetailAgent, an experimental framework in which an LLM observes anonymized intraday equity price histories and permitted state, then repeatedly chooses long (hold the stock) or flat (stay out) before the subsequent interval return is revealed. We compare returns during long and flat intervals along the same stock's intraday path after removing the overall fraction of long decisions. This exposure-matched measure reveals persistent negative timing across modality, horizon, state, and model family. Shuffling saved action sequences substantially attenuates the effect, showing that alignment between actions and subsequent returns drives the negative score. Feeding self-authored memories into decisions further increases policy persistence, while timing becomes more negative among stock-days on which the agent uses both actions. These results reveal stable, recoverable directional structure in sequential LLM financial decisions and a behavioral signal for studying how another participant could respond to a predictable policy.