多智能体LLM团队中的领导协调控制:行为特征与恢复优势边界

Leadership as Coordination Control: Behavioral Signatures and the Recovery-Advantage Boundary in Multi-Agent LLM Teams

精选理由

这篇论文用实验证明多智能体团队里领导不是万能的,只有在初始投票不靠谱且能补救的特定条件下才有用,比如情境领导在Llama-4-Scout上提升了8个点。挺扎实的研究。

AI 摘要

该论文研究多智能体LLM团队中过程级协调控制的价值,通过行为签名(多数锁定、探索、恢复)和逐动作消融实验,将交易型、变革型、情境型三种领导风格作为控制器。在四种任务制度和三个开源模型族(包括Llama-4-Scout)的12种组合中,没有控制器在准确率上占优,交易型控制与共享第0轮投票的差距在1.3个百分点内。情境型控制在Llama-4-Scout social任务上比平坦基线高出8个百分点,仅当初始多数不可靠且任务可恢复时才有效。结果表明协调控制是权变,而非排行榜驱动,与团队科学的权变理论一致。

原文 · arXiv cs.AI

Leadership as Coordination Control: Behavioral Signatures and the Recovery-Advantage Boundary in Multi-Agent LLM Teams

Team science holds that leadership is contingent: it helps only under specific conditions, and capable, autonomous teams may need none at all. We ask the analogous question for multi-agent LLM teams: under what measurable conditions does process-level coordination control add value, and do those conditions match what team science predicts? We use behavioral signatures (majority lock-in, exploration, recovery from an incorrect round-0 consensus) and per-action ablations, clean because each controller is an explicit action set, not a monolithic prompt. We operationalize three classical leadership styles (transactional, transformational, situational) as controllers over a shared action vocabulary (explore, revise, accept, synthesize). A matched controller with the same actions but an arbitrary rule recovers no better than majority voting, so the theory-derived rule, not the vocabulary, does the work. Across four task regimes and three open-weight model families, no controller dominates by accuracy, as the contingency view predicts: transactional control matches a shared round-0 vote on all 12 (model, regime) combinations to within 1.3pp, and gains appear only on the one combination where the round-0 majority is unreliable (llama-4-scout social; situational +8pp over flat). A recovery-advantage account, tested with four boundary probes, says a controller beats plain interaction only where the round-0 majority is unreliable, the task is recoverable, and undirected interaction does not already repair it. These regions map onto contingency theory (leadership substitutes, path-goal redundancy, the situational readiness gap), so a largely null accuracy result is what the theory predicts, not a failure of the controllers. We read process-level coordination control as a contingency to be measured and theory-mapped, not a leaderboard to be topped.