这篇论文讲怎么用LLM自动造认知决策场景,还验证了复杂度靠谱,做AI评估或心理学实验的可以看看。
该研究开发了一个自动化流程,用大语言模型生成结构化决策场景,并通过基于任务复杂度理论的复合框架验证其复杂度。研究评估了4,238个场景,五个独立模型家族的一致性近乎完美,组内相关系数达0.997,kappa值为0.971。已知组效度显示各复杂度层级间分离明显,eta-squared为0.587,所有配对比较p值小于0.001。因子分析显示一个主导的复杂度构念,载荷在0.87到0.96之间,交互性构成较弱的次要维度。Llama 4 Maverick生成速度最快,每分钟134个场景,而DeepSeek Chat V3.2在领域覆盖和模式合规性上更均衡。
Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models
Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and biased. We developed an automated pipeline that uses LLms to generate structured decision scenarios and validates their complexity through a composite framework rooted in established task-complexity theory. We evaluated 4,238 scenarios across multiple domains and complexity tiers. Measurement validation met rigorous psychometric standards. Agreement among five independent model families was nearly perfect, with an intraclass correlation coefficient of 0.997 and a kappa of 0.971. Known-groups validity demonstrated large separation between tiers, with an eta-squared of 0.587 and all pairwise comparisons significant at p less than .001. Factor analysis revealed a dominant complexity construct, with loadings between 0.87 and 0.96 across three frameworks, while interactivity formed a weaker secondary dimension at 0.34. Discriminant validity was limited by a strong relationship between complexity and text length that persisted after controlling for tier, yielding a partial correlation of 0.86. This constrains construct purity but does not undermine the instrument's tier-grading function. Model analyses showed a negative association between throughput and schema pass rate (r = -0.967, p = .007, n = 5), suggesting a speed-quality trade-off, though largely driven by one high-throughput model. Llama 4 Maverick generated scenarios fastest at 134 per minute versus 25 for DeepSeek Chat V3.2, but underproduced complex-tier scenarios, whereas DeepSeek Chat V3.2 balanced domain coverage with high schema compliance. The system demonstrated strong psychometric properties, enabling reliable classification into Simple, Moderate, and Complex tiers and providing the measurement infrastructure needed for downstream cognitive assessment of AI systems