做创造力评估或 AI 教育对话系统的研究者值得关注——IntElicit 解决了静态测试无法捕捉真实创造力的痛点,用对话策略优化让评估更贴近实际场景。
IntElicit 是一个用于评估情境化创造力的框架,它通过对话策略优化来减少认知能力和参与意愿等非创造性因素的干扰。该框架作为自适应 AI 面试官,在多轮交互中提供非指导性知识和参与支持,同时保留参与者生成创造性内容的责任。它引入分解过程奖励机制,避免奖励作弊,鼓励引导参与者推理而非直接给出答案。实验表明,IntElicit 能比专家设计的基线方法更好地激发创造性成果,揭示静态评估可能遗漏的创造潜力。这为 AI 辅助学习中的情境化创造力评估提供了形成性和诊断性视角。
IntElicit: Eliciting and Assessing Contextualized Creativity via Dialogue Policy Optimization
Contextualized assessment offers high ecological validity for evaluating creativity but introduces a critical challenge: observed performance may be confounded with cognitive proficiency (domain knowledge) and agency (willingness to engage). Meanwhile, in the age of generative AI, creative problem solving increasingly occurs in tool-mediated and human--AI interactive environments, making fully static assessment less aligned with contemporary creative practice. To address these issues, this paper proposes IntElicit, a framework for eliciting and assessing contextualized creativity via dialogue policy optimization. IntElicit functions as a constrained adaptive AI Interviewer: it provides non-directive knowledge and agency scaffolds in multi-turn interaction to reduce non-creative confounders, while preserving participants' responsibility for generating the creative content being evaluated. Specifically, to tackle sparse rewards and potential reward hacking (e.g., answer dictation) in open-ended educational dialogue, IntElicit introduces a decomposed process reward mechanism. This mechanism aligns the policy with pedagogical elicitation, rewarding prompts that draw out participant reasoning rather than producing optimal answers on their behalf. Extensive experiments, including participant simulation and a human subject study (N=64), show that IntElicit improves elicited creative outcomes over expert-designed baselines. Together, the results suggest that interactive elicitation can reveal creative potential that static FPSP-style assessment may miss, providing a formative and diagnostic lens for contextualized creativity assessment in AI-mediated learning contexts.