Curriculum Brain:用11个智能体自动构建课程知识图谱
Curriculum Brain: Constructing Curriculum Knowledge Graphs as a Substrate for Cognitive Diagnosis
教育领域一个很实际的问题:手工标注 Q-matrix 太贵,作者用 11 个智能体流水线自动建课程知识图谱,每章只要 1.19 美元,论文还诚实标注了评估局限。
这篇论文提出 Curriculum Brain,一个两仓库系统,用版本控制知识库加 11 个单一职责智能体从官方课程文档自动生成概念-技能映射,用于替代手工编写 Q-matrix。在 241 次章节运行(168 个不同章节)中,41.5% 的生成结果一次通过双重校验,67.6% 无需人工介入;第 77 次运行后的单一配置改动将两项数字提升到 57.0% 和 91.5%。每章 API 成本为 1.19 美元,不含人工审核开销。论文明确说明评估基于系统自产标准,仅反映内部一致性而非外部一致性。
Curriculum Brain: Constructing Curriculum Knowledge Graphs as a Substrate for Cognitive Diagnosis
Cognitive Diagnostic Models (CDMs) identify which specific skills a student has and has not mastered, the signal a personalized learning path needs and a single aggregate score cannot give. Yet they are rarely deployed. The obstacle is their precondition: the Q-matrix, a mapping from every assessment item to the skills it requires, historically authored by hand. We separate the task into two stages: first construct the curriculum's own knowledge graph, the full space of concepts and skills it contains, independent of any item; then map items against that graph on demand. This paper addresses the first stage only. The item-mapping stage is designed but not implemented here, so the claim that this shifts judgment cost from once per item to once per curriculum is a design rationale rather than a finding. We present Curriculum Brain, a two-repository system pairing a version-controlled knowledge base with an agentic pipeline of eleven single-responsibility agents under a thin deterministic orchestrator. It generates candidate concept-skill mappings from official curriculum documents, checks them against accumulated rules, and compares them with a concept-skill map extracted independently from the textbook, repairing its own failures and escalating to a human only when it cannot resolve a case itself. Across 241 chapter runs (168 distinct chapters), 41.5% produced a Generator output passing both checks without a patch, and 67.6% resolved without escalation. Both are measured against criteria the system itself produced, so both describe internal consistency rather than agreement with an external standard, and both pool two pipeline configurations separated by a single change at run 77; after it the figures are 57.0% and 91.5%. Observed spend was $1.19 per chapter, API spend only, excluding human review. We release both the framework and the resulting curriculum dataset.