这篇论文评估了LLMs在语法工程中的实用性,特别是针对Cantonese ParGram资源,与GPT-5.4相比,gpt-oss-120b表现稍逊。研究有助于了解LLMs在语法工程中的优势和局限性。
本文提出新的Cantonese ParGram资源,并评估了LLMs在控制实验范式下的知识驱动语法工程。使用Cantonese ParGram资源作为黄金标准,与相应的英语基准进行对比,研究OpenAI的gpt-oss-120b和GPT-5.4是否能在系统变化的提示条件下从句子和目标正式结构生成可机器处理的语法。GPT-5.4优于gpt-oss-120b,而从目标正式结构生成的语法通常优于从句子生成的语法。尽管两种模型都能生成局部合理的短语结构规则、词汇条目和模板,但它们在多构造设置中往往难以协调相互作用的正式约束。研究结果描述了当前LLMs的潜在集成到AI辅助专家工作流程的能力和局限性:LLMs可能支持语法发展的中间阶段,但人类语言学专业知识在分析、验证和细化中仍然至关重要。该研究还贡献了新的Cantonese符号语法资源。
How Useful are LLMs for Grammar Engineering? Cantonese ParGram Resources and Controlled Experimental Evaluation with English Baselines
This paper presents new Cantonese ParGram resources and evaluates LLMs for knowledge-driven grammar engineering within a controlled experimental paradigm. Using Cantonese ParGram resources as gold standards, with corresponding English baselines, we investigate whether OpenAI's gpt-oss-120b and GPT-5.4 can generate machine-processable grammars from sentences and target formal structures under systematically varied prompting conditions. GPT-5.4 outperformed gpt-oss-120b, while grammars generated from target formal structures generally outperformed those generated from sentences. Although both models could generate locally plausible phrase-structure rules, lexical entries, and templates, they often struggled to coordinate interacting formal constraints, especially in multi-construction settings. The results characterize both the capabilities and limitations of current LLMs for potential integration into AI-assisted expert workflows: LLMs may support intermediate stages of grammar development, but human linguistic expertise remains central to analysis, validation, and refinement. The study also contributes new Cantonese symbolic grammatical resources.