评估大语言模型在反犹事件分类中的表现

Evaluating Large Language Models for Antisemitic Incident Classification

精选理由

这篇论文告诉你GPT-4o和Llama-3.2-3B在处理反犹事件时谁更强,还给出了提升分类效果的实用技巧。

AI 摘要

研究者引入仇恨事件检测任务,测试GPT-4o和Llama-3.2-3B-Instruct在多个专家标注数据集上的反犹事件分类能力。结果表明GPT-4o潜力较大但需显著改进。提供术语定义或上下文示例可提升性能:定义对修辞类事件(如经典反犹陈词)最有效,示例对行动类事件(如人身攻击)更有帮助。以大学校报为案例,LLM可辅助识别相关真实事件,支持早期监测。

原文 · arXiv: OpenAI

Evaluating Large Language Models for Antisemitic Incident Classification

Addressing hate and violence in society requires timely detection of hateful events from public reporting, but automated identification of hateful events remains underexplored. We introduce the task of hateful event detection and investigate the ability of AI systems, specifically large language models (LLMs), to discover and classify reports of antisemitic events with fine-grained labels. We evaluate OpenAI's GPT-4o and Meta's Llama-3.2-3B-Instruct on multiple expert-annotated datasets containing antisemitic event descriptions from news articles, civil society reports, and official records. We show that LLMs, particularly GPT-4o, have potential for this task, but substantial improvement is needed. Providing clear term definitions and in-context examples in prompts can improve performance: definitions are most helpful for rhetoric-oriented events (e.g. classical antisemitic tropes), while examples help label action-oriented events (e.g. physical assault). A case study of college newspapers demonstrates that LLMs can help surface relevant real-world events, supporting early monitoring and intervention. Overall, our findings highlight both opportunities and critical gaps in AI's ability to recognize complex harms and underscore the need for collaborative efforts among AI developers, policymakers, and civil society to design models, implement robust evaluation, and develop policy frameworks for defining and combating hate efficiently and effectively.