匈牙利语 ASR 研究者终于有了更大规模的对话数据集——BEA-Dialogue+ 将可用训练数据从85小时扩展到200小时,做低资源语言语音识别的团队可以直接用于模型评估和微调。
匈牙利语对话自动语音识别(ASR)因公开对话式训练数据有限而受限。BEA-Dialogue 语料库虽填补了空白,但其严格的说话人分离划分导致可用数据仅85小时。本文提出扩展版 BEA-Dialogue+,放宽划分标准,保留主要说话人完全分离,将可用数据增至200小时。研究评估了 Whisper 和 FastConformer 模型,发现更大语料库对未微调模型更具挑战性,而基于序列化输出训练(SOT)的微调在词错误率、字符错误率等指标上持续提升。该语料库为匈牙利语对话 ASR 提供了更大且更具挑战性的基准。
Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus
Conversational automatic speech recognition in Hungarian is constrained by the limited amount of publicly available dialogue-style training data. The BEA-Dialogue corpus addresses this need, but its strictly speaker-disjoint train/dev/eval split reduces the usable material to only 85 hours. In this paper, we introduce BEA-Dialogue+, an expanded version of the corpus that relaxes the split criterion for experimenters and dialogue partners while preserving complete separation of the primary speakers. This results in 200 hours of transcribed natural conversations and enables a controlled study of the trade-off between additional training data and speaker overlap across the splits. We evaluate several Whisper- and FastConformer-based models on both corpus versions, including Serialized Output Training (SOT)-based fine-tuning for dialogue transcription. Our results show that the larger corpus is more challenging for models without fine-tuning, whereas SOT-based adaptation yields consistent improvements in WER, CER, cpWER, and cpCER. Overall, BEA-Dialogue+ provides a substantially larger yet still demanding benchmark for Hungarian dialogue ASR, and a practical resource for training and evaluating dialogue transcription systems.