遗忘语言在神经语音模型中的痕迹

Lost but not erased: Finding traces of a forgotten language in neural speech models

精选理由

这篇论文揭示了语言习得中的关键期效应,通过神经语音模型的研究,让我们更深入地理解了语言习得的过程。

AI 摘要

国际收养儿保留出生语言的语音痕迹,研究发现神经语音模型在训练第二语言时,第一语言的痕迹主要存在于底层,这些痕迹有助于模型更快地重新学习第一语言,表明关键期效应反映的是基础表征的根深蒂固,而非成熟期塑性损失,经验在语言习得的关键期中起核心作用。

原文 · arXiv cs.LG

Lost but not erased: Finding traces of a forgotten language in neural speech models

International adoptees retain phonological traces of a birth language they can no longer speak or comprehend, a persistence typically attributed to a biologically-timed critical period. We asked whether it could instead reflect the ordinary dynamics of learning, using automatic speech recognition models that simulate the international adoptee experience without maturational confounds. Models were trained on one language and then abruptly switched to a second. We found that traces of the first language persisted throughout second-language training, but mainly in the lowest, pre-phonemic layers. These traces were functional, as models with early exposure re-learned their lost first language 14% faster than naive models; this advantage held even against models adopted early from a related language and disappeared when the earliest layers were substituted from a non-adopted model. We argue that these critical-period effects reflect entrenchment of foundational representations rather than a maturational loss of plasticity, and that experience plays a central role in critical periods in language acquisition.