Jeff Dean 亲自梳理了 Google Translate 从统计方法到神经网络的两次关键跃迁,做 NLP/翻译系统的开发者能从中看到技术选型的真实演进逻辑,值得一读。
Google Translate 迎来20周年,Jeff Dean 回顾了其关键里程碑:2006年首次部署基于5-gram语言模型的系统,使用了万亿词级训练数据,是早期大语言模型实践;2016年转向深度神经网络,结合序列到序列模型和自研TPU,推理性能提升30-80倍,延迟降低15-30倍,使服务可覆盖数亿用户;近期又借助Gemini模型进一步优化。这些技术迭代持续提升了翻译质量和全球连接性。
Google Translate is turning 20! 🎉. There are 20 fu…
Google Translate is turning 20! 🎉. There are 20 fun facts and tips in the thread below.
Translate is one of my favorite Google products because it brings us all closer together!
I've been involved with a couple of things over the years. The first was our deployment of the initial system in 2006, which provided a huge leap forward in quality because it used a much larger 5-gram language model trained on trillions of words of text (indeed, probably the first trillion token language model training in the world: paper has some nice heads showing scaling-law-like quality improvement from scaling to more data/compute).
See "Large Language Models in Machine Translation", Thorsten Brants, Ashok C. Popat, Peng Xu, Franz J. Och and Jeffrey Dean, https://t.co/QnK7lllpoj
The second major collaboration was in 2016 when we moved Translate over from a statistical machine translation approach to using deep neural networks. This approach relied on two key innovations. The first was Google's work on Sequence-to-Sequence models (https://t.co/W9c0a0PXoV). The second was our development of TPUs, custom cups that improved the performance of inference for deep neural networks by 30-80X over existing CPUs and GPUs of the day (and reduced latency by 15-30X). This made launching compute-intensive language model services like Translate feasible for hundreds of millions of users. See "In-Datacenter Performance Analysis of a Tensor Processing Unit", Norman P. Jouppi et al. https://t.co/qpJl7FM6EO
GNMT paper: "Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation", Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean, https://t.co/YasV0MEpxM
Most recently, we have advanced Translate further using Gemini models.
Each of these advances relied on research that have major quality leaps over the existing status quo translation approaches, bringing better quality and connectedness to all of our Translate users! 🎉