Percy Liang的实验室以开放的方式发布了Marin 535B-A23B模型,展示了AI训练的开放性,值得关注。
Marin项目展示了AI模型训练的开放性,包括开源代码、数据、配方和实验结果。本周开始训练Marin 535B-A23B模型,整个训练过程公开。预训练和中期训练将在11个GB200 NVL72上使用18.75T标记进行,预计约3个月完成。之前进行了从1.6B-A61M到27.7B-A1.2B的4级缩放训练,以调试问题和预测主要运行。这是迄今为止最大的运行,预计会有意外。
In the fight to defend openness in AI, the Marin project is a precious demonstration of openness in ...
In the fight to defend openness in AI, the Marin project is a precious demonstration of openness in model training, with open code, data, recipes, even experimental results. Releasing AI research openly used to be the norm; I'm grateful for @percyliang 's open lab approach. Percy Liang @percyliang 🚢 Marin 535B-A23B started training this week! As usual, the whole process is open. Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow. Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected. 🔗 View Quoted Tweet 💬 12 🔄 6 ❤️ 51 👀 8668 📊 12 ⚡
- Stanford AI Lab08-23 15:46原文