Cerebras 让万亿参数模型跑出千 token 每秒
Cerebras 正在企业测试中运行 Kimi K2.6,这是一个万亿参数模型。据 Artificial Analysis 测量,其推理速度约为每秒1000个 token,是迄今最快的前沿模型性能。这反驳了此前认为开源大模型无法快速运行的质疑。
I remember when people were saying "It's useless to open-source big models because nobody will be ab...
I remember when people were saying "It's useless to open-source big models because nobody will be able to run them fast".... Cerebras @cerebras Cerebras is now running Kimi K2.6 – a trillion parameter model – in enterprise trials. At ~1,000 tokens/s, this is the fastest frontier model performance ever measured by Artificial Analysis @ArtificialAnlys . 🔗 View Quoted Tweet 💬 6 🔄 3 ❤️ 75 👀 5213 📊 10 ⚡