80张RTX 5090运行Kimi K3实现20 tok/s

这哥们儿拿 5090 跑 Kimi K3,总共用了 80 张卡,速度是 20 tok/s,其实可以用了。 然后我基于他的内容和官方的一些信息,重新让 Codex 做了一张图,展示在不同的显卡平台上跑...

精选理由

80张游戏卡就能跑2.8T参数的Kimi K3,20 tok/s挺实用,比之前GLM-5.2快了好几倍。

AI 摘要

开源模型Kimi K3(2.8T参数)在80张RTX 5090上以单流20 tok/s运行。该配置使用GDDR7游戏显卡和普通以太网,无需HBM,权重为官方MXFP4格式。此前同一集群运行GLM-5.2时从30 tok/s优化至110 tok/s。任何实验室或大学均可使用该配置复现。

原文 · 歸藏(guizang.ai)

这哥们儿拿 5090 跑 Kimi K3,总共用了 80 张卡,速度是 20 tok/s,其实可以用了。 然后我基于他的内容和官方的一些信息,重新让 Codex 做了一张图,展示在不同的显卡平台上跑...

这哥们儿拿 5090 跑 Kimi K3,总共用了 80 张卡,速度是 20 tok/s,其实可以用了。 然后我基于他的内容和官方的一些信息,重新让 Codex 做了一张图,展示在不同的显卡平台上跑 Kimi K3 需要多少张卡和对应的带宽。 Ning @totheagi we got the full Kimi K3, 2.8T params, running on 80x RTX 5090s. 20 tok/s single stream, day one, untuned. Last week we took GLM-5.2 from 30 to 110 tok/s on this same fleet. This number will climb. A first for open weights: frontier intelligence served with zero HBM, the scarcest silicon in AI. Just GDDR7 gaming cards, plain ethernet, and the official MXFP4 weights, nothing requantized. The most powerful open model on Earth, on the most abundant GPUs on Earth. Any lab, startup, or university can now own it, probe it, fine-tune it, run agents on it. @Kimi_Moonshot 🔗 View Quoted Tweet 💬 7 🔄 0 ❤️ 6 👀 4802 📊 6 ⚡

80张RTX 5090运行Kimi K3实现20 tok/s · AI 热点