本地跑大模型?这份DGX Spark指南从1台到4台,告诉你选哪款模型、速度多少。
NVIDIA AI官方转发了一份来自MiaAI_lab的DGX Spark本地模型运行指南。1台DGX Spark可跑DeepSeek v4 Flash(1M上下文,26 tok/s)或Qwen 3.6 35b NVFP4(256k上下文,81 tok/s)。2台DGX Spark是性价比甜点,DeepSeek v4 Flash速度提升到82 tok/s,还能跑Inkling-Small等全模态模型。3台以上可跑GLM-5.2 with Vision,348k上下文、25 tok/s,被称为本地最强智能。指南还给出了4台时的多卡分工方案,比如2台跑DeepSeek v4 Flash,另外2台跑图像模型和ComfyUI。
The time for local AI is here. So many great models have been released recently, here's a guide fr...
The time for local AI is here. So many great models have been released recently, here's a guide from @MiaAI_lab on chaining DGX Sparks together to run them. Drop a pic of your setup and let us know what you're running 👇 Mia @MiaAI_lab What are the best models you can run on your @NVIDIAAI DGX Spark? ✨ Aug 2026 Edition 1× DGX Spark • DeepSeek v4 Flash - 1M ctx, 26 tok/s - recommended! • Qwen 3.6 35b NVFP4 - 256k ctx, 81 tok/s • Qwen 3.6 27b NVFP4 - 256k ctx, 33 tok/s Qwen 3.8 27b - should be released soon and this could change my recommendation! 2× DGX Sparks ← sweet spot! • DeepSeek v4 Flash 0731 - 1M ctx, 82 tok/s • Inkling-Small - 1M context, Full Omni, 33 tok/s • MiMo-V2.5 - 1M ctx, Full Omni, 31 tok/s • Step-3.7-Flash — 256K ctx, 30 tok/s 3× DGX Sparks • GLM-5.2 with Vision - 348k context, 25 tok/s - still the best intelligence you can run locally if you have 3 sparks. • DeepSeek v4 Flash on 2 units + smaller models on the 3rd spark for images, ComfyUI and other things. This setup gives you DeepSeek v4 Flash speeds for coding, plus strong agentic workflows and image support from smaller models. 4× DGX Sparks • GLM 5.2 NVFP4 across all 4 units - still think this is the way to go if you have 4 units! • The alternative is running DeepSeek v4 Flash on 2 units and any other 2x setup on the other units. The choice is your. Links and repos below 👇 🔗 View Quoted Tweet 💬 4 🔄 4 ❤️ 43 👀 3405 📊 7 ⚡