腾讯Hy4-preview模型压缩至200GB GGUF

We compressed Hy4-preview from 1.5TB to ~200GiB GGUF and it still works well ! Meet MIX-STQ1_0.The...

精选理由

腾讯把1.5TB的Hy4-preview压缩到200GB还保持性能,MIX-STQ1_0技术真厉害

AI 摘要

腾讯将Hy4-preview模型从1.5TB压缩至约200GB的GGUF格式,模型性能保持良好。该模型采用MIX-STQ1_0技术,通过校准数据为每层选择不同比特宽度,最低达1.31-bit STQ1_0,最高至2.06-bit IQ2_XXS。在MCP Atlas基准上,模型得分从83.7降至83.2;在SWE-Bench multi上从82.9降至81.3;在MRCR上从81.3降至81.1;在IFBench上从73.5降至72.5。

原文 · Hunyuan

We compressed Hy4-preview from 1.5TB to ~200GiB GGUF and it still works well ! Meet MIX-STQ1_0.The...

We compressed Hy4-preview from 1.5TB to ~200GiB GGUF and it still works well ! Meet MIX-STQ1_0.The trick isn’t just going low, it’s deciding where: calibration data picks each layer’s bit-width, some down to 1.31-bit STQ1_0, some up to 2.06-bit IQ2_XXS. Same budget, lower error. Accuracy barely moves vs BF16 📊 MCP Atlas 83.7→83.2 📊 SWE-Bench multi 82.9→81.3 📊 MRCR 81.3→81.1 📊 IFBench 73.5→72.5 See the details on HF : AngelSlim/Hy4-preview-GGUF Weights & low-bit GGUF huggingface.co/AngelSlim/Hy4-… T9 #LLM #Quantization a #llamacpp m #Hy p #Hy Tencent Hy @TencentHunyuan 🚀 Hy4 preview is here. 770B, 49B active, 1M context. Built for productivity. Open source frontier. Consistent affordable price. Use it. Tell us what breaks. More on Hy blog hy.tencent.ai/research/hy4-p… C HuggingFace huggingface.co/tencent/Hy4-pr… R Github github.com/Tencent-Hunyua… L 🔗 View Quoted Tweet 💬 13 🔄 15 ❤️ 158 👀 8424 📊 35 ⚡