腾讯混元Hy4预览版已适配vLLM,支持MoE和HPC-Ops内核,可直接部署使用。
腾讯混元Hy4预览版可在vLLM环境中运行,已在NVIDIA GPU上验证。该模型总参数770B,激活参数49B,包含256个路由专家和一个共享专家。上下文长度为1M,但每个查询仅关注2048个token。模型中仅21层计算稀疏索引,其余57层复用同一索引。
Hy4 preview runs in vLLM
Hy4 preview runs in vLLM vLLM @vllm_project @TencentHunyuan 's Hy4-preview runs in vLLM from day 0, verified on NVIDIA GPUs. 🎉 - 770B total, 49B active, 256 routed experts plus one shared - 1M context, but each query attends to just 2048 tokens - Only 21 of the 78 layers compute their own sparse index, the other 57 reuse one - A 10B MTP layer ships inside the checkpoint, 0.7B of it active, draft depth 3 Tencent's HPC-Ops attention and MoE kernels have been in vLLM main since Hy3. VLLM_ENABLE_HPC_OPS=1 vllm serve tencent/Hy4-preview-FP8 -tp 8 Thanks @TencentHunyuan for the preview weights! 🙌 recipes.vllm.ai/tencent/Hy4-pr… yZb 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 20 👀 1090 📊 2 ⚡
- NVIDIA AI08-28 21:03原文