Qwen模型升级,带来更高效的注意力机制和优化器,提升模型性能和稳定性。和同类模型相比,Qwen在处理长序列时更具优势。
Qwen模型进行四大升级,包括GDN+QSA混合注意力机制、Gated Residual增强信息流、N-gram Embedding扩展容量和Muon优化器提升效率。
Model Architecture Four core upgrades for maximum capability, efficiency, capacity, and stability: ...
Model Architecture Four core upgrades for maximum capability, efficiency, capacity, and stability: - Attention: GDN + QSA Hybrid. Gated DeltaNet (GDN) compresses history. Qwen Sparse Attention (QSA) uses a lightweight indexer for micro-block context selection. Lower the cost of attention on long sequences. - Residual: Gated Residual (GR) widens the residual stream to 4 branches with a dynamic read and write gating, strengthening cross-layer information flow and significantly improving training stability. - Embedding: N-gram Embedding uses local context lookups to expand model capacity at minimal compute cost, while keeping the embedding table in host memory with asynchronous prefetching. - Optimization: Muon optimizer. Refines Muon through improved orthogonalization, smarter parameter assignment between Muon and AdamW, and fused-parameter splitting, with scaling laws refitted for the new architecture. 💬 4 🔄 31 ❤️ 349 👀 57074 📊 53 ⚡