Qwen-Image 2.1 发布:单张 RTX 4090 生成 1024×1024 图像仅需 18.7 秒
Thanks @sgl_project for the day-0 support! 🙌 SGLang-Diffusion now serves Qwen-Image-2.1: text-to-im...
Qwen-Image 2.1 一张 4090 就能跑,18.7 秒出图还带透明底输出,SGLang 当天就给出了部署命令。
Alibaba Qwen 发布图像模型 Qwen-Image 2.1,SGLang-Diffusion 在发布当天提供支持。在单张 RTX 4090 24GB 上配合 CPU offload,1024×1024 文生图耗时 18.7 秒,图像编辑耗时 21.7 秒,峰值显存 22.7 GiB,且不做量化。换到 RTX PRO 6000 96GB 后,生成降至 8.0 秒,编辑降至 9.6 秒。单一 checkpoint 同时覆盖文生图、多图编辑和透明 RGBA 输出,原生推理支持 TP/SP 与 LoRA,提供 OpenAI 兼容 API,默认 40 步去噪。
Thanks @sgl_project for the day-0 support! 🙌 SGLang-Diffusion now serves Qwen-Image-2.1: text-to-im...
Thanks @sgl_project for the day-0 support! 🙌 SGLang-Diffusion now serves Qwen-Image-2.1: text-to-image generation, multi-image editing, and transparent RGBA output. Try it out! 🎨 SGLang @sgl_project Day-0 support for @Alibaba_Qwen ’s Qwen-Image 2.1 is here in SGLang-Diffusion! 🖥️ Native precision on a single RTX 4090 24GB with CPU offload - 1024×1024 generation in 18.7s and image editing in 21.7s with 22.7 GiB peak GPU memory during requests. - On an RTX PRO 6000 96GB: 8.0s generation and 9.6s editing. 🎨 Text-to-image, multi-image editing, and transparent RGBA output—all with one checkpoint. ⚡ Native inference with TP/SP, LoRA, and OpenAI-compatible APIs. 40 denoising steps, one image per request, warmed HTTP latency including PNG output. No quantization. Cookbook and GPU-specific commands below 👇 🔗 View Quoted Tweet 💬 7 🔄 4 ❤️ 68 👀 9300 📊 11 ⚡