商汤的架构创新解决了多模态模型常见的模块间信息丢失问题,做视觉内容生成或信息图设计的团队可以直接用这个开源模型,生成效率翻倍值得一试。
中国 AI 实验室商汤开源了 SenseNova U1,这是一个统一的多模态模型,能在单一模型中理解、推理并生成图像和文本。其架构去除了传统的视觉编码器和变分自编码器,在共享表示空间中处理图像和语言,减少了模块间切换和信息损失,提升了生成一致性。该模型在生成信息图、指南、海报、漫画等密集视觉内容时表现出色,据客户基准测试,生成信息图的速度约为 Qwen-Image-2.0 / Seedream-4.5 的两倍,且质量相当。
Chinese AI lab SenseTime just open-sourced SenseNo…
Chinese AI lab SenseTime just open-sourced SenseNova U1, a unified multimodal model that can understand, reason, and generate images + text inside 1 model.
The interesting part is the architecture: it removes the usual visual encoder and variational auto-encoder setup, then handles image and language inside a shared representation space, instead of being passed between separate modules.
That means less handoff between modules, less information loss, and better consistency when creating dense visual content like infographics, guides, posters, comics, and image-text workflows.
That’s how the model can generate coherent text and images together in one flow, which is why it is strong for infographics, guides, comics, posters, and step-by-step visual content.
For infographic generation specifically, it is also around 2x faster than Qwen-Image-2.0 / Seedream-4.5 while staying in the same rough quality band, based on the client benchmark chart. 1/n