UniCAD:统一多模态多任务CAD基准与通用模型

UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD

精选理由

CAD 研究者终于有了统一的多模态基准和通用模型,做3D设计、CAD生成或问答的团队可以直接用 UniCAD-MLLM 替代多个专用模型,建议关注开源资源。

AI 摘要

UniCAD 是一个面向计算机辅助设计(CAD)的多模态学习基准,涵盖点云到CAD重建、文本/图像到CAD生成以及CAD问答等任务。同时提出的 UniCAD-MLLM 是一个通用多模态大语言模型,能端到端处理文本、图像、草图和点云,在单一框架内完成异构任务。实验表明,UniCAD-MLLM 在 UniCAD 和 Fusion360 基准上均达到最先进水平,超越现有任务专用和多任务基线。该工作填补了CAD领域缺乏统一多模态基准的空白,将开源数据集、代码和预训练模型。

原文 · arXiv cs.AI

UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD

Computer-Aided Design (CAD) underpins modern engineering and manufacturing by enabling the creation of precise, editable 3D models. However, CAD research typically studies tasks in isolation, and multi-modal, multi-task learning for CAD is hindered by the absence of a unified benchmark. To address this gap, we introduce UniCAD, a comprehensive benchmark for multi-modal CAD learning that covers point-to-CAD reconstruction, text/image-to-CAD generation, and CAD question answering across diverse input modalities. Alongside the benchmark, we present UniCAD-MLLM, a universal multi-modal large language model that ingests text, images, sketches, and point clouds and performs these heterogeneous tasks in an end-to-end fashion within a single framework. Extensive experiments on the UniCAD and Fusion360 benchmarks demonstrate that UniCAD-MLLM achieves state-of-the-art performance across all tasks, outperforming existing task-specific and multi-task baselines. We will release the dataset, code, and pretrained models to accelerate future research.