语言指导的嗅觉表征学习SCENT:用语义桥接视觉与气味

What Images Cannot Say: Language-Guided Olfactory Representation Learning

精选理由

想用语言帮AI'闻'气味?这篇论文用VLM生成场景描述,在纽约气味数据集上检索效果最牛,还能拆解混合气味。

AI 摘要

SCENT框架利用视觉-语言模型(VLM)生成场景描述,捕捉对象、环境和可能的气味线索。该方法训练气味编码器将电子鼻信号映射到共享嵌入空间,并对齐视觉与文本表征。在New York Smells数据集上,SCENT在气味-图像和气味-文本检索任务中达到SOTA。框架还通过语言引导的潜在分解,分离对象气味与背景环境贡献。实验证明了语义信息对多模态嗅觉感知的重要性。

原文 · arXiv cs.LG

What Images Cannot Say: Language-Guided Olfactory Representation Learning

Images tell us what a scene looks like, but rarely what it would feel like to be there. While recent datasets pair visual scenes with electronic-nose measurements, aligning smell signals with images remains challenging because many olfactory cues arise from contextual environmental factors that are not directly visible in pixels. We introduce SCENT, a multimodal framework that uses language guidance as a semantic bridge between vision and olfaction. Our approach leverages Vision-Language Models (VLMs) to generate scene descriptors capturing objects, environmental context, and plausible ambient smell cues suggested by the visual scene. These descriptors provide semantic guidance for learning olfactory representations. We train a smell encoder that maps electronic-nose signals into a shared embedding space aligned with both visual and textual representations, and introduce a languageguided latent decomposition that separates object-specific odors from contextual environmental contributions. Experiments on the New York Smells dataset demonstrate that SCENT significantly improves crossmodal retrieval compared to vision-only baselines, achieving state-of-theart performance on smell-to-image and smell-to-text retrieval tasks. In addition, our framework produces interpretable olfactory representations that enable the disentanglement of complex smell mixtures. Our results reveal the importance of contextual semantic information for grounding olfactory perception in multimodal learning and pave the way for future research in this area.