论文

NeuronEye框架提升视觉语言模型推理能力

NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning

精选理由

研究人员提出NeuronEye框架,让视觉语言模型能根据查询激活相关视觉概念,在多个基准测试中显著提升推理性能。

NeuronEye是一种新型插件框架,从Qwen2.5-VL-7B中间表示中构建稀疏概念级神经元词汇表。该框架在CV-Bench基准测试中将整体准确率提高3.1个百分点,距离任务提升9.5个百分点,BLINK Multi-view提升8.3个百分点。类似趋势也出现在LLaVA-1.6-7B模型上。

原文 · arXiv cs.AI

NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning

Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and modulate the specific visual evidence required by a given language query. Inspired by sparse population coding and top-down modulation in biological vision, we introduce NeuronEye, a plug-in framework that constructs a sparse, concept-level neuron vocabulary from intermediate VLM representations and selectively activates query-relevant visual concepts during inference. NeuronEye decomposes vision-token states into an overcomplete sparse basis organized by concept-level clusters, uses the language query to activate relevant clusters and localize the patches where selected concepts are expressed, and injects the focused evidence back into vision tokens. A complementary suppression mechanism attenuates dominant perceptual directions to preserve weaker but relevant cues. All operations run in a single forward pass over a frozen VLM backbone. On Qwen2.5-VL-7B, NeuronEye raises CV-Bench overall accuracy by +3.1 with gains of +9.5 on Distance, and improves BLINK Multi-view by +8.3, with similar trends on LLaVA-1.6-7B. These results suggest that sparse neuron vocabularies can serve not only as post-hoc interpretability tools but also as active interfaces for concept-level visual reasoning.