论文

BrainWideBench:基于 139 只小鼠多脑区记录的跨动物迁移基准

BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings

精选理由

把 139 只小鼠 276 个脑区的神经记录做成基准,专测预训练模型跨动物迁移,做神经+AI 的能直接拿来跑。

BrainWideBench 基于 International Brain Laboratory 的 Brainwide Map 数据集构建,涵盖 139 只小鼠在感觉引导决策任务中的记录,覆盖 276 个脑区。基准设三个任务套件,分别考察学习到的表示能否解码动物行为、预测被掩码或未来的神经活动、以及恢复有生物学意义的解剖结构。作者用它在下游微调和零样本迁移到未见动物等设定下系统评测了多种预训练方法。结果显示预训练优于匹配的单会话基线,但收益高度依赖预训练目标与下游任务的对齐程度,没有一种方法在三个套件上全面占优。

原文 · arXiv cs.LG

BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings

Advances in large-scale neural recording have made it possible to collect data across many animals and distributed brain regions, raising the question of whether this scale can be exploited to learn general-purpose neural representations transferable across diverse downstream tasks. Yet, progress toward this goal has been limited by fragmented evaluation protocols and a narrow focus on individual task domains. Here, we present BrainWideBench, a benchmark for evaluating across-animal transfer on multi-region neural recordings, built on the International Brain Laboratory Brainwide Map dataset of neural and behavioral recordings spanning 276 brain regions from 139 mice performing a sensory-guided decision-making task. The benchmark is organized around three complementary task suites that evaluate whether learned representations support downstream decoding of behavior, can predict masked or future neural activity, and can recover biologically meaningful anatomical organization. With this benchmark, we systematically evaluate pretraining methods across transfer settings, including finetuning on downstream objectives and zero-shot generalization to unseen animals. Our results confirm pretraining improves performance over matched single-session baselines, but we show current methods exhibit heterogeneity in transfer capabilities: gains depend strongly on the alignment between pretraining objectives and downstream tasks. No single approach performs uniformly well across all three suites, and most methods are designed to only address a subset of them. Together, these findings suggest that learning representations that jointly generalize across behavior, dynamics, and anatomy remains an open challenge. By providing a unified and reproducible evaluation suite, BrainWideBench establishes a framework for measuring progress toward general-purpose models of the mouse brain.