TerraNova把地球和社会数据一起训练,能几分钟适配新变量,比地理模型多覆盖海洋和不确定性。
TerraNova 是一个面向人类世的地球与社会耦合基础模型,在 arXiv 2607.29527 发布。它用 1,024 条记录训练,其中 512 个格点化地球系统场和 512 个国家指标。跨模态 Transformer 融合位置、国家、时间和任务,两个对比目标分别对齐国家与境内坐标、预训练地理空间嵌入。TerraNova 在消费级硬件上数分钟内即可重构稀疏观测的密集场并适配未见变量。与专用地理空间编码器相比,TerraNova 还覆盖时间、海洋和不确定性维度。
TerraNova: A Foundation Model for the Anthropocene
A defining problem of the Anthropocene is to model the physical Earth and human societies as one coupled system, yet no learned representation spans their observational breadth. We argue the obstacle is geometric: the physical Earth is measured as continuous fields that ignore political borders, whereas societies are reported for administrative units. Earth-system foundation models serve the first geometry; coupling it to the second has required lossy averaging over borders. We introduce TerraNova, a foundation model trained on 1,024 physical and societal records in their native geometries: 512 gridded Earth-system fields and 512 national indicators. Dedicated encoders represent location, country, time and task, cross-modal transformers fuse them into a shared spatiotemporal state, and a hypernetwork generates a per-query decoder whose evidential head returns a predictive distribution. Two contrastive objectives couple the representation: a population-weighted alignment between each country and coordinates in its territory, and one to pretrained geospatial embeddings carrying image-derived semantics. Read out through that decoder, the representation is competitive with purpose-built geospatial encoders while spanning axes they do not represent (time, oceans and uncertainty) and supporting country-level capabilities. The frozen backbone reconstructs dense fields from sparse observations and adapts to unseen variables in minutes on consumer hardware.