论文

MAMHOI:通过功能分解场景感知人机交互

MAMHOI: Factorizing Scene-Aware Human-Object Interaction through Affordances

精选理由

MAMHOI解决了人机交互数据稀缺问题,通过功能分解实现更真实的三维场景交互生成。

MAMHOI是一种新型人机交互生成模型,通过显式的运动-功能接口分解场景感知HOI生成过程。该模型首先预测交互可行性,然后生成相应的人机运动。在复杂室内环境中测试,MAMHOI减少了物体-场景穿透问题,同时保持了人机交互质量,生成了更真实且物理可行的场景感知交互。

原文 · arXiv cs.AI

MAMHOI: Factorizing Scene-Aware Human-Object Interaction through Affordances

Generating realistic human-object interactions (HOI) in complex 3D scenes requires two complementary capabilities: reasoning about interaction feasibility in the environment and synthesizing realistic human-object motion. However, supervision for these capabilities is rarely available jointly at scale. Human-scene datasets provide rich information about environment-aware motion, while human-object datasets capture detailed interaction dynamics, yet paired human-object-scene data remain scarce. We present MAMHOI, an affordance-mediated factorization for scene-aware human-object interaction generation. MAMHOI factorizes scene-aware HOI generation through an explicit motion-affordance interface between scene understanding and motion synthesis: a scene-conditioned model first predicts where and how an interaction can be feasibly executed, and an affordance-conditioned HOI model then generates the corresponding human-object motion. This factorization allows scene understanding and interaction dynamics to be learned from complementary sources of supervision without requiring paired human-object-scene data. Experiments in complex indoor environments show that MAMHOI reduces object--scene penetration while better preserving human--object interaction quality, yielding more realistic and physically feasible scene-aware interactions. Project page: https://leimingyuan.github.io/MAMHOI-project-page/