多模态可穿戴传感器融合检测重复性身体聚焦行为

Deep Multimodal Wearable Sensor Fusion for Detection of Body-Focused Repetitive Behaviors

精选理由

这个研究用腕带传感器加深度学习,检测拔毛、抠皮这类小动作,F1到0.985,比单模态准多了,对心理健康监测挺有参考价值。

AI 摘要

Child Mind Institute使用Helios腕戴设备收集数据,结合惯性测量单元、热电堆传感器和飞行时间传感器,捕捉运动、热感和距离信息。研究团队开发了结合卷积神经网络和门控循环单元的深度学习框架,并加入模态专用自编码器和晚期融合分类器。该框架在二元检测中F1分数达0.985,AUC为0.997,九类分类中宏平均F1为0.700,AUC为0.963。Shapley解释显示飞行时间和惯性模态主导判别力,层级聚类表明误分类主要由手势解剖区域驱动。研究证明多模态融合可实现准确、客观、连续的行为监测,为可穿戴辅助心理健康诊断奠定基础。

原文 · arXiv cs.LG

Deep Multimodal Wearable Sensor Fusion for Detection of Body-Focused Repetitive Behaviors

Body-focused repetitive behaviors, such as hair pulling and skin picking, are compulsive motor actions commonly associated with obsessive-compulsive and anxiety disorders. Their early, objective detection remains difficult because the movements are subtle and overlap with ordinary, non-pathological gestures. We developed and evaluated a multimodal deep learning framework to detect and classify these behaviors from wrist-worn sensor data. The data, collected by the Child Mind Institute using the Helios wrist-worn device, combine inertial measurement units, thermopile sensors, and time-of-flight sensors, capturing kinematic, thermal, and proximity information. The framework combined a convolutional neural network with a gated recurrent unit, alongside modality-specific autoencoders and a late-fusion classifier, to exploit temporal and spatial dynamics. It achieved an F1 score of 0.985 and an area under the receiver operating characteristic curve of 0.997 for binary detection, distinguishing these behaviors from other activities, and a macro-averaged F1 score of 0.700 with an area under the curve of 0.963 across a nine-class scheme that distinguished each individual behavior from a single grouped Non-Target class, improving over single-modality baselines. Post-hoc interpretability based on Shapley additive explanations showed that the time-of-flight and inertial modalities dominated discriminative power by capturing spatial proximity and dynamic movement, while hierarchical clustering indicated that misclassifications were driven primarily by the anatomical region of the gesture. These findings demonstrate that multimodal sensor fusion enables accurate, objective, and continuous behavioral monitoring. This work establishes a foundation for real-time, wearable-assisted mental health diagnostics and personalized interventions in biomedical research and clinical care.