论文精选

一种基于信号理论的机器学习框架用于早期识别劳动剥削

Detecting Deceptive Recruitment: A Signal-theoretic Machine Learning Framework for Early Identification of Labour Exploitation

精选理由

这个研究挺有意思,用机器学习框架来识别劳动剥削,特别是招聘广告里的欺骗行为,挺实用的。

本文提出了一种多模态检测模型,结合计算机视觉、自然语言处理和语义嵌入,用于识别欺骗性的招聘广告。该模型在包含464个案例的数据集上进行了测试,其中164个是欺骗性的,300个是合法的。研究通过特征消融实验和分层交叉验证,证明了单个模态(ROC-AUC: 0.87--0.97)和集成模态的有效性。SHAP分析显示,文本质量和特定领域的风险语言是主要区分因素,包括可读性指数、风险关键词密度和签证赞助提及。这些生产质量差距反映了剥削者资源有限,无法在所有沟通渠道同时保持专业标准。

原文 · arXiv cs.LG

Detecting Deceptive Recruitment: A Signal-theoretic Machine Learning Framework for Early Identification of Labour Exploitation

Deceptive online job advertisements have emerged as a primary pathway into forced labour, yet systematic detection methods remain underdeveloped due to data scarcity and absence of empirically validated indicators. We formalise this detection challenge as a classification problem under signalling theory, where exploiters transmit costless signals mimicking legitimate communications across textual, visual, and structural dimensions. Using 464 verified cases (164 deceptive, 300 legitimate) collected through anti-slavery charities across nine origin countries and 21 industries, we develop multimodal detection models combining computer vision, natural language processing, and semantic embeddings. Through systematic feature ablation experiments and repeated stratified cross-validation, we demonstrate that individual modalities achieve substantial discriminatory power (ROC-AUC: 0.87--0.97), whilst their integration yields modest further gains. SHAP-based analysis reveals that text quality and domain-specific risk language are the primary discriminators, with readability indices, risk keyword density, and visa sponsorship mentions ranking highest, followed by visual colour and texture features. These production quality gaps reflect resource constraints that prevent exploiters from maintaining professional standards across all communication channels simultaneously. We operationalise findings through a proof-of-concept decision support system providing interpretable risk scores for practitioners. This work demonstrates how rigorous analytical frameworks can address complex humanitarian operations challenges characterised by information asymmetry and limited ground-truth data.