论文

Quantile-TabDDPM:用分位数正则化扩散模型应对高不平衡表格数据

Usefulness of Quantile-Aware Diffusion Modeling for Highly Imbalanced Tabular Data

精选理由

一篇针对表格数据类别不平衡的扩散模型改进论文,在信用卡欺诈数据上验证,做金融风控或少数类生成的朋友可以看看思路。

论文提出 Quantile-TabDDPM,在标准扩散模型的二次误差损失之外加入分位数损失项,构成分位数正则化的去噪目标,用于捕捉少数类中的稀有极端值。标准扩散模型的二次损失对高度偏斜和重尾数据缺乏结构敏感性,难以保留少数类细节。作者在真实世界信用卡交易数据集上评估,该数据集具有极端类别不平衡特征。结果显示该方法能有效生成合成数据,改善欺诈检测这类高不平衡场景下的分类效果。

原文 · arXiv cs.LG

Usefulness of Quantile-Aware Diffusion Modeling for Highly Imbalanced Tabular Data

Classification problem in the context of highly imbalanced data is a major challenge in many real-world applications (e.g., FinTech, healthcare, etc.). In these cases, the vast majority of instances belong to a single class and a small fraction represent the minority class (often the most critical class). Recently, diffusion models have emerged as powerful approaches to reduce the degree of ``imbalanced-ness'' in the dataset; they work by generating synthetic data by capturing complex data distributions using iterative transformations. However, standard diffusion models are not inherently suited to highly skewed or heavy-tailed data, due to inbuilt quadratic error loss, which lacks the structural sensitivity to capture rare, extreme values, and minority-class nuances. We propose a novel approach, namely, Quantile-TabDDPM, based on a quantile-regularized denoising objective that combines the standard quadratic error loss with a quantile loss term to explicitly capture rare events while preserving the theoretical grounding of the original denoising objective. We extensively evaluated our approach on a real-world credit card transaction dataset characterized by extreme class imbalance. The results demonstrate that the integration of diffusion-based synthetic data generation with a quantile-regularized denoising objective provides a robust and effective framework for fraud detection in highly imbalanced datasets.