论文提出全局平均精度 gAP 与可微替代损失 gSAP
Global Average Precision for Representation Learning
一篇写损失函数的论文:gSAP 直接换掉 InfoNCE,同样精度下能多捞 4 倍正样本,做检索和预训练的可以试试。
论文提出 Global Average Precision(gAP)指标,把所有 query-candidate 对放进同一个列表排序并计算单一 AP,区别于逐 query 评估的 mAP。作者同时推出可微替代损失 gSAP,输入只需相似度矩阵和正对标记矩阵,与 InfoNCE 等现有损失输入一致,可直接替换。gSAP 在 batch 内联合考虑所有成对比较,因此在低温度区间仍可训练,而逐 query 替代损失此时梯度耗尽。在监督式度量学习、跨模态对齐和自监督预训练三类任务中,gSAP 是已知首个替换 InfoNCE 的排序损失,同等精度下检索到的正对数量最高达最强 AP 替代损失的 4 倍。
Global Average Precision for Representation Learning
Standard information retrieval metrics, such as mean Average Precision (mAP), assess performance one query at a time, based on how the similarities between a query and its positives compare against those with its negatives. The same holds for common representation learning losses, such as InfoNCE and per-query AP surrogates. None of them considers whether similarities are comparable across queries, which any system with a single decision threshold relies on. Global Average Precision (gAP) does, by ranking all query-candidate pairs in one list and computing a single AP. We introduce gSAP, a differentiable surrogate of gAP. It needs only a similarity matrix and a binary matrix marking the positive pairs, the same input as existing losses, so it is a drop-in replacement for them and agnostic to the encoder, the modality, and the source of supervision. Since it considers all possible pairwise comparisons in the batch jointly, it also remains trainable at low temperatures, a regime where per-query surrogates run out of gradient. Swapping it into established recipes improves supervised metric learning, cross-modal alignment, and self-supervised pretraining, where, to our knowledge, it is the first ranking loss to replace the community standard InfoNCE in the latter two. Its similarities are more consistent across queries, which drives the gains under a universal threshold. gSAP retrieves up to four times as many positive pairs as the strongest AP surrogate at the same precision, and it degrades the least when queries with no positives in the database are added. Beyond thresholding, models trained with gSAP also learn better representations, with higher transfer, $k$NN and zero-shot classification accuracy.