这篇论文教你如何不让合成临床数据看起来像模板:两个修订方法在保住效用下限的同时把真实感拉上来,适合做医疗AI基准的人参考。
论文提出在效用约束下改进合成临床基准的真实性。基线基准的配对缺失率达79.44%,仅12.75%的行可操作,38.94%的患者无可操作指标。两种确定性修订方案在保持效用下限的同时改善了上述指标。与朴素加密控制相比,修订避免了不现实的模板化。研究还发现内部真实性与对聚合操作参考的源保真度是不同目标。
Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints
Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access. We study how to improve such benchmarks without breaking the downstream utility checks already used in practice. We formulate benchmark revision as utility-constrained realism improvement: dataset changes should increase realism while staying above an operational utility floor. We instantiate this idea on a care-gap benchmark derived from Synthea-generated patients exercised through demonstration electronic health record workflows and then processed by the same downstream pipeline as operational data. Realism is measured through missingness structure, simplicity, structural plausibility, and population alignment. The baseline benchmark is extremely thin: sampled-pair missingness is 79.44%, only 12.75% of rows are actionable, 38.94% of patients have zero actionable measures, and top-three token concentration reaches 100.0%. Two deterministic revisions improve these panels while remaining above the current utility floor, whereas a naive densification control preserves unrealistic templating. We further show that internal benchmark realism and source fidelity to an aggregate operational reference are related but distinct objectives. These results suggest that synthetic benchmark quality should be optimized explicitly, with utility treated as one constraint rather than as sufficient evidence of realism.