论文多源确认精选

黑盒大模型秘密提取框架研究

Practical Secrets Extraction against Black-box LLMs

精选理由

研究人员发现黑盒大模型可能泄露训练数据中的机密信息,提出新框架能从API接口提取出真实密钥。

研究人员提出针对商业API大模型的黑盒秘密提取框架。该框架包含交叉验证秘密知识蒸馏和代理引导秘密提取两部分。在API密钥基准测试中,该框架提高了恢复有效率和真实密钥率,同时降低了提取延迟。研究团队从OpenAI和Claude Code三个独立部署的黑盒系统中成功恢复了被屏蔽的提供商特定凭证。

原文 · arXiv: OpenAI

Practical Secrets Extraction against Black-box LLMs

Large language models (LLMs) increasingly power autonomous coding agents such as Codex and Claude Code, yet their training corpora may contain confidential credentials exposed in public repositories or collected from private development artifacts, creating risks of memorization and subsequent leakage. Existing extraction audits, however, largely assume access to model weights or token probabilities. In this work, we present a black-box secret extraction framework for commercial, API-based LLMs under output-only access. It comprises (i) \emph{Cross-Validated Secret Knowledge Distillation}, which uses semantics-preserving prompt variants, response cross-validation, and provider-specific format filtering to distill secret-relevant behavior into a local white-box proxy; and (ii) \emph{Proxy-Guided Secret Extraction and Candidate Filtering}, which combines truncated top-$p$ sampling with local token entropy, $N$-gram frequency profiling, and provider-specific structural priors. On controlled API-key benchmarks, our framework improves recovery effectiveness and real-key rates over representative baselines while reducing extraction latency. A responsible real-world evaluation further recovers masked provider-specific credentials from three independently deployed black-box LLM systems spanning OpenAI and Claude Code, showing that memorized secrets can be exposed under output-only access.