谷歌用加密技术实现AI模型双盲评估,解决基准测试信任问题
Google Deepmind首次对前沿AI模型进行双盲评估测试。通过Confidential Space加密技术,谷歌无法看到测试问题,评估者也无法接触模型权重。新加坡AI安全研究所参与试点项目,使用Gemini Flash Lite模型。这一方法可能为防篡改AI基准测试树立新标准。
AI benchmarks have a trust problem and Google wants to fix it
Google Deepmind is testing a double-blind evaluation of a frontier AI model for the first time. Cryptographic protection through Confidential Space is meant to keep Google from seeing the test questions and keep evaluators from seeing the model weights. The pilot project with the Singapore AI Safety Institute uses a Gemini Flash Lite and could set a new standard for tamper-proof AI benchmarks. The article AI benchmarks have a trust problem and Google wants to fix it appeared first on The Decoder .