Vals AI CEO谈主权AI需冷战式验证
Vals AI CEO Rayan Krishnan on why sovereign AI makes Cold War-style verification necessary: "From m...
Vals AI CEO分享主权AI验证难题,借鉴核武器冷战经验,探讨AI评估框架建设。
Vals AI CEO Rayan Krishnan指出主权AI投资增加,需要建立共享评估语言。他借鉴核武器验证经验,提出"信任但验证"机制。当前AI模型能力主要依赖自我报告,缺乏独立评估体系。随着模型优化测试能力增强,需要持续演进的独立评估方法。
Vals AI CEO Rayan Krishnan on why sovereign AI makes Cold War-style verification necessary: "From m...
Vals AI CEO Rayan Krishnan on why sovereign AI makes Cold War-style verification necessary: "From my very idealistic perspective, I'm surprised to see so much investment in sovereign AI." "You could probably consolidate a lot of these efforts. But it seems like that's not the world we're in or the one we're headed towards. There are actually increased efforts to build AI in a sovereign way." "So that takes having a shared language to communicate about what the framework for evaluations is, and where we're going to collectively align around the risks." "There's a lot to learn from nuclear here. Reagan had this line, 'Trust but verify.' We're starting to see signs of trust in that Xi Jinping and Trump are going to be meeting... But there is no clear way to actually do the verification part of this." "Having the shared language of evals will allow us to say things like, you have the right number of nuclear warheads. In that example, there were also flyovers, a mechanism by which a country could audit another country's nuclear stockpile." "If there's concern about the societal or even existential risk of AI, it will necessitate us constructing this shared language of evaluations to do the verification process." @RayanKrishnan @JenniferHli Your browser does not support the video tag. 🔗 View on Twitter a16z @a16z Vals AI co-founder and CEO Rayan Krishnan with a16z's Ben Horowitz and Jennifer Li on grading AI, what it costs, and who gets to make the rules: Every big industry eventually grows an independent testing layer. AI has credit ratings to learn from and Enron to avoid. Model capability today is still mostly self-reported. As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. The harder problem is geopolitical. Reagan's "trust but verify" worked during the Cold War because you could fly over and count the missiles. No simple equivalent for AI models exists. In this conversation with Erik Torenberg, they get into how you measure a model's ability to improve itself, why every good benchmark eventually has to be retired, and what happens when token spend begins to rival employee salaries. 00:00 Intro 02:20 Llama 4 on public vs private benchmarks 05:24 Nobody agreed how to test humans either 06:55 What movie ratings teach us about AI 08:55 The Enron problem in benchmarking 11:36 Why a good benchmark has to be retired 13:22 Evals that run for weeks, not seconds 16:20 Where the real workday starts at 4pm 18:08 A firm really is just its evals 20:35 Why Sonnet can cost more than Opus 22:42 One engineer, 6 billion tokens in a day 25:05 Who should set the rules for models 28:55 Public sector enforces, private verifies 33:32 Why sovereign AI is inefficient and happening anyway 35:00 The AI version of trust-but-verify 37:15 Where cyber evals have to go next YouTube: youtube.com/watch?v=WO9c9q… @RayanKrishnan @ValsAI @bhorowitz @JenniferHli @eriktorenberg Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 2 🔄 0 ❤️ 4 👀 3017 📊 2 ⚡