a16z投了做AI评估的Vals,四千万美元。他们不搞排行榜,直接测真实工作流,还能用Vals Smith自己建编码基准。
Vals宣布完成4000万美元A轮融资,估值达4亿美元,由a16z领投。该公司推出Vals Smith工具,用户可从任意GitHub仓库创建自定义编码基准,并获120个免费积分。Vals还发布了与CoreWeave合作的RSI指数,以及联合多所大学推出的网络基准ReverseEngBench。公司称其收入较2025年全年增长8倍,客户数量在6个月内翻倍。
We're thrilled to invest in Vals. A frontier model can look brilliant on a leaderboard and still st...
We're thrilled to invest in Vals. A frontier model can look brilliant on a leaderboard and still struggle with the messy work that actually matters in the real world. @ValsAI takes a fundamentally different approach to evaluation: test models on the work people actually want them to do, not on contrived exams. The team works with domain experts to turn real workflows into rigorous benchmarks, then builds automated grading systems that can evaluate the final work product to an expert standard. We're thrilled to partner with @RayanKrishnan , @langstonnashold , and the entire Vals team as they build the trust layer to underpin the AI economy. By @JenniferHli , @stuffyokodraws , @RaghuRaghuram , and @shangdaxu Vals AI @ValsAI Today we're announcing our $40M Series A at a $400M valuation, led by @a16z , with participation from existing investors @8vc , @pearvc , and @BloombergBeta and new investors @HRTVentures and @nextladder . Alongside the fundraise, three more announcements: - Vals Smith: is now generally available. Anyone can create a custom coding benchmark from any GitHub repo with 120 free credits to get started. - Frontier Risk Benchmarks: We are releasing the RSI Index in collaboration with @CoreWeave and just launched ReverseEngBench, a new cyber benchmark built with Columbia University, Tufts University, UC Berkeley, and UCLA. We are also sharing our initial work in mental health, with more to come across environmental impact, military, and biosecurity. - Website + Vals Index 2.0: We completely rebuilt the Vals website and have released Vals Index 2.0, with coverage of more of the economy. Our revenue has already grown 8x compared to all of 2025. Our customer base doubled and the team tripled in 6 months. Our results have been cited in model cards from OpenAI, Anthropic, Google, Meta, and xAI. The AI economy runs on self-reported grades. When a model ships, the scores come from the company that built it. No other trillion-dollar industry works this way. Finance has ratings agencies. Medicine has the FDA. AI has vibes and vendor benchmarks. We built Vals to be the independent evaluation layer the industry is missing. Check out our new website and try Vals Smith! 🔗 View Quoted Tweet 💬 7 🔄 2 ❤️ 67 👀 14948 📊 11 ⚡