Logo
FrontierNews.ai

The AI Benchmarking Problem That's About to Get Solved

AI companies have figured out how to game traditional benchmarks, and a well-funded startup is now racing to fix the problem by measuring what actually matters: whether AI models can do real work. Vals, founded in 2024 and recently backed by a $40 million Series A investment led by Andreessen Horowitz, is positioning itself as the gold standard for AI model evaluation at a time when the industry desperately needs trustworthy measurement systems.

The core issue is straightforward but consequential. As AI models have become more capable, the academic benchmarks used to measure them have fallen behind. Companies can now train their models specifically to perform well on publicly available tests, essentially allowing them to study the answers in advance. Meanwhile, investors, regulators, and the public have no reliable way to verify whether the capabilities AI companies claim actually translate into useful work.

Why Traditional AI Benchmarks Are Becoming Obsolete?

For years, AI evaluation relied on standardized tests borrowed from academia. These benchmarks measured general knowledge and reasoning ability, similar to how a bar exam tests whether someone knows enough law to practice. But this approach misses something crucial: whether an AI model can actually perform complex, real-world tasks that produce value.

Rayan Krishnan, Vals' 25-year-old co-founder, observed this gap firsthand. A Stanford graduate who previously interned at Palantir and worked at Microsoft and Stanford's AI lab, Krishnan explained the fundamental shift his company is pursuing. "Historically, I think evaluation has been done to evaluate intelligence in a very abstract way," he said. "Like, do models know enough information to be able to take a bar exam type test? What we're doing is actually looking at what are the real impacts of the models. Can they do work that produces a product of the same quality as a human within every domain?"

Rayan Krishnan, Vals' 25-year-old co-founder

"We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks were not keeping up with that frontier advance," Krishnan explained.

Rayan Krishnan, Co-founder at Vals

Vals differentiates itself by keeping its test materials private, preventing companies from training their models to pass specific evaluations. Instead of measuring abstract knowledge, the startup evaluates models on their ability to complete complex tasks in specific industries.

How Vals Is Building a New Standard for AI Evaluation

  • Industry-Specific Testing: Vals evaluates models on real-world tasks in law, finance, coding, and other specialized domains, measuring whether AI can produce work quality equivalent to human professionals.
  • Hidden Test Materials: Unlike legacy benchmarks with publicly available tests, Vals keeps its evaluation criteria confidential to prevent companies from gaming the system through targeted training.
  • Negative Impact Assessment: The company doesn't just measure positive outcomes; it also evaluates potential harms, asking what would happen "if these models ran wild in the world," including testing for biosecurity, cybersecurity, mental health risks, and compliance with international law like the Geneva Convention.
  • Emerging Capability Benchmarks: Vals is pushing into cutting-edge territory, including benchmarks for recursive self-improvement and applications in areas like law of armed conflict.

The business model itself is unconventional but logical. Companies pay Vals to evaluate their models, similar to how students pay the College Board to take the SAT. While it might seem counterintuitive for a company to pay for potentially negative results, effective measurement helps organizations identify weaknesses and improve their systems over time. These evaluations are increasingly becoming key decision-making factors for companies considering acquiring or deploying new AI models.

The startup's growth trajectory reflects the market's hunger for trustworthy evaluation. Vals started 2026 with only eight employees and has already tripled to 25 people. Revenue is currently eight times what it was last year. Krishnan plans to relocate to a significantly larger office and hire an additional 10 to 15 people as the company scales.

Why AI Benchmarking Matters for the Entire Industry?

As AI companies prepare to go public, benchmarking is becoming a critical component of corporate governance and investor confidence. Anthropic is slated to go public later in 2026, and Krishnan expects OpenAI to follow. When AI models become central to economic activity and are deployed across industries, the benchmarks and evaluations that measure their capabilities will directly influence how companies present themselves in public filings and attract investment.

Vals has also recently launched a program providing model evaluations to federal agencies, signaling that government bodies recognize the importance of independent, rigorous AI assessment. This positions the startup at the intersection of private sector innovation and public sector oversight at a moment when both are racing to understand and manage AI capabilities.

"AI companies are starting to go public. SpaceX went public. Anthropic is slated for later this year. I suspect OpenAI will be public soon. I think as AI models become a core part of the economy and are diffused more broadly, the types of benchmarks and evaluations that we do are going to drive their usage and be a central part of how these companies submit public filings or talk about the prospective investments they're going to make in AI," Krishnan stated.

Rayan Krishnan, Co-founder at Vals

The emergence of Vals and its rapid funding success reflects a broader recognition that the current benchmarking landscape is broken. As AI models become more powerful and more integrated into critical systems, the ability to measure what they can actually do, what they cannot do, and what harms they might cause has shifted from a nice-to-have to an essential infrastructure requirement. For OpenAI, Anthropic, and other frontier AI labs, the question is no longer whether their models will be evaluated, but by whom and according to what standards.