Newsclip — Social News Discovery

Business

Vals AI: A New Standard for Trust in AI Benchmarking

September 19, 2026
  • #AI
  • #Benchmarking
  • #Techinnovation
  • #Startupnews
  • #Artificialintelligence
  • #Trustinai
0 views0 comments
Vals AI: A New Standard for Trust in AI Benchmarking

Reimagining AI Evaluation

When I first encountered the term "benchmarking" in the context of artificial intelligence, I was struck by its simplicity—yet profound complexity. At its core, benchmarking is supposed to be a way to test and measure how well an AI model performs compared to others. But in practice, it's often a game of appearances. Companies are quick to tout their models' scores on publicly available tests, even if those tests were designed years ago and no longer reflect real-world capabilities.

"We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance," said Rayan Krishnan, co-founder of Vals AI.

This is precisely where Vals steps in. Formed in 2024, the company has rapidly established itself as a serious contender in the AI evaluation space, having recently raised $40 million in Series A funding led by Andreessen Horowitz. The mission? To create a more accurate, ethical, and trustworthy method of evaluating AI models.

Why Traditional Benchmarks Fall Short

The current landscape of AI benchmarking is riddled with issues that undermine its credibility. Many of these tests were designed in the early days of AI, when models were far less sophisticated than today's systems. They're often static and easily gamed by companies that optimize their models specifically for these benchmarks.

Consider this: if a company's model excels on a test that has been publicly shared and optimized for, it may not necessarily perform well in real-world applications. This discrepancy is what Krishnan identifies as the core problem—AI systems are being evaluated on their ability to game old tests rather than on their ability to actually solve complex problems.

As AI continues to integrate into sectors like finance, law, and cybersecurity, the stakes of accurate evaluation rise exponentially. A model that performs well in an academic setting may fail spectacularly when tasked with a real-world challenge such as interpreting complex legal documents or securing sensitive data.

Introducing Vals: Beyond the Test

Vals AI's approach to benchmarking is fundamentally different from traditional systems. Rather than relying on general knowledge tests that can be easily circumvented, Vals focuses on domain-specific tasks. The idea isn't just to see how much a model knows, but rather how well it can apply its knowledge in meaningful ways.

This means that instead of testing whether a model has memorized facts, Vals evaluates whether it can produce outcomes equivalent to human performance in specific industries. Whether it's coding, legal analysis, or financial risk assessment, the company measures models on their ability to complete tasks with real-world impact.

But what makes Vals particularly interesting is its forward-thinking approach to evaluation. Rather than only looking at positive outcomes, Vals also evaluates potential negative implications of AI models running loose in society. This includes scenarios such as cybersecurity vulnerabilities or ethical concerns like bias in decision-making.

A Strategic Approach to Trust

As Krishnan puts it, "What we're doing is actually looking at what are the real impacts of the models. Can they do work that produces a product of the same quality as a human within every domain?" This approach reflects a deeper understanding of how AI systems interact with our lives.

One aspect of Vals' methodology that deserves attention is its proprietary test materials. Unlike many competitors, Vals does not make its test details public, preventing companies from gaming the system by optimizing models specifically for those tests. This creates an environment where true capability is measured rather than performance on a particular set of metrics.

This strategy aligns with a broader trend in AI development: companies are beginning to recognize that trust is just as important as performance. Vals' revenue model—where companies pay to get their models evaluated—is somewhat counterintuitive at first glance, but it's actually quite logical. Just as students pay for the SAT, companies are paying for an independent assessment of their technology's real-world effectiveness.

The Growing Impact of AI Benchmarking

The rise of Vals isn't just about one company's innovation—it's part of a larger shift in how we think about AI development. As more AI companies go public and become central to the economy, benchmarking will play an increasingly critical role in investor decision-making and regulatory oversight.

Think about it: SpaceX went public, and now Anthropic is preparing for its own IPO. OpenAI could be next. In this environment, how we evaluate AI systems becomes a question of public trust. Vals' approach suggests that the future of AI evaluation will be less about marketing claims and more about demonstrable results.

Moreover, Vals' expansion into areas like mental health, cybersecurity, biosecurity, and even international law demonstrates a commitment to evaluating AI not just in terms of what it can do, but also in terms of how it might affect society. This is especially important as AI models become more autonomous and capable of making decisions that impact human lives.

Real-World Applications and Future Directions

As Vals continues to grow—currently employing 25 people after starting with just eight—the company has plans to expand its scope even further. One exciting development is their recent launch of a program focused on providing model evaluations to federal agencies. This indicates that Vals' approach may soon become a standard for government oversight as well.

The company's trajectory reflects an important trend in the AI space: the increasing need for independent, trustworthy evaluation mechanisms. As artificial intelligence becomes more embedded in society, the responsibility of ensuring it functions ethically and effectively becomes paramount. Vals is positioning itself at the forefront of this movement.

Looking ahead, I believe we'll see a proliferation of companies adopting similar approaches to benchmarking. The current system, while functional in theory, has shown its limitations when faced with rapid technological advancement. Vals represents one possible solution to this challenge—but more importantly, it signals that the industry is finally beginning to take responsibility for how AI systems are evaluated and deployed.

Conclusion: A New Standard for AI Trust

The emergence of Vals AI is not just another tech startup story—it's a critical step in establishing a more responsible future for artificial intelligence. As we continue to integrate these powerful tools into our lives, the need for accurate, transparent, and ethical evaluation systems cannot be overstated.

While there are certainly risks associated with this kind of advanced AI benchmarking—such as potential misuse or overreliance on any single system—it's clear that companies like Vals are taking steps toward a more nuanced understanding of what makes an AI model truly valuable. Their approach suggests that the next generation of AI evaluation will be less about scores and more about societal impact.

As I reflect on this development, one thing is certain: we're entering a new era in which AI's success will be measured not just by its computational prowess, but by its ability to enhance rather than harm human life. Vals AI may well become the gold standard for that measurement.

Key Facts

  • Company Name: Vals AI
  • Founded: 2024
  • Series A Funding: $40 million
  • Lead Investor: Andreessen Horowitz
  • Co-founder: Rayan Krishnan
  • Headquarters: San Francisco
  • Team Size: 25 people
  • Benchmarking Approach: Domain-specific tasks with real-world impact

Background

Vals AI is a startup company founded in 2024 that aims to revolutionize how artificial intelligence models are evaluated and benchmarked. The company was established by Rayan Krishnan, who previously interned at Palantir and worked at Microsoft and Stanford's artificial intelligence lab. Vals addresses the limitations of traditional AI benchmarking methods, which often rely on outdated academic tests that can be easily gamed by companies optimizing their models specifically for those benchmarks. The company's approach focuses on evaluating AI models based on their ability to perform complex tasks in specific industries rather than general knowledge tests.

Quick Answers

What is Vals AI?
Vals AI is a startup company formed in 2024 that specializes in AI benchmarking with a focus on real-world impact and ethical considerations.
Who is Rayan Krishnan?
Rayan Krishnan is the 25-year-old co-founder of Vals AI who previously interned at Palantir and worked for Microsoft and Stanford's artificial intelligence lab.
When was Vals AI founded?
Vals AI was founded in 2024.
How much funding did Vals AI raise?
Vals AI raised $40 million in Series A funding led by Andreessen Horowitz.
Where is Vals AI headquartered?
Vals AI is headquartered in San Francisco.
What makes Vals AI's benchmarking approach different?
Vals AI's approach focuses on domain-specific tasks and real-world impact rather than general knowledge tests, and it does not publicly disclose its test materials to prevent gaming the system.
What is Vals AI's revenue model?
Vals AI's revenue model involves companies paying to have their models evaluated by the company, similar to how students pay for SAT testing.
Why is Vals AI important for AI development?
Vals AI is important because it provides more accurate, ethical, and trustworthy evaluation of AI models that can measure real-world performance rather than just test scores.

Frequently Asked Questions

What problem does Vals AI solve?

Vals AI solves the problem of outdated AI benchmarking systems that don't accurately reflect real-world capabilities of modern AI models.

How does Vals AI evaluate AI models?

Vals AI evaluates models on their ability to complete complex tasks in specific industries rather than general knowledge tests, and it measures both positive outcomes and potential negative implications.

What industry sectors does Vals AI focus on?

Vals AI focuses on sectors such as law, finance, coding, cybersecurity, biosecurity, mental health, and even international law including the law of armed conflict.

Does Vals AI publish its test materials?

No, Vals AI does not make its test details public to prevent companies from optimizing models specifically for those tests.

Source reference: https://techcrunch.com/2026/09/19/vals-backed-by-andreessen-horowitz-is-looking-to-become-the-gold-standard-for-ai-benchmarking/

Comments

Sign in to leave a comment

Sign In

Loading comments...

More from Business