San Francisco-based Vals has closed a $40 million Series A led by a16z at a $400 million valuation. The company builds evaluations and benchmarks that test AI models on real-world tasks. Its revenue is reported to have grown eightfold since 2025.

Why a market formed

This round demonstrates something that would have looked odd a few years ago: measuring models has become an industry in itself.

The cause is a concrete problem organisations face. When a company puts AI into its operations it has to answer three questions: which model do I use, does this model actually work on my work, and did something break when the version changed?

General model leaderboards do not answer those. Rankings measure on standard exams; they do not contain an insurer's policy documents or a bank's compliance filings.

What \"real-world tasks\" means

That distinction is where Vals positions itself. It builds benchmarking not as an academic exam but as professional work: tasks in law, finance or healthcare of the kind specialists in those fields solve.

Building a benchmark of that sort is not cheap. It needs domain experts, correct answers have to be prepared, and the set has to stay private — a published exam enters the next model's training data and loses its measuring value.

That cost explains why the work gets outsourced to a company.

Eightfold growth

Revenue growing eightfold in a year shows where the demand comes from. An auditing tool growing at that rate means the thing being audited is spreading fast.

That is another indicator of enterprise AI moving from pilot to production. Nobody spends money measuring a tool they are trialling; the need to measure arises once the tool starts doing work.

The independence question

The underlying issue is who performs the evaluation. Today most information about model performance comes from figures the model's own builder publishes.

An independent evaluation company raising at this scale shows there is a commercial answer to that gap. That independence has its own limit, though: the evaluator's revenue also comes from the ecosystem the evaluated companies inhabit.

The industry's quiet layer

Vals's round is part of a pattern that keeps repeating. Observability, evaluation, governance, security — all of these are companies that build no models but do business around them.

The growth of that layer is a sign the AI market is maturing. A new technology is first discussed in terms of its capability; then the tools that make that capability usable reliably form a market of their own. Cloud computing followed the same sequence.

What small teams should do

A tool at this scale is not necessary for everyone. The idea behind it is:

  • Prepare a set of questions drawn from your own work, with known answers.
  • Re-run that set whenever you change model or version.
  • Record the result — so that where quality is heading over time becomes visible.

Setting that up takes a few hours and moves model selection from guesswork to measurement.