The Pulse
Vals Raises $40 Million to Build a New AI Benchmarking Standard
Vals has raised a $40 million Series A led by Andreessen Horowitz to expand private, industry-specific evaluations of AI models. The San Francisco startup tests systems on legal, financial, coding and other real-world tasks rather than rely

AI.info Team ·
“What we’re doing is actually looking at what are the real impacts of the models.”
Rayan Krishnan, co-founder, Vals
Vals has raised $40 million in a Series A led by Andreessen Horowitz, placing a young AI evaluation company at the center of a growing argument over how model performance should be measured. The San Francisco startup says public, academic tests no longer provide a reliable guide to whether an AI system can complete the professional work companies want to automate.
Founded in 2024, Vals keeps its test materials private and evaluates models on tasks drawn from fields including law, finance and software development. The funding announcement came on August 13, 2026, while the company’s wider benchmarking strategy was detailed in an interview published by TechCrunch on September 19.
Public tests are becoming easier to optimize
Rayan Krishnan, Vals’ 25-year-old co-founder, told TechCrunch that the company emerged from his view that academic benchmarks were falling behind the pace of model development. “We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance,” Krishnan said.
Vals argues that openly available test sets create a basic measurement problem. Once benchmark questions enter training data or become targets for model optimization, a high score can say less about general capability and more about familiarity with the exam. The company therefore withholds the specific materials used for its published evaluations while providing descriptions of the tasks and the domains they cover.
Andreessen Horowitz made a similar case in its investment announcement, writing that public datasets can become saturated, leak into training corpora or become targets for explicit optimization. The firm said Vals tests whether a legal model can perform legal research, whether a finance model can analyze complex documents and whether a coding model can build a working application.
From bar-exam scores to work products
Vals’ approach focuses on the output of a task rather than a model’s performance on a general knowledge examination. “Historically, I think evaluation has been done to evaluate intelligence in a very abstract way,” Krishnan told TechCrunch. “Like, do models know enough information to be able to take a bar exam type test?”
The company instead asks whether a system can produce work comparable to a human professional. Its public benchmark catalog includes evaluations for legal research, financial analysis, tax work, medical administration, software engineering and formal mathematics. The site also lists tests for cybersecurity, games, public benefits, voice transcription and scientific research.
Recent entries show how far Vals is extending beyond conventional question-and-answer testing. Its benchmark page lists the Vals RSI Index, updated September 18, for assessing whether a model can conduct research that contributes to building the next model. Other recent evaluations examine binary reverse engineering, agentic biology investigations, security-bug discovery, web application development and autonomous work inside a simulated space program.
A private examiner with public consequences
Companies pay Vals to evaluate their models, a business model Krishnan compares with students paying the College Board to take the SAT. The value, he argues, comes from finding weaknesses that developers can address before a system reaches customers or is used in a high-stakes setting.
Vals says its evaluation infrastructure can collect criteria from subject-matter experts and run tests across large numbers of models. Its methodology also reports uncertainty: single-run benchmarks use standard error over instance-level scores, while selected multi-run tests calculate uncertainty across independent runs. The company says those error estimates do not capture every source of variation, including changes in prompts, deployment conditions or model generation.
The arrangement creates a tradeoff. Private tests reduce the risk that models have been trained against the questions, but outsiders cannot inspect the full test set or independently reproduce every result. Vals’ standing will therefore depend not only on the difficulty of its tasks, but also on whether model developers, buyers and researchers trust its methods and reporting.
Eightfold revenue growth and a federal program
TechCrunch reported that Vals’ revenue is now eight times what it was last year. The company began 2026 with eight employees and had grown to 25 by September, with plans to move into a larger office and hire another 10 to 15 people.
Vals has also launched a program to provide model evaluations to federal agencies. Its stated scope reaches beyond measuring useful outputs: Krishnan said the company is working on evaluations involving mental health, cybersecurity, biosecurity and the law of armed conflict, alongside a benchmark examining recursive self-improvement.
Those projects reflect the company’s broader claim that evaluation should measure both what an AI system can accomplish and how it might fail when placed in consequential environments. “Can they do work that produces a product of the same quality as a human within every domain?” Krishnan said.
The test Vals must pass itself
Andreessen Horowitz’s investment gives Vals money to expand its evaluation catalog and sell more testing to model developers and institutional buyers. The company’s own growth figures suggest demand is rising as businesses compare systems for procurement rather than treating benchmark scores as a purely technical exercise.
Its harder task is establishing authority. A benchmark that stays private can avoid contamination, but its credibility rests on the examiner’s process, the quality of its expert reviewers and the clarity of the evidence it publishes. Vals is now asking the market to accept that its confidential tests are a better guide to AI performance than the public exams they are designed to replace.