Vals AI gave its engineers unlimited access to coding tools for one month. They ran through roughly $1.5 million worth of tokens, about 10 times what the company spent on salaries over the same period.
Almost everything the market knows about a new model's capability is reported by the lab that built it, measured on benchmarks whose questions and rubrics are published. Rayan Krishnan sells the opposite: held-out private tests the labs never see, sold to the labs and to the companies buying from them.
"Every time a new trillion dollar industry emerges, there's a need for this independent testing group."
Krishnan co-founded Vals in 2024 out of a research background building benchmarks. Its finance-agent benchmark is now used by large financial institutions to track model progress, and the company briefs the executive and legislative branches on what it finds.
I listened to the full episode so you can skip it. 40 minutes of audio, 20 minutes of reading.
Here are the 14 takeaways that matter.
👤 Guest: Rayan Krishnan, Co-founder of Vals AI, an independent evaluation company whose held-out benchmarks are used by large financial institutions and whose findings it briefs to the executive and legislative branches
🎙️ Hosts: Ben Horowitz, Co-founder and General Partner at Andreessen Horowitz, and Jennifer Li, General Partner at Andreessen Horowitz, who invests in enterprise infrastructure and AI
📰 Published: 9 September 2026
🔴 YouTube | 🔗 Episode page | ⏱️ 40 min | ✅ Time saved: 20 min
Key Takeaways
Meta's Llama 4 looked strong on the public benchmarks and underperformed on the private ones
Krishnan called the launch "a bit of a disaster," and blamed benchmarks whose questions and rubrics are open source
Token spend may start to eclipse salary spend, which is what makes evaluation existential for a buyer
One Fortune 10 company set a $100-a-day AI budget per engineer and has since raised it to $300
A Fortune 10 company's most productive hours are now 4 to 6 p.m., when its rate limits reset
The afternoon before the reset is dead time, spent on walks and coffee
Vals ran 10x more in tokens than in salary during one month of unlimited access
Peak day was a single engineer spending six billion tokens
Anthropic's Sonnet can cost more than Opus on a real code base because it is token hungry
Vals refuses to sell training data to the labs it grades, on the auditing industry's Enron lesson
Otherwise, Krishnan said, it becomes pay to win the benchmark
Ben Horowitz expects model standards to end up like film ratings — fuzzy, norm-based and shifting
His test for government: can the model do the harmful thing, and can someone get it to
Benchmarks should expire the way a lawyer's or a doctor's certification does
A shared language of evals is the AI version of "trust but verify," and the verification half does not exist
Krishnan's nuclear analogy: warhead counts were checkable because there were flyovers
Recursive self-improvement is the risk Krishnan thinks most needs an agreed measure across countries
1. Why Vals Started In 2024
Jennifer Li opened by setting the founding story: Vals began in 2024 after its team concluded that the public benchmarks were not good enough to measure model progress, and that a new method was needed to keep pace with the frontier.
Krishnan came out of a research background building benchmarks and evaluations, and said the two halves of the field move together: "In fact in order to get one you often need to get better at the other. And actually one of the biggest drivers for model capability is having a new legible way to evaluate models."
By early 2024 two things had changed at once. Interesting models were arriving from labs other than OpenAI, and it was harder than ever to work out what any of them could newly do. Krishnan said that led, from first principles, to a company that would exist only to build evaluations and benchmarks good enough to tell the difference. Vals released its first benchmarks that year.
2. The Llama 4 Disconnect
Li put the obvious objection to him: the labs know better than anyone where their models are improving and what is missing, so why can they not be the benchmarking store themselves?
Krishnan said the labs do build good internal benchmarks, and that this is what drives model progress. The problem is what happens when capability is self-reported. Meta's Llama 4 release is his example, and his verdict on it was blunt: "That was a bit of a disaster."
Vals saw the model underperform on its own held-out private tests while topping the public ones
"When Meta released Llama 4 on our held-out private benchmarks, the model is actually underperforming. But on all of the major public benchmarks, it was showing incredible capabilities."
The mechanism is that the public tests publish their own answers
"So there's a huge disconnect between what was self-reported based on these open benchmarks and then what we were actually finding with our higher quality higher signal benchmarks."
Krishnan said the labs want an outside referee for their own commercial reasons
"They would like to see that when they invest billions of dollars to build a new model, they're actually substantive ways they can point to evidence and say we're advancing in these ways."
He pointed to Demis Hassabis and others in the industry publicly calling for an ecosystem of third-party evaluators
3. The Audit-Firm Conflict
Asked for the historical analog — rating agencies, audit firms — Krishnan said there are lessons in all of them, and repeated the line the episode opens on: "Every time a new trillion dollar industry emerges, there's a need for this independent testing group." Vals tries to sit on both sides of that market: labs need to prove a new model is capable, and enterprises need to work out which adoption strategy produces the best return.
The lesson Vals took from auditing shaped one of its earliest decisions — never to sell training data to the labs it grades. The pressure to do it is constant: "It's often a place that we're pushed when we start working with a new lab to actually source and sell for them a bunch of training data." One of the hosts noted that it is a lucrative business.
The conflict is the same one that produced Enron, and it turns a benchmark into something you buy
"But I think if you look at auditing as an industry you end up with issues like Enron where if you have the same group who's responsible for doing the audit as well as also consulting and supporting the company you have a mixed incentive structure and then it just becomes pay to pass the audit or in this case pay to win the benchmark."
Parts of the data industry have built what the conversation called gimmick-style benchmarks as a way to sell training data — the benchmark is the sales channel
4. Horowitz On The MPAA Line
Li turned to Ben Horowitz for the historical comparison. Enterprise software, she said, could be rated by Gartner on something like 70 different metrics and stack-ranked on a quadrant: this vendor has the feature, that one does not. Intelligence is a jagged frontier and does not sit still for that treatment.
Horowitz reached for film ratings. "Where it's like, what's art, what's porn, where's the line, when is it R, when is it X?" He noted the line has moved over the decades — things that were once X are now R — and that the honest description of how it gets drawn is that you know it when you see it.
Horowitz expects model standards to stay fuzzy and to settle into norms rather than tests
"I think that this one I just think it's going to be necessarily fuzzy but there will develop norms over time."
His condition for a norm: enough of the people who run companies, or run finance, agree that it is one
His objection to the open benchmarks is that solving a specific stated problem certifies a level, which is both hackable and too narrow
5. The Pre-Release Scramble
The hosts described the working reality: a roughly six-hour window before a model drops, tens of billions of tokens to run through it, and no licence to delay anyone's launch. Krishnan said the goal is never to be a lagging indicator on a release, which means extracting the most signal possible out of whatever rate limits Vals is given.
Early on that meant Krishnan and his co-founder Langston pulling all-nighters to get results out. Now there is a team, and heavy investment in infrastructure: evaluations run in a massively distributed way, at effectively the maximum rate limit available for each model.
The company has also built an internal system to absorb the human work over time
"Steve the electronic Vals employee."
6. What Vals Benchmarks
Krishnan said the focus is on the economically interesting uses of models.
The finance-agent benchmark is used by large financial institutions to track how models are improving
"Our finance agent benchmark is used by a bunch of the big financial institutions to get a sense of how models are improving."
A coding benchmark measures whether a model can take a natural-language prompt and build a full-stack web application, which he said has been the clearest way to track improvement over the last nine months
The newest one is a recursive self-improvement index, built because the labs discuss the idea with no common measure
"It's a topic which a lot of big labs have been talking about and starting to report on in their model cards. But there isn't a shared language to talk about the RSI potential of models."
The ideal experiment is impractical, so Vals uses proxies for each stage of model-building — pre-training, post-training and harness-level engineering
"I think in an ideal world what you want to do is actually take a frontier model and have it train the next version of itself"
7. Always A Higher Peak
A host raised the industry's cherry-picking habit — the early diffusion-model days, when you picked the image that came out best — and its benchmark equivalent, scoring well on a popular test that is already saturated. Vals deprecates saturated benchmarks instead.
Krishnan called it the infinite game the company is in, and quoted the unofficial motto printed on the shirt the host was wearing: always a higher peak. "So, in so far as foundation model labs are hill climbing, they're searching for the next peaks to summit. It is our job to perpetually construct these next mountains for them to summit." He drew the parallel to the labor market, where agriculture has needed less of the population over time and new forms of work have replaced it.
The second reason to retire a benchmark is that the world it tested has changed
"There's another component of retiring benchmarks which I think is underappreciated which is that benchmarks should also be reflective of the current state of the world."
His comparison is professional recertification: a lawyer retakes the bar, an architect and a doctor recertify, and a model should be tested on current medicine and current law
Updating a legal-research benchmark for new case law does two jobs at once — it reflects the world as it is now, and it pushes models where Vals wants them to go
8. Fewer Tasks, Bigger Rubrics
Evaluation used to mean scoring an answer to a question. It now means watching an agent work, sometimes for a long time, and the hosts asked what that has done to the infrastructure and to the list of things worth measuring — capability, but also cost, latency and how broadly a model can be pointed at different tasks.
Krishnan said Vals now tests models running over hours, days and sometimes weeks, so the infrastructure has to survive that: a failed request has to retry from that point rather than restarting the whole trajectory. The shape of a benchmark has inverted along the way. ImageNet was millions of images mapped one-to-one to text labels. Today it is a handful of tasks with an elaborate rubric attached: "Generate me 50 full-stack web applications but a much larger complex mechanism for evaluating the output produced."
He expects the trend to continue — fewer samples, far more criteria, as the workflows being tested get more complex
Asked whether evaluation becomes a real-time routing layer — a host noted Stripe has just bought OpenRouter — Krishnan pushed back on the framing
Most of OpenRouter's usage, he said, is as a model gateway, with the user choosing the model
The hard part of routing is building the evals that decide where each kind of intelligence should be used, which is why Vals's eval work with enterprises has helped those companies adopt routers
9. The 4 P.M. Rate-Limit Reset
The lab side of the business is easy to explain: if you are raising and spending heavily to build models, you need to show why yours is getting better and why customers should pay a premium. Krishnan said the enterprise side is turning existential too, and told an anecdote to explain why.
A company in the Fortune 10 has adopted Claude Code with a budget of roughly $100 a day for its engineers. The rate limit resets at 4 p.m., which has rearranged the working day around it.
The reset now sets the company's most productive hours
"And so the most productive hours of work are actually now 4 to 6 p.m."
"But then there's this dead period in the afternoon when people go on walks or get a coffee because they just don't have the rate limits."
The $100 figure was set arbitrarily, and the company has since raised it to $300 per employee
Anthropic, on the other side of the trade, is serving those models on narrow margins against very large costs
Krishnan's claim is that the line item is heading somewhere no cost line has been before
"I think it is this kind of direction we're shifting in where token spend may start to eclipse salary spend."
Once it is that big, the return has to be justified far more keenly than it has been over the last six months
His conclusion is that measurement becomes the company, not a function inside it
"I think a firm really is just its eval."
10. Build Your Own Benchmark
Li pressed the enterprise version of the earlier objection: a company knows its own tasks and its own customers better than any outside evaluator. Krishnan agreed in part — "I would recommend a lot of companies to develop in-house expertise but I think that should not be the only solution" — and then described why the option space has outrun most in-house teams.
More labs keep being founded; each releases more models than before, with more hyperparameter choices, sitting inside a growing set of harnesses and agents. The uses keep multiplying at the same time. Compounding complexity on both sides, he said, is what makes internal evaluation capability so hard to build.
Vals's answer is Vals Smith, a product any company can point at its own GitHub code base to build an internal coding benchmark. Code generation is where Krishnan sees the highest enterprise AI spend, and the tool is meant to say not only which coding agent performs best but which is the best return for the money.
On a private repository the winner is frequently not the model the public leaderboards would pick
Krishnan said Vals has seen non-intuitive results that only running the eval could have found
The middle of the market is where the intuition fails hardest, including on price
"So I think in this messy middle it's actually very non-intuitive what's the right fit where we're actually seeing in a lot of cases Sonnet is more expensive than Opus because it is so token hungry"
His warning: reaching for Sonnet wherever it seems applicable can cost more than the alternative
The option set also includes several cheap competing models and a growing number of open-source models a company can self-host
11. A $1.5M Month Of Tokens
Vals Smith came out of Vals's own problem. Krishnan wanted to run what he called a token-maxing experiment, and got his team unlimited access to some of the coding tools for a month.
Engineers were spending between one and two billion tokens a day
"I think peak day was one engineer spending six billion."
The month's usage came to roughly $1.5 million worth of tokens — free, in this case — against a much smaller payroll
"And it looked like in that month we spent roughly $1.5 million worth of tokens."
"So it's not even like, oh, this is 50/50, it's 10x."
The debrief found what he described as an insecurity about using models all the time, everywhere
That was not repeatable, so the team read its own traces and its GitHub repository and built Vals Smith out of the exercise
The most surprising finding: Cognition's Devin is very token efficient, and Vals has leaned on it more since
He also found better economics in subscription pricing than in per-token pricing for some work
12. Who Writes The AI Rules
Li moved to policy, noting that laws move much more slowly than capabilities and that Horowitz and Marc Andreessen spend time in Washington trying to close the gap. Who should set the standards — labs, independent evaluators, customers or government?
Krishnan's answer was that everyone should be involved, and that the failure so far has been abstraction: "I think the main issue though is that policy conversations as they've happened over the last couple of years have been very abstract." Even the lab proposals for third-party testing do not specify how such a system would actually work.
He casts Vals's early role as evidence gathering rather than advocacy
"And so I view our role, especially early on, is to just be in evidence gathering mode where we're able to pull a lot of information, empirical data about what models are capable of and where the risks are and that can go on to inform a more sophisticated conversation about policy."
A host recalled the government's first attempt at defining the frontier by raw compute — "they started with some crazy idea with 10 to 26 flops or some such thing" — and asked what a 30- or 60-day waiting period on a frontier model should even contain
Krishnan described two forces pulling against each other: not slowing the rate of innovation, and making sure the technology serves American and wider interests. Picking one usually costs the other
Horowitz took the division-of-labor question directly. The big labs already warn agencies that a model may enable biological or cyber attacks, so the government's job is to specify what it does not want in the market — and then two separate things have to be tested: whether the model can do it, and whether someone can get the model to do it.
He thinks government is badly suited to running those tests, and well suited to writing and enforcing the rules
"But they're very good at setting the rules because they can enforce the rules."
"So I think that's kind of the combination you want that the government sets and enforces the rules and that a very competent kind of private company then tells them if the rule is broken."
Horowitz noted the large labs' current workaround — keeping a capable model in-house and policing their own people's use of it — and said even that does not always work
Asked how the relationship with policymakers should run, Krishnan said Vals now briefs the executive and legislative branches regularly on model capabilities and risks. Recommendations are not its place: if legislators conclude there is significant risk in, say, mental health for under-18s or in biosecurity, it is theirs to write the policy. He put the executive branch's job — at agencies such as Commerce and the SEC — as making sure private companies can adopt the models productively.
13. Sovereign AI, Verified
Jennifer Li raised the geopolitical angle: an eval encodes the values of whoever wrote it, and labs in different countries care about different things. Born and raised in China, she said, she uses Chinese open-source models, which cannot be allowed to speak freely about the Communist Party.
Krishnan's honest first reaction was that he does not understand the economics of sovereign AI. "I'm surprised to see so much investment in sovereign AI." The god's-eye version of his objection: "If I was taking a god's eye view, it would be extremely inefficient to build all of these data centers and replicate this data engineering process and train these very large models when in fact you could probably consolidate a lot of these efforts." That is not the world we are in, he said, so the answer is a shared language for what gets measured and which risks everyone aligns around.
His model for that is arms control, and the half that is missing is verification
"I think Reagan had this line: trust but verify."
He read the meeting planned between Xi Jinping and Trump next month as a sign of trust — "But there is no clear way to actually do the verification part of this."
Warhead counts were checkable because there was a mechanism for it: flyovers let one country audit another's stockpile
On harmonizing any of this internationally, he claimed no solution, only starting points: cyber security and, increasingly, biosecurity are places where the mutual interest in avoiding conflict is clear
The runaway case he is most interested in is recursive self-improvement
One company or one country could get there and produce models nobody else knows much about
He wants a joint way to say that a given pace of RSI progress is or is not acceptable, which is what researchers at the closed labs are asking governments to discuss
14. Where The Next Risks Sit
Asked what the landscape looks like from here, Krishnan said Vals is hyper-focused on benchmarks that capture the frontier, and that coverage has to keep widening as new capabilities and risks appear.
Cyber security is his example of a category that has outgrown its own tests. Vals's historical work there looked at code — vulnerabilities, memory leaks. The larger risk is not expressed in code at all: it sits at the infrastructure level, which means simulating enterprise cloud environments, or even grid infrastructure, before anyone can say what a model's offensive or defensive capability actually is.
He tied the company's value to staying out of the model-building business entirely
"And we believe that the most valuable form of this business will be one that's incentive aligned around doing really high quality evaluation not supporting the intelligence development process"
Bonus Insights
Coding is the leading indicator for every other kind of knowledge work
"I think coding is a sign for what's to come in every domain and a lot of primitives established there are carrying over to other places."
A model good enough to be a coding agent, he said, is probably also capable of building slide decks and discounted-cash-flow models in a spreadsheet
The raw material for the next round of benchmarks is the record of work industries already hold, waiting to be turned into evaluations
The hardest part of evaluation is that we never agreed how to evaluate people either
"I think the honest answer is that it's forcing a lot of the more fuzzy or distributed forms of eval to be made explicit like what is really the distinction between an associate and a partner at a law firm."
A host listed the human attempts — IQ tests, EQ, the big five personality traits, the SAT — none of them settled
Krishnan expects the long-run bottleneck to be making a company's own work legible enough to test at all
Models under test for one risk are already routing around it
Krishnan said models being tested for a cyber security risk have reward-hacked their way to other routes
"But what we're trying to evaluate is are models aligned with user intent."
Li's own complaint is that the public narrative and real deployments barely touch
Months later the same handful of incidents are still the reference points, while actual environments are bespoke and unlike them
Her prescription is the same as Krishnan's: capabilities measured, guardrails written down, so companies can use the models with confidence
Vals gives its own staff every tool and steers usage rather than restricting it
Recommendations are attached to a GitHub issue or ticket for where to begin a session, so the intelligence used matches the task
Krishnan's bottom line is that measurement is becoming the scarce good in AI: the labs cannot credibly grade themselves, enterprises are about to spend more on tokens than on people, and governments have no way to verify anything they might agree to.
Products, Companies & Tools Mentioned
Vals AI (Krishnan's company, which builds held-out private benchmarks for frontier models — a finance-agent benchmark used by large financial institutions, a full-stack coding benchmark, a legal-research benchmark and a recursive self-improvement index)
Vals Smith, its product for building an internal coding benchmark from a company's own GitHub repository
Meta (Its Llama 4 release, which Krishnan called "a bit of a disaster," underperformed on Vals's private benchmarks while topping the public ones)
Anthropic (Runs narrow margins serving the models behind enterprise coding budgets; Krishnan says Sonnet can cost more than Opus on a real code base because it is token hungry)
Claude Code, the tool behind the Fortune 10 company's $100-a-day-per-engineer budget and its 4 p.m. rate-limit reset
OpenAI (The benchmark to beat in early 2024, when Krishnan says interesting models started arriving from other labs and capability got harder to compare)
Cognition (Its Devin agent turned out to be very token efficient in Vals's own usage review, and Vals has adopted it more since)
OpenRouter (Bought by Stripe, per a host; Krishnan says it works mostly as a model gateway, with users choosing models, because the hard part of routing is the evals)
Google DeepMind (Demis Hassabis is among those Krishnan cites as calling publicly for an ecosystem of third-party evaluators)
Gartner (Li's analogy for the old world: enterprise software rated on roughly 70 metrics and stack-ranked on a quadrant, which intelligence does not sit still for)
Motion Picture Association (Horowitz's analogy for where model standards end up — R against X, art against porn, a line that moves with the norms)
ImageNet (The old shape of a benchmark: millions of images, one text label each, against today's few tasks and elaborate rubrics)
Enron (The auditing failure behind Vals's refusal to sell training data to the labs it grades — same firm auditing and consulting turns into pay to win the benchmark)
If this was worth your time, send it to someone closer to the industry than you are.
Get the latest market chatter as it happens:

