Rayan Krishnan's firm extrapolates its own benchmark to models eclipsing human AI researchers around August 2027, and he gave that date on air as a prediction rather than a warning.
The week's essays and resignations have been read as a crisis inside the labs. Krishnan tests those labs' unreleased models for a living and described the opposite experience: companies being forthcoming, and a coordination problem that market-based solutions can reach.
"And so our ability to create better models now outpaces our ability to actually understand them. That's the issue at hand."
Krishnan runs Vals AI, which independently tests and benchmarks frontier models, including unreleased models from OpenAI and Anthropic, so he sees the systems before the public does.
The full segment is covered here so you can skip it.
Here are the 4 arguments that matter.
👤 Guest: Rayan Krishnan, Founder and CEO of Vals AI, which independently benchmarks frontier models, including unreleased ones, and publishes an index tracking recursive self-improvement
🎙️ Host: Ed Ludlow, who anchors Bloomberg Tech from San Francisco
📰 Published: 14 September 2026 on Bloomberg Tech
🟣 Apple Podcasts | 🔗 Episode page | ⏱️ length not available
Key Takeaways
The money went into making models, not into testing them, and that gap is the actual problem
Krishnan says the ability to build better models now outpaces the ability to understand them
His index extrapolates to models passing human AI researchers around August 2027
Today's models execute experiments well and operate as engineers, but do not originate them
What the index cannot see is what the labs keep inside
Vals tests publicly available models, not unreleased systems, specialized agents or multi-agent systems
An evaluator that also fixes what it finds is the Enron failure
Vals publishes its methodology, keeps its test sets private, and does not sell training data or solutions to the labs it tests
He thinks the labs cooperate with or without regulation, because they are rational actors
1. Optimistic, Not Alarmed
Ludlow set up the segment on the one concrete item in Dario Amodei's essay — independent oversight — and Anthropic's offer to bring in third-party evaluators with near-employee-level access to scrutinize not just finished models but how they are being developed. Krishnan had responded publicly to the essay, and Ludlow asked for his conclusion.
His first move was to push back on the temperature of the week rather than the substance. He said there have been a lot of extreme statements over the last week, that he understands where they come from, but that they can also cause undue anxiety for the general public
What he sees from inside the process does not match the public argument. He said that from what Vals AI sees with boots on the ground, the firm has been able to see the labs be very forthcoming
He said the firm engages across all the major foundation model labs, as well as with members of Congress, government agencies and some of the largest enterprises
The conclusion he drew is a distinction between the stage and the work: "So I actually think there's a lot of conflict happening on the public stage, but on the whole, we have reason to be optimistic that coordinated effort is possible."
He summarized what the essays are asking for as slowing the pace of capability development so that safety alignment can catch up
2. Capability Outruns Testing
Ludlow checked the premise of Krishnan's standing — that he sees these models in the training phase, before release.
Krishnan's answer names an imbalance in where the industry's money has gone. "Yeah, well, I think what we see is that there's really been an outsized investment towards the generation or the capabilities of the models without that same parallel investment in the testing and evaluation," he said
The consequence is the line the whole segment turns on: "And so our ability to create better models now outpaces our ability to actually understand them. That's the issue at hand."
He gave the reason the argument arrived this month rather than last year. He said what Amodei wrote about is the ability of AI to aid in the research and development of future generations of AI, so-called recursive improvement
3. The RSI Index
Ludlow brought the firm's index up on screen and asked Krishnan to explain the methodology and what it says about an AI's ability to contribute to the development of AI.
The index measures autonomy, not intelligence. Krishnan said it is the mainline figure tracking the progress of the major frontier models toward recursive self-improvement. "And that describes the model's ability to make the successor version of itself autonomously. So wholly without human input," he said
His illustration: "And so this would be akin to GPT-6 being responsible for making GPT-7, or Claude making the next version of itself without any human researchers."
Where the models are now is early and well short of people. He said the plot shows some early signs that models are exhibiting recursive self-improvement, but that they are still far away from the human frontier
The forecast is an extrapolation of those trends, and he framed it as a prediction: "And in extrapolating out these trends, we're actually making the prediction that we'll see models eclipse human researchers probably around August of 2027."
The capability gap he describes is specific rather than general. He said the models are very good at executing experiments and operating as engineers, but that they lack the intuition to develop new experiments and have fundamental breakthroughs
The measurement has a boundary that matters for the whole safety debate. He said Vals is testing a set of proxies and the publicly available models: "We're not testing any unreleased model systems. We're not testing specialized agents or even multi-agent systems"
That is the argument for embedding evaluators inside companies, which is what Anthropic proposed — he said the role of evaluators over time could be to get closer to these internal systems
4. Who Pays the Auditor
Ludlow put the commercial problem to him directly: the proposal creates demand for exactly what Vals sells, Anthropic would be the one paying, and he asked how that squares with an independent assessment of an unreleased model.
Krishnan did not claim his model is the answer. He said it is going to take a diversity of approaches, and that it would be a failure to say that this will be the end-all be-all
He reached for the precedent rather than arguing from principle. He said ratings agencies and auditing firms have been for-profit companies responsible for testing and evaluation in other industries, and that natural conflicts arise
The specific conflict he says has to be excluded is remediation: "And so the main one that we want to avoid is a situation where the same group auditing is also responsible for remediating or fixing. And then you end up in situations like Enron."
The firm's own boundaries are structural. He said Vals makes its methodologies public but keeps its test sets private, and does not sell training data or the solution to the labs while it runs the testing process
None of it works unilaterally. He said that people have taken different approaches, but "it seems that this is only going to work if all of the frontier labs are able to cooperate with one another to set a standard"
Bonus Insights
Asked what happens next, Krishnan said he has a lot of reason to be optimistic, "I think that people at these labs are rational actors, and so they see that the way things are headed, it will take some coordinated effort"
He does not think the outcome depends on Washington: "And so I think with or without regulation, I think market-based solutions will have a part to play"
Ludlow's introduction carried Krishnan's own framing of the Anthropic proposal, that embedded evaluators will promote an even playing field and public trust
Krishnan's bottom line is that the testing industry has been outspent by the model-building industry for years, that his own index already shows machines closing on human AI researchers, and that the fix is a common standard with the auditing kept strictly separate from the fixing.
Products, Companies & Tools Mentioned
Vals AI (Krishnan's firm, which benchmarks frontier models including unreleased ones and publishes the recursive self-improvement index discussed on air)
Anthropic (Proposed embedding third-party evaluators with near-employee-level access, the concrete item Krishnan was responding to)
OpenAI (One of the labs whose unreleased models Vals tests, and the source of the GPT-6 to GPT-7 illustration)
If this was worth your time, send it to someone closer to the industry than you are.
Get the latest market chatter as it happens:

