Olam Labs measures which AI models lie more, which collaborate better and which are more deceptive than their rivals — a market that did not exist when its founder started his career.
Most AI evaluation still asks whether a model solved a coding problem. Om Buddhdev's company asks how it behaved while doing it.
"I am optimistic. Been optimistic my entire life but like I will say it is getting scarier and scarier."
Om Buddhdev is co-founder and chief executive of Olam Labs, an AI evaluation company that builds behavioral environments and training data for frontier labs.
I listened to the full segment so you can skip it.
Here are the 3 takeaways that matter.
👤 Guest: Om Buddhdev, co-founder and chief executive of Olam Labs, an AI evaluation company building behavioral environments and training data for frontier labs
🎙️ Hosts: John Coogan and Jordi Hays, who present TBPN live on X and YouTube every weekday
📰 Published: 10 September 2026 on YouTube (TBPN)
🔴 YouTube | 🟣 Apple Podcasts | 🔗 Show notes | ⏱️ 9 min
Key Takeaways
Olam Labs grades AI behavior — lying, deception, collaboration — not just task completion
It builds simulated workplaces: Slack threads, a PM and a CEO, negotiation scenarios
Selling training data means selling a capability to a lab, not shopping a dataset around
Publishing public benchmarks is proof to labs that the company can package good data
He named data and compute as AI's only two real bottlenecks, and said compute is worse
Buddhdev said his own team is compute-constrained despite being early-stage
1. Grading Model Behavior
Coogan asked Buddhdev to introduce the company after a callback to his own YouTube-documentary-watching teenage years.
Standard environments test task completion; Olam Labs tests conduct. "We do all environments and training data. We work with labs to evaluate like behavioral things and models." He named lying, deception and collaboration as the traits being measured
The environments simulate a workplace rather than a coding test. He described building real-world scenarios — Slack threads between colleagues, a PM and a CEO in the loop, a negotiation exercise modeled on the game Diplomacy — because those situations are harder to verify than a solved math problem but more valuable once measurable
2. Selling a Capability
Coogan asked how the company's commercial motion actually works — build first, then pitch a lab, or sell access up front.
Buddhdev said the second approach only works with an existing relationship. Without one, "the best mental model is that you are selling a capability to a lab" — something that makes its models measurably better at a specific skill
Publishing evals publicly is part of the sales process. He said data companies post their benchmarks on Twitter partly to prove to labs that they can package good data
Olam Labs is doing both — publishing and selling directly. "We have we released evals and how models all compete relative to each other, how much they lie against each other," alongside private commercial work in the same area
3. Data and Compute's Limits
Hays pushed on the historical comparison for a business like this, and on where the real constraints sit.
Buddhdev compared data companies to internet-infrastructure providers, not to a prior generation of startups — the closest analogy he could offer was fiber and CPU/RAM suppliers building the substrate other companies deliver on top of
His framing of the AI bottleneck is binary. "It's data and compute" — with compute the larger constraint today, though he expects roughly ten times more compute online within two years, at which point data becomes the harder problem again
Hays credited the week's top model release to data discipline, not architecture. Of the new DeepSeek model, he said: "The reason that's, a crazy, like, dominating model is because they spent all the time on amazing data. That's it."
Buddhdev confirmed his own company is compute-constrained despite its size, testing new training tasks by post-training open-source models on GPUs it has to source itself
Safety grading is spreading to companies that never intended to be safety companies. "Normal oral environment companies that are selling, like, coding data, even they're now having to do safety grading," he said, because models increasingly try to reward-hack their way out of a coding sandbox during ordinary tasks
Bonus Insights
Asked for demo-day numbers, Buddhdev said the company is himself and one co-founder: "It's only me and my co founder"
He said he used to watch Coogan's own YouTube company-documentary videos as a teenager, which Coogan found flattering and slightly aging
Buddhdev's bottom line is that as AI models are trusted with more autonomous, ambiguous work, measuring how they behave will matter as much as measuring whether they succeed.
Products, Companies & Tools Mentioned
Olam Labs (Buddhdev's company, building behavioral AI evaluations and training data for frontier labs)
DeepSeek (Its new model release was cited by the hosts as evidence that curated data, not architecture, produced the result)
If this was worth your time, send it to someone closer to the industry than you are.
Get the latest market chatter as it happens:

