Wafer raised $40 million at a $200 million valuation on a product that writes the GPU code almost nobody can write.
Most inference providers buy Nvidia hardware and tune their software for it. Wafer points AI agents at whatever chip is cheapest to run, and says it has already made AMD match Nvidia on two open-source models at around half the cost.
"There's genuinely probably less than I don't want to exaggerate but probably less than like 500 people on earth that are extremely good at doing this type of GPU code writing and kernel writing."
Emilio Andere co-founded Wafer and runs it, and The Information broke the round the same morning this segment aired, along with the acquisition offers that came with it.
I listened to the full segment so you can skip it.
Here are the 7 takeaways that matter.
👤 Guest: Emilio Andere, co-founder and chief executive of Wafer, an AI inference cloud whose agents rewrite GPU code so open-source models run faster on cheaper chips
🎙️ Host: Akash Pasricha, who anchors TITV, The Information's live weekday news show at 10 a.m. Pacific
📰 Published: 1 September 2026 on YouTube (The Information)
🔴 YouTube | ⏱️ 12 min
Key Takeaways
Agents now write the GPU code that almost nobody can write by hand Andere puts the number of people who are genuinely good at it at fewer than 500 worldwide
AMD matched Nvidia on two open-source models at around half the cost GLM 5.2 ran at 80 to 90% of an Nvidia chip's throughput; Kimi K3 beat Nvidia's B200s
Wafer optimizes open-weight models only, and is deliberately chip-agnostic
The chip market is splitting into specialists rather than producing a second Nvidia Cerebras for ultra-low latency, SambaNova for reconfigurable dataflow, and buyers now mixing two vendors in one system
He will not rule out selling, and says everyone has a price
OpenClaw's second release landed with far less noise than its first He uses a rival assistant called Instinct instead, and twice said he has no affiliation with it
OpenAI's own chip contradicts the specialization trend and still performs Stood up in about nine months, with performance charts against Nvidia's and AMD's latest flagships
1. What Wafer Actually Sells
Pasricha opened by pointing readers at The Information's story that morning and asking Andere to state the proposition himself.
Wafer sells inference capacity and uses AI to make that inference faster. "Wafer is an AI inference cloud and the proposal is simply to use AI to Optimize AI." The mechanism is agents doing engineering work: "So we've been building from the beginning systems so that agents can do the menial work and very complicated work of optimizing GPUs so that LLMs run very efficiently on those GPUs." His argument for why that is obvious: a sufficiently capable AI should be able to optimize itself
The company works only on open-weight and open-source models. Asked whether that was the whole focus, Andere said "We are fully focused on that."
Pasricha checked on air that the company does not make silicon wafers. It does not — Andere said Wafer sits at the software layer
The name is a marketing problem and he knows it. "It's not very good for SEO, I can tell, or you know, for AEO, whatever, whatever we're using these days to find companies." He added: "Uh, googling wafer is just makes you hungry. That's how I know."
2. AMD at Half the Cost
Andere said Wafer was the first to show publicly that AMD silicon can match Nvidia when the software is right. "So, this was one of the first time that anyone has ever publicly showed that AMD if you optimize the software correctly can actually run at similar performance than um Nvidia." He framed the same point as a cost-of-ownership claim, not only a speed one
Two open-source models carry the evidence. On GLM 5.2, Wafer reached about 80 to 90% of an Nvidia chip's performance at half the cost. On Kimi K3, a model he said had come out recently and become popular, it ran on AMD at similar performance to Nvidia — better than Nvidia's B200s, he said — again at around half the cost
The unit behind the claim is throughput, not a benchmark score. "Yeah, we're measuring um tokens per second. So how much throughput the chip actually outputs per uh node of GPUs which is eight GPUs, right?"
Pasricha put the commercial version of the claim to him directly: that a buyer could take cheaper AMD parts over Nvidia's Rubin generation and, with Wafer's software, get Nvidia-class output. Andere's answer was "Exactly. Exactly."
The host then asked whether the results had been shown in testing, which is what produced the two model examples
3. Agents Replace Kernel Devs
The bottleneck Wafer is attacking is a labor shortage, not a hardware one. "There's genuinely probably less than I don't want to exaggerate but probably less than like 500 people on earth that are extremely good at doing this type of GPU code writing and kernel writing." The people who can do it are expensive and already spoken for: "These are humans that are paid so much money and are paid so much money by the labs by like Anthropic, OpenAI, Nvidia because they're so valuable."
The bet was that coding agents would get good enough to do that specific job. "So our thesis from the beginning was well agents are actually getting really good at coding." Wafer's version of the work is automating what he called the grueling and menial part of making models run fast on chips
4. Not Selling, He Says
The Information reported that Wafer has received acquisition offers. Pasricha asked who had approached the company.
Andere would not say: "I yeah no comment on that."
Asked whether he would sell, he left the door open rather than closing it. "No, not really. I mean or the right price is everyone has a price."
The test he set for a buyer is a shared goal, not a number. "I mean I guess the thing that we really care about is like what we call maximize intelligence per watt." He said he is excited about "a future of intelligence too cheap to meter," and that there is probably some future where Wafer aligns with a company that believes the same thing
His stated reason for staying independent was that the technology is working and the job is enjoyable
5. The Chip Field Is Widening
Pasricha said customers are looking at a mixed set of chips and asked Andere to size up the market, naming Cerebras, SambaNova and Google's TPUs as the alternatives the show keeps returning to.
Andere did not dispute Nvidia's position. "So Nvidia is definitely still dominant to be clear like they're they're amazing at what they do and they've been the market leaders for many years now, but you do see this sort of gap getting um shorter and shorter into as you're saying um as to what the other chip providers can do."
AMD is the closest challenger, and still behind out of the box. "So I would say the second biggest sort of closest contender to what Nvidia can give you today is AMD."
The more interesting competitors are not trying to be Nvidia at all. Cerebras builds a wafer-scale chip and is his example of ultra-low latency; SambaNova uses a reconfigurable dataflow design He described their pitch as accepting a trade: "Hey, maybe I'm not as good as a GPU in all these general case scenarios, but I can be very specialized in these particular directions."
Buyers are already combining vendors inside one system — AMD with Cerebras, Nvidia with Cerebras — and he pointed to Nvidia's purchase of Groq as the same idea, running different parts of a model on different silicon
His forecast is more chips and more choice, not consolidation. "I'm very very excited about the heterogeneous world um that will become in the next like couple years of GPUs and also just of people having more choice um to choose their chips."
Wafer's own roadmap follows that: he named Nvidia, AMD, Google TPUs, AWS Trainium, SambaNova and Cerebras as chips he expects to use
6. OpenClaw's Buzz Faded
Pasricha said OpenClaw 2.0 had shipped that week to noticeably less attention than the original, and asked whether the tool has fallen out of favor.
Andere said he has not spent much time on OpenClaw itself, but uses that style of agent daily
The one he named is Instinct, and he disclosed twice that he has no stake in it. "Uh I have no affiliation to Instinct." He described it as an agent you text on your phone, connected to your accounts, that runs the same kinds of workflows OpenClaw supports The example he gave was asking it to find a standing Saturday reservation for two people
Set-up cost is the reason he thinks the hosted version wins. "It's still pretty hard to set up OpenClaw uh for most people in the world, right?"
On which one lasts, he separated the projects by who owns them. OpenClaw was open source and its creator, he said, was bought by OpenAI to work on internal projects. Instinct is "a startup whose sole purpose is to make this happen for the world," which is why he is more confident about its durability
7. OpenAI's Own Chip
Asked about OpenAI's jalapeno chip, Andere said it cuts against the argument he had just made for specialized silicon.
The design puts the whole inference job on one chip. "They basically just decided to um put it all in the same chip, similar to how um GPUs used to do it."
He said the published performance holds up against the newest parts from both incumbents, comparing it to Nvidia's and AMD's latest flagships
The schedule impressed him more than the silicon. "So, first of all, like the quick things are first of all, insane performance by the OpenAI team and the fact that like they were able to stand up the hardware." He put the timeline at "I think it was in like 9 months, some absurdly fast number."
Bonus Insights
Pasricha opened the programme by trailing all three of the day's segments: the Wafer round, Nvidia's $3.5 billion convertible investment in MediaTek, and a venture capitalist's case for raising less money in the AI era
He described the Wafer funding round as The Information's exclusive reporting that morning, and credited the article to the reporter who wrote it rather than to the show
This was the first of three interviews in one episode; the other two are written up separately
Andere's bottom line is that the scarce input in AI inference is not chips but the handful of engineers who can make chips fast, and that automating those engineers is what lets a buyer treat AMD, Cerebras or a TPU as a substitute for Nvidia.
Products, Companies & Tools Mentioned
Wafer (Andere's company: an AI inference cloud whose agents rewrite GPU kernels, newly funded at a $200 million valuation and fielding acquisition offers)
Nvidia (Still dominant on his account, and the benchmark for everything else — his AMD results are quoted against its B200s and the host raised its Rubin generation)
AMD (The closest challenger, which he says ran two open-source models at Nvidia-class performance for roughly half the cost once Wafer's software was applied)
Cerebras and SambaNova (His examples of chipmakers attacking Nvidia from a specialized angle — wafer-scale and ultra-low latency in one case, reconfigurable dataflow in the other)
Groq (Bought by Nvidia, in his telling, and combined into a single system that runs different parts of a model on different silicon)
Google TPUs and AWS Trainium (Named alongside AMD and the specialists as chips Wafer expects to support)
OpenAI (Its own chip puts the whole inference job on one die, which he said runs against the specialization trend while still posting strong numbers)
Anthropic (Named with OpenAI and Nvidia as the labs bidding up the few hundred engineers who can write GPU kernels well)
OpenClaw (The open-source agent whose 2.0 release the host said landed quietly; Andere says it remains too hard for most people to set up)
Instinct (The hosted alternative he uses and recommends, twice noting he has no affiliation with it)
GLM 5.2 and Kimi K3 (The two open-source models Wafer ran on AMD hardware to make its cost claim)
Books & Resources Mentioned
Wafer, An Inference Provider That Uses Non-Nvidia Chips, Lands Acquisition Offers and $200 Million-Plus Valuation (The Information's story that morning, which the host told viewers to read and which this segment walks through)
If this was worth your time, send it to someone closer to the industry than you are.
Get the latest market chatter as it happens:

