All-In Sep 21, 2026 23m 8m saved
With Naveen Rao, co-founder and CEO of Unconventional AI
Google alone crosses 3.2 quadrillion tokens a month. At 10 joules a token, which Naveen Rao called the low end of the range, that is 12 gigawatts of continuous power for one company's AI services, against about 40 gigawatts going into every data center in the United States.
Most of the response to that arithmetic is to build more power stations. Rao's company is trying to build a computer that does not need them, by throwing out the memory interface that every machine since 1945 has been organized around.
"So, we're going to run out of energy pretty fast in like 3 years or so is my estimate."
Rao founded the first AI chip company in 2014 and sold it to Intel, then built the model-training platform that became a quarter of Databricks' revenue. He used this talk to show the first results from a chip his team sent to the fab on June 1, which he says is the first physical dynamical computer ever built.
The full episode is covered here so you can skip it. 23 minutes of audio, 15 minutes of reading.
Here are the 10 insights that matter.
Key Takeaways
Energy is now about half the cost of serving a token, and it is what data center operators sign for first Floor space, then networking, then GPUs, now the power contract
One company's AI services need 12 gigawatts, against roughly 40 gigawatts for all US data centers
A monkey's brain runs on one watt, the same as a mobile phone A human cortex moves about 16 billion bits a second; a high-end computer moves nearly 30 trillion
Moving bits, not computing on them, is where the energy goes — which is why a faster chip is not a cheaper one
The company built a working dynamical chip in five months, taped out on June 1 and already producing images Rao says it is the first physical dynamical computer ever built, and this was the first public mention of it
Throwing away connections made the system better, not just cheaper — sparsity improved trainability and performance at once
The target is 1000x power efficiency, and the timeline moved in, from five years to three and a half
Animal brains sit within one or two orders of magnitude of the thermodynamic limit; today's computers are about 10 billion times away
A product is about two years out, sold as a rack that takes tokens in and tokens out over a network cable
Jevons paradox is the business case — cut the cost of intelligence by 1000x and consumption rises more than 1000x
1. From Neuroscience To Chips
Rao opened with his own route into the problem, which is unusual enough to be the reason he is on stage. His family had a computer in 1978; he learned to program as a child because it looked like a puzzle; he became an electrical engineer because he read science fiction and wanted to build an intelligent machine. Then, after a career building computers, he went back and took a PhD in neuroscience to attack the same question from the other end.
The company history is short and specific. He founded the first AI chip company, Nervana Systems, in 2014, when there was no AI in common usage and hardware for it was a hard sell. He notes that Jensen Huang had been on the same stage earlier that day, running the largest company in the world, which is a hardware company because of AI.
On the exit
I sold the company way too early to Intel but I ran I started and ran the AI group at Intel.
Naveen Rao
After Intel he moved to the infrastructure problem, which was how to build large language models at all. He started packaging GPUs into something other people could use to train their own. ChatGPT's release in 2022 made that the best game in town, and the business joined Databricks in 2023.
What that business is worth inside Databricks now
that's a quarter of the total revenue of Databricks today
Naveen Rao
Unconventional AI, the current company, is organized top to bottom around one number. It starts with theorists, meaning mathematics PhDs and theoretical neuroscientists, who propose ways to move less information around. Those ideas become models trained on real data and tested against real criteria, then circuit designs, then physical systems and eventually a product.
The target, and the date that moved forward
I've actually revised this to three and a half years because things have gone faster than we anticipated.
Naveen Rao
The original goal was 1000x power efficiency within five years. He said the schedule came in because deep scientific problems were solved faster than planned, partly with the help of AI itself.
2. Energy Is The Constraint
Rao used Google as his worked example, because Google has published the figure.
The token volume he is scaling from
Per month they cross 3.2 quadrillion tokens.
Naveen Rao
At 10 joules per token, which he said is at the lower end of the energy range for models, that comes to 12 gigawatts. For comparison, he put United States data centers as a whole at about 40 gigawatts, and said the US is about half the world's data center capacity. So one company's AI services account for a substantial share of American data center power on his arithmetic, and both sides of it are still growing: bigger models raise the energy per token, and demand raises the token count.
His estimate of when the ceiling arrives
So, we're going to run out of energy pretty fast in like 3 years or so is my estimate.
Naveen Rao
He framed it as two lines on a chart — a market growing exponentially toward roughly a trillion dollars by 2030, and energy supply growing in a straight line underneath it. The gap between them is what he wants to fill with hardware rather than with generation.
The shift in how data centers are procured is the part an investor can check. It used to be floor space, then networking equipment, then GPUs. Now the power contract comes first and everything else is fitted to it.
Energy is roughly half the unit cost of inference
So every time you try something on ChatGPT 50% of that cost is energy.
Naveen Rao
The rest is the capital cost of the hardware and the building. That is why his pitch is stated as a multiple on a power contract rather than as a chip specification.
The business case, in one sentence
our business case is pretty easy we're going to monetize that a thousandx better than existing hardware
Naveen Rao
3. What Biology Costs
The evidence that a 1000x gain exists at all is that nature already built it.
The number everyone half-knows
the human brain, you may have heard this, runs on about 20 watts of energy
Naveen Rao
What he finds more striking is the animal end of the scale. A monkey's brain, scaled by neuron count, runs on about a watt — the same as the phone in a listener's pocket. Rats and bats run on milliwatts. His chosen example was the squirrel, which jumps between branches and lands it a thousand times out of a thousand on a brain running on single-digit milliwatts.
What that means about a mobile phone
You could run over a 100 squirrel brains on your phone and it has very precise and accurate behavior.
Naveen Rao
His conclusion is not that brains are mysterious but that they are the right physical substrate for the job, and that synthetic systems have reached similar capability by a far more wasteful route.
The line he built the company around
I don't feel like we truly understand something until we can create it.
Naveen Rao
4. Bits Are The Energy Bill
The inefficiency has a specific location: most of the energy in a computing system goes into moving information, not into the computation itself. Rao put two numbers side by side to show the scale of it.
What a cortex moves
the human cortex, the squiggly part of your brain, the outside of it, only moves about 16 billion bits per second
Naveen Rao
He noted that is a small number against the 13 or 14 billion neurons in the cortex — roughly one bit per neuron per second.
What a computer moves
A GPU or, you know, a high-end computing system moves nearly 30 trillion bits in and out of memory per second.
Naveen Rao
That is traffic outside the chip. Inside it, he said, the figure is probably ten to a hundred times more. The ratio between those two lines is the energy problem stated exactly.
5. Why Computers Stalled
Rao's account of how the industry arrived here is that the machine was optimized for the wrong thing for eighty years. Computers were mechanical, became analog around the turn of the century and digital in the 1930s and 1940s, and a 1945 machine operates in substantially the same way as a 2026 one: memory on the outside, computation in the middle, bits moving back and forth.
He used ENIAC to make the point about what the design was for. It was built to calculate artillery trajectories faster than the alternative, and the alternative was human beings doing the arithmetic. Machines have been sold on speed against the previous machine ever since, and nothing in that sales pitch contemplates energy.
Transistor counts kept rising while clock frequency and single-thread performance stopped scaling, and efficiency has now stopped scaling too. Making transistors smaller, he said, has largely ended as a source of gains.
His diagnosis of the remaining waste is the stack of abstractions. Digital ones and zeros are themselves an abstraction — a transistor has states in the middle, and engineers force it to behave as though it does not. Every layer built on top of that is lossy, because an abstraction by definition does not carry the complexity beneath it.
What his team is doing instead
We find an abstraction of the physics of the semiconductor and connect that to the neural network.
Naveen Rao
The analogy is to the brain, which has no linear algebra and no floating-point arithmetic in it. The physics of the neurons is what produces the behavior.
6. Metronomes That Compute
The theory underneath the chip is dynamical systems, which Rao pointed out has been around for a hundred years. The idea is that simple rules followed by individual components produce complicated collective behavior: birds in a flock each look left and right and the flock emerges; ant colonies do intelligent things without any ant being intelligent.
His physical demonstration is metronomes. Put several on a plank that can roll slightly and they synchronize, because each one pushes against the plank and the plank couples them. Scale it to hundreds and they still synchronize. Vary the coupling and you can get half in one phase and half in another.
The company's first public artifact from this was a model called Un-0, an image generator built on a set of oscillators, simulated and released open source.
What the model proved
this was a f the first demonstration that I can actually scale something up, train it and actually get useful output like image generation
Naveen Rao
The analysis he showed was what he calls a state space trajectory: characterize the state of the system as the phases of all the oscillators, condition it on what you want generated, and watch it take a different path through that space for an airplane than for a car or a bird.
7. Sparsity Buys Performance
The scaling problem in any all-to-all connected system is quadratic. Ten elements fully connected is a hundred connections; a thousand elements is a million. Rao's team asked whether connections could be discarded and the behavior of the whole system rescued.
The answer was better than that. Throwing connections away did not merely preserve the behavior — it improved it, and made the system more trainable, in simulation and in physical hardware.
Why he calls this the unusual result
So, it's one of these rare things where you get something that's more efficient, that's actually more scalable and even gives you more performance.
Naveen Rao
He described it as a holy grail that had been a problem for a long time, and said the breakthrough came from framing the question correctly rather than from new equipment.
8. The Chip Is Back
This is the part Rao had not said publicly before, and he said he chose this venue for it.
The claim
This is actually the first physical dynamical computer ever built. We did this in 5 months.
Naveen Rao
The company started in earnest in January without a full team. The design went to the fab on 1 June, and the silicon is now in the lab with results from it.
The timeline, in his own words
we taped out the design, meaning we sent it to the fab in June on June 1st. The chip is back in our lab, and we actually have results from it.
Naveen Rao
What he showed
So, these are the first ever images generated from such a computer.
Naveen Rao
Images are the demonstration rather than the point — he said the same machine can do sequence modeling or language models. The number that matters is the energy per image, which he put at roughly 500 nanojoules against the order of millijoules on a conventional GPU.
Why the gap is that large
it's many orders of magnitude more efficient than a standard computer and it's because it just doesn't move information around
Naveen Rao
And his summary of the result
This is proof positive this works.
Naveen Rao
9. What 1000x Would Do
The architectural claim is that this is not another step along the same line. CPUs became GPUs and then compute-in-memory designs, each more parallel and finer grained, but all of them are von Neumann machines with a memory, a compute unit and traffic between them.
What is different about this one
We don't have a memory interface. Each individual computing element is a memory.
Naveen Rao
He calls the result 4D computing: three physical dimensions, including stacking dies vertically, plus time as a fourth, used in the dynamics of the system.
The metric he wants the industry judged on is intelligence per watt, and he anchored it to a physical ceiling. There is a thermodynamic limit that cannot be exceeded. Mammalian brains sit within one or two orders of magnitude of it.
How far today's machines are from that limit
Today we're on the far left of this graph and we're about 10 billion times away.
Naveen Rao
His three-and-a-half-year target is to reach the limits of 2D lithography.
The company's stated ambition
the overarching goal of this company is to beat biology
Naveen Rao
What he expects that to change is the shape of the industry. Gigawatt data centers give way to many small ones, closer to where they are used and more adaptive, which he argues is also the more environmentally friendly outcome. Beyond that he pointed at robotics — billions of robots that dynamically assemble to solve problems.
On the market, he turned the disruption argument the opposite way to the usual one. If AI is a trillion-dollar market and the cost of it falls by 1000x, the revenue does not fall with it.
Jevons paradox as the business plan
if you make something half the price, you'll consume more than 2x
Naveen Rao
Where he thinks that ends up
I think this will create the largest market that humanity's ever seen.
Naveen Rao
10. The Ecosystem Question
Chamath Palihapitiya took the stage afterward and started with what he called the thing on everybody's mind: a chip this different needs fabs, packagers and an ecosystem around it before anyone can use it.
His reaction to the talk
I mean that was extremely unexpected. I got to say that was pretty amazing.
Chamath Palihapitiya
The timeline Rao gave
Yeah, timewise we're within two years of getting it to a full product.
Naveen Rao
Asked what the product actually is, Rao described a rack-scale system rather than a chip anyone buys loose — something that sits in a data center under someone else's management.
The interface is deliberately ordinary
tokens in, tokens out through a network cable, but the inner guts are completely different than an existing computer
Naveen Rao
Palihapitiya pressed on the switching cost, which is the real question for anyone holding the current architecture: the industry has built years of work on top of these abstractions, from KV caches down. Rao's answer was that migration effort scales against how much better the alternative is, and that he chose to make the gain large enough to be worth the pain.
Where the port happens
We actually don't port at the operations layer, you port the model layer.
Naveen Rao
Existing models will run, but getting them across takes real compute, and basic primitives do not exist in the same form — matrix multiplication is not implemented as matrix multiplication but as time-varying behavior, which can be analyzed as a current state matrix multiplied by a transition matrix.
On hiring, Rao named the specific organizational problem.
Two populations who have never worked together
we got people from that world and then we got people who actually build chips and they don't talk to each other
Naveen Rao
Getting that span of talent to coordinate on one thing, he said, is among the hardest parts of running the company.
Asked what plays the role CUDA plays for Nvidia, he pointed at a set of libraries his team wrote.
The developer layer
It's not CUDA, but it's a language of sorts that allows you to kind of express time varying elements
Naveen Rao
Bonus Insights
Why he was on this particular stage
Rao opened by placing himself against the room's theme rather than hedging, and it set the tone for a talk whose whole argument is that the constraint is physics rather than risk.
I'm the opposite of a doomer.
Naveen Rao
The claim underneath that
I think AI is one of the most transformational technologies that humanity's ever created and will enable us to get to that next level of evolution which I'm here for
Naveen Rao
A point about timing that is easy to miss
Rao founded an AI chip company in 2014, when the phrase barely existed outside research groups and, on his account, convincing anyone the field mattered was harder than building the hardware. The same observation cuts both ways on his current company: he is again selling a hardware thesis before the market has agreed there is a problem.
Rao's bottom line is that the industry has been buying speed for eighty years and is now paying for power, and that the way out is not more generation but a machine whose computing elements are their own memory — which he says his team has now built and run.
Products, Companies & Tools Mentioned
Unconventional AI (Rao's current company, rebuilding the computer from first principles for power efficiency; target is 1000x, timeline revised from five years to three and a half)
Nvidia (The largest company in the world, and a hardware company because of AI — Rao's evidence that the 2014 thesis was right; CUDA is the comparison for his own Python libraries)
Google (The worked example for the energy argument, because it has published the figure: 3.2 quadrillion tokens a month, which he converts to 12 gigawatts)
Intel (Bought his first AI chip company, which he says he sold way too early; he then started and ran Intel's AI group)
Databricks (Acquired the model-training platform he built in 2023; he says it is a quarter of Databricks' revenue today)
ChatGPT (The 2022 release that made his training platform the best game in town, and his example of where half the cost of a query is energy)
Books & Resources Mentioned
Un-0 (The open-source image generation model built on coupled oscillators, released so people could play with it; the first demonstration that a dynamical system could be trained and produce useful output)
Listen to the full episode
🔴 YouTube | 🔗 Episode page
Watch the full episode:
If this was worth your time, send it to someone closer to the industry than you are.
Get the latest market chatter as it happens:


