OpenAI took a quarter of its production engineers off whatever they were building and reassigned them to security, pointed its own models at its own infrastructure, and kept going until the models stopped finding critical bugs.
Most of the argument about AI risk is about what future models might do. Greg Brockman's argument is that the capability already exists, that it will be widely available shortly, and that the only advantage defenders have is the gap before it is.
"Sorry, all your projects are on hold. You are now defending. You are now upleveling our security architecture."
Brockman co-founded OpenAI, is its President, helped build Stripe before that, was in the room at the November 2015 offsite where the company's three-step plan was written, and spent the Christmas of 2019 alone with a freshly trained GPT-3 instead of going on holiday.
The full interview is covered here so you can skip it. 51 minutes of audio, 23 minutes of reading.
Here are the 15 insights that matter.
👤 Guest: Greg Brockman, Co-founder and President of OpenAI, who helped build Stripe early in his career
🎙️ Hosts: Ben Horowitz, Co-founder of Andreessen Horowitz, and Erik Torenberg
📰 Published: 14 September 2026 on YouTube (a16z)
🔴 YouTube | 🟣 Apple Podcasts | ⏱️ 51 min | ✅ Time saved: 28 min
Key Takeaways
He and Ilya Sutskever did the compute arithmetic in 2016 and landed on roughly 15 years to AGI, or 10 if someone spent hundreds of billions of dollars
The constraint he expects to bind is not model capability but the compute to serve it affordably to everyone
Safety, security and alignment standards are the other bottleneck, and he says people are underestimating both
The Hugging Face incident is his evidence that cyber-capable AI in the hands of attackers is a near-term certainty, not a scenario
His conclusion is that defenders have a window, because a defender who finds a vulnerability can patch it
OpenAI reassigned 25% of its production engineers to security and ran its own models against its own systems until the findings saturated
He says other chief information security officers have told him the same thing: serious issues found, and fixable
OpenAI ran 10,000 agents at the Navier-Stokes problem and formalized the result in Lean, which he treats as a route to formally verified software
He calls the current phase the AGI era on the strength of computer use, with models running coherently for 24 hours
He still calls the capability jagged: the writing is the first that is not slop, and is still not good
Employment has gone up as AI has improved, on the numbers he has seen, and he expects a wave of new firms rather than fewer jobs
About 1.5 billion people have tried ChatGPT and stopped, against 1.1 billion who use it weekly
His conclusion is that the product should tell users what it can now do for them, rather than making them discover it
OpenAI canceled Sora this year to concentrate on the agentic coding cycle, which he called very painful and necessary
1. The 2016 AGI Timeline
The hosts opened by asking whether Brockman, ten years ago, would have predicted today's breakthroughs. His answer was that he and Ilya Sutskever had actually done the arithmetic.
They modeled it on compute, not on research intuition. "And we kind of came to the conclusion that if you look at Moore's Law progress, that kind of thing, 15 years felt like about the timeline to AGI"
The faster branch of that estimate assumed a spending decision. If you were willing to scale up, build massive supercomputers and spend the hundreds of billions of dollars, he said, maybe it would be 10
He described what has happened since as the conclusion of converging forces rather than a surprise. Taking the macro view, he said, it makes sense that it is happening now
He was careful not to treat the moment as finished. He called it a moment for everyone to be a part of and to help shape collectively
2. Compute Is the Real Limit
Asked whether the timeline still holds given the shortages now appearing in the supply chain, Brockman reframed the question away from capability.
He accepts that compute cannot keep up with demand already visible in the market. The difficulty is not making models more capable but scaling their raw potential to everyone
The models will keep improving; the distribution will not keep pace. He said they will be plenty powerful and will continue apace, but it will be hard to get them to everybody in an affordable way because there will not be enough compute to serve it all
He named that as the underrated problem. "I think that's going to be a huge challenge people are underestimating"
He tied it back to the company's stated mission — empower everyone, ensure it benefits everyone — and said that half of the problem deserves a lot more airtime than it has had
3. Pacing the Frontier
The same answer introduced the phrase Brockman used repeatedly through the interview, and the argument for coordination between labs that follows from it.
"And I do think we're at a point now where we have to really start thinking about what we call pacing the frontier"
The standards have to rise with the models. "And so thinking about as we move to more capable models, you really have to make sure that safety, security, alignment, those are all standards that you're constantly upleveling"
Those standards, not compute, become the limiting step in his account. He said they almost become the bottleneck to progress — the part you have to spend a lot of effort on to make sure you have gotten right
On coordination between the frontier labs, he was direct. "I think there's nuance here and I do think the coordination is going to be a very important theme"
The hosts pushed the point, asking whether safety techniques would be shared across labs on related architectures or each developed independently
He argued the operational questions are genuinely new. Making safety cases for training, developing and evaluating these models is unprecedented: "No one's ever really had to operationalize this before"
His framing of why no single company controls the outcome is about where the technology comes from. He described what they are building as something that falls out of compute progress, which in turn falls out of technological progress — a wave that has been building a long time, whose leading edge labs can see slightly early
"But we can't do that alone"
4. Safety Ideas From 2017
Asked whether making a model structurally safe — rather than filtering its outputs — is a different and harder category of problem, Brockman said the groundwork is older than people assume.
"I absolutely think we can and are making very rapid progress on this problem"
He pointed back to 2017 for two results he thinks are underappreciated. One is an early paper laying out the first inklings of modern language models; the other is reinforcement learning from human preferences, created the same year to align a model to what people want by taking feedback from them
The scalable-oversight ideas are from the same period. In 2017 and 2018, he said, they were already asking how you supervise something very capable and provide feedback that keeps it aligned with you — and had ideas including debate and iterative amplification
His point is that these were developed before the systems they were meant for existed, and are now visibly trickling down into modern systems
He described the public conversation as having moved away and then back. Early OpenAI communication was squarely about AGI safety; a middle period pushed questions such as whether the AI is politically neutral to the front; and the older ideas are now taking the main stage again
"And I think we've been sort of thinking about this moment for a long time"
The hosts characterized the early approach as surface-level — filters and reinforcement learning from human feedback around the edges, on the reasoning that someone determined to extract bad words could do so and it did not much matter — and contrasted that with a model good at cyber attack, where the model itself has to know not to reward hack in a dangerous way
5. The Defender's Window
Asked to explain why he called the OpenAI–Hugging Face incident a watershed and said the defender window is now open, Brockman gave two takeaways and then the strategic argument.
The first lesson was internal. He said it forced changes to how OpenAI monitors, sandboxes and controls models during evaluation, and that his team has changed a great deal of its internal standards and implemented controls he considers critical for more capable future models
The second was a preview. He called it an insight into what future capabilities look like once they are broadly diffused and in the hands of threat actors — which he said will happen
He was explicit that broad diffusion is also good, because concentration of that power in one or a few entities is its own large risk
What the incident actually showed: "And in the case of Hugging Face, you saw both an AI that was able to hack out of a secure environment and hack into a company's production environment" — and he said what it found was quite sophisticated
The window exists because the capability is dual use. An attacker finds vulnerabilities; a defender finds the same vulnerabilities and patches them
"If you're a defender, you control the battleground, right? You control the setup of your systems"
His description of the typical defender is unflattering and is the reason the window matters. By default, he said, an organization's security is probably static and has been for the past five or ten years
The play is to use differential access now. Defenders who can reach frontier capability through trusted access programs should use it to move themselves up, so that as frontier capability improves they are pulled along with it
The hosts added a third lesson and a harder problem. The good news is that ten thousand agents can organize themselves and do useful work; the bad news is fifty years of code, architecture and deployment practice built for a different world, plus large centralized stores of consumer data that individuals cannot protect themselves
They asked whether a decentralized consumer architecture becomes necessary, and whether today's centralized data repositories remain viable at all
6. The Navier-Stokes Result
Brockman answered the ten-thousand-agents point with what OpenAI did with ten thousand agents.
OpenAI used 10,000 agents to solve the Navier-Stokes problem, which he said matters both for its own applications — fluid dynamics, ocean currents — and for what it represents
What it represents, in his account, is new knowledge created by AI, and with it a wave of scientific discovery and medicines he says are now on the table
The part he thinks generalizes is the formalization. The problem was formalized into Lean, so the AI could write verifiable code
He connects that directly to formal verification of software, a long-standing ambition the hosts described as a dream that never took off because it was intractable for people
"But we have these that are solving these crazy impossible math problems"
7. OpenAI's Defense Factory
The same answer turned to what OpenAI did to itself, which is the interview's most concrete claim.
A quarter of production engineering was reassigned. "Sorry, all your projects are on hold. You are now defending. You are now upleveling our security architecture"
The models were pointed at OpenAI's own systems, found a number of serious issues, and those were fixed
The finding he treats as encouraging is that the search terminated. When they pointed Astra at their systems it found new problems and then saturated — to their knowledge, every critical problem Astra is capable of finding was found
The hosts made the obvious objection: there will be a new model, and a new round. He agreed
The conclusion is a permanent loop rather than a one-off audit. Every time a new cyber capability drops, you deploy it against your own systems and find the new holes
OpenAI is automating that loop internally and calls it the defense factory — find the vulnerability, triage it, remediate, deploy, validate, end to end
"And if you can do that at machine speed, I think the defenders will be advantaged in deeply significant ways"
He paired the optimism with a deadline. "So I think there's real hope but I think that our view is that the world needs to act with urgency because we're in a very dangerous window right now"
On why the sector is behind, he was blunt about resourcing. "I've never met a CISO who felt that they were appropriately resourced, right? That they were appropriately prioritized. Never" — and the hosts added that this is especially true in the public sector
He agreed that society has let technical debt pile up, and said this should have been fixed years ago but that there is now both the motivation and the ability
8. A Pen Test of His Own Site
Asked what the public narrative about the incident got wrong, Brockman named access, then illustrated it with something he did himself.
His first answer was about who can use these capabilities. The capability exists but sits inside a small number of frontier companies, reachable only through trusted access programs — so scaling the number of defenders with access is, in his view, a task for the field and for society
His argument for urgency is that every day matters and using those days requires the tools
He cited Hugging Face's own response as evidence. They said they used frontier models to analyze the attack logs, because that is the only way to analyze an attack of that kind, and that the frontier models refused — but, he said, they did not try OpenAI's, and believe OpenAI's would have permitted it
His point from that is about the default stance of model providers, not about any one company
The personal example is a security test of gregbrockman.com, which he described as a very simple static site and not the most popular website
"So I took my Codex and asked it go check out gregbrockman.com. Tell me if there's any vulnerabilities. So I did a pen test and it came back with 13 findings"
The findings were individually small — an SPF record that did not stop email spoofing, pages served over HTTP without forcing HTTPS — but his concern is an AI chaining small vulnerabilities into a large one, and he did not want a hole that let someone scoop his email
"So 15 minutes for it to find these 13 findings. But then I asked it, can you fix these?"
"And so 45 minutes it opened up my Cloudflare control panel" — it set the headers, migrated the site to Cloudflare Pages, and started the DMARC process, which requires a 48-hour wait
It then set an automation to check back in 48 hours and complete the process on its own
9. Astra and Computer Use
Asked what is most groundbreaking about Astra and what is left to do, Brockman started with why the version number moved at all.
The numbering had been stuck because progress was incremental. OpenAI wanted GPT-6 to represent something worthy of the number, and "And the problem we always have is that our models are kind of incrementally getting better"
This was the first time several research bets landed at once and produced an almost discontinuous step
Computer use is the headline, and the reason is tooling. For agents, everything comes down to whether the model is smart enough to use the tools and whether it can reach the context it needs through them
His criticism of the current approach is that it retools the world for the machine. Building MCP servers and command-line interfaces makes software accessible in a stilted way that was not meant for humans, and adds another layer with its own security problems
The alternative was in OpenAI's founding plan. "We actually laid out this three-step plan that basically is what we ended up following for the next 10 years" — and at that November 2015 offsite in Napa they also discussed reinforcement learning where the environment is screen pixels, keyboard and mouse
The appeal then and now is coverage: anything a human can do with a computer is in distribution
Early attempts to build agents that could do it were aborted, and it took until now
The uses he finds most striking are ordinary. People handing the model a screenshot and getting a 3D model out of Blender, or redesigning a living room
His larger claim is about the work that disappears. Clicking through menus and typing into spreadsheets is not something anyone did a hundred years ago, and he does not think it is crazy that nobody is doing it in five or ten years
The framing he used is physical: no more carpal tunnel or hunched shoulders from people contorting to the machine
10. Why Employment Went Up
The hosts noted that OpenAI has been more sober than most on employment, and asked how he thinks about it.
The empirical claim came first. "And so far, at least in the numbers, the better AI gets, the higher employment goes, not the lower"
His prior is that AI does not play out the way the logic suggests. He said OpenAI put that in its 2015 launch post, and that the history so far has been that the obvious conclusion does not arrive
He thinks people underrate the depth of most jobs — relationships, accountability, setting goals and owning outcomes are things he says feel fundamental and worth preserving
He expects a wave of new firms rather than fewer jobs. He said he has heard from someone whose industry is seeing people quit to start their own firms, specifically because the AI tools let them do so much more, and that barriers to entry are falling
The hosts gave the version from inside their own business. Junior staff get the grunt work by default; if the AI takes the grunt work they can develop faster and get to the real part of the job — the relationship with the entrepreneur — instead of spending a weekend writing an investment memo
They added that Astra is very good at writing investment memos
He refused the rosy version. He said it will be a nuanced story, it will be hard, and there will be change, but that the world can be much better
His historical analogy is the plow, which he noted put a great deal of human labor out of work and produced the Luddite movement: "Nobody here wants to go back to 1870"
He was equally clear that the speed of change is frightening people and that the company spends a lot of time trying to understand how people feel
11. Why US Sentiment Is Worst
Asked why sentiment toward AI is higher in Asian and European countries than in the United States, Brockman put the blame on the field's own communication.
His diagnosis is that the benefit has been argued at the wrong level. The field and the company, he said, need to explain far better why this is good for an individual person and not only for the country
On the strategic framing he was unambiguous, calling the technology rapidly the single most important strategic priority and resource for the United States
The usage numbers are the scale argument. "You look at ChatGPT 300 million health queries or 300 million people every single week using it for help, right? That's a huge deal"
"And we're at a billion almost, 1.1 billion weekly active users" — of which he put the US at around 100 million, roughly a third of the population, using it weekly
The case he thinks goes untold is medical. He described his wife's health conditions and said the family does not really know how they would have managed them before ChatGPT — the toil of reaching the right answer, and of understanding what a doctor has told you
He told one story at length. A friend in hospital was about to be injected with an antibiotic, typed the situation into ChatGPT, and was told not to take it because of a condition she had had a year earlier. She showed the doctor, who agreed: "I only had five minutes to read your chart"
He said he hears such stories every day, alongside people running small businesses on the product who could not otherwise do it, and argued they need to be in public consciousness
The hosts turned that into a positioning problem. The narrative is a teacher, a doctor, a lawyer and a therapist in your pocket — while also telling teachers, doctors, lawyers and therapists that the same tool makes their own work better
On why other countries are warmer, he pointed at demographics. In many of them, he said, it is more keenly felt that a large older generation will have to be supported by a smaller younger one, which makes the potential of the technology more urgent
12. The Data Center Argument
The conversation moved to data centers, and here the transcript does not reliably separate the two voices, so the policy argument below is reported without attribution to one speaker.
The argument made on the show is that banning data centers would cost the US its lead by driving them overseas, the way silicon manufacturing moved in the 1980s
The employment case offered was blue-collar. One data center provider, Switch, was said to employ around 45,000 people on a union contract basis to build them
The proposed policy was conditions rather than prohibition — requiring a data center to be well behaved on power, water and noise instead of banning the category, on the reasoning that AI continues without the US and the US then has no say
Brockman's own contribution was OpenAI's commitments, which are unambiguous in the transcript. "So we've made commitments on not increasing people's electricity bills. Our data centers are all closed loop water"
"So the amount of water used by Abilene which is the data center that actually trained Astra that it uses about the same amount of water as an office building"
He listed community commitments in Ohio and Georgia, where OpenAI has data centers, and credits for every college student to access Codex
"But again, I think that we need to do even more"
13. $1B for Frontline Defenders
Asked about OpenAI's billion-dollar commitment to frontline defenders, Brockman described who it is for.
The principle is that everyone should be using the window. "So we believe that every organization, every company, every government, critical infrastructure such as water service providers, hospitals should all be using this defender window to secure themselves"
The commitment exists because not every organization can pay. The billion dollars gives organizations communities rely on daily access to OpenAI's models to secure themselves
"We think this is the beginning. This is not the end." He named CrowdStrike as a partner in providing discounted access to defenders, and called for a global effort
The hosts framed the stakes in terms of what is already happening. Hospitals were being broken into and held hostage before AI, and water supplies have been hacked by foreign and state actors — critical infrastructure was neither built nor maintained with cybersecurity in mind
14. The AI We Were Promised
Asked whether users notice when a problem gets solved — the hosts had not seen a hallucination in some time, without anyone announcing it — Brockman turned it into a product argument.
He called continual education one of the most important problems OpenAI has. Users should not have to extract from the model what it is capable of; the model should tell them
The number behind that is the lapsed user base. "I think we have something like another maybe 1.5 billion people who have used ChatGPT and don't use it anymore" — against a billion-plus weekly actives, which he noted is a significant fraction of the planet
His point is that those people can be told what has changed and how it can now help them
His illustration of the gap is the interface itself. "It's like this new text box is way better than the old text box, but there's still some things that the old text box is better at"
Which is the setup for the section's claim. "The AI we were promised should be an AI that you talk to over voice primarily" — with persistence, memory, context, knowledge of you, trustworthiness and proactive help in both personal and work life
He extended the analogy to how people learn to work with each other. A new colleague takes time to read, and people do not come with an instruction manual — though consulting firms lean on Myers-Briggs, resumes and back channels to shortcut it
The open question he posed is what the equivalent is for an AI that keeps changing as new models and surfaces arrive
His stated North Star is simplicity: one unified AI, so that the user spends less of their time wrapping themselves around the computer
15. The Year of Focus
The final stretch was about how OpenAI decides what not to build, and how Brockman spends his own time.
"So this year the theme was focus." He said the company realized it cannot do everything, and that the test is which work reinforces the mission of ensuring AGI benefits all of humanity
Deployment and productization passed that test, on the grounds that bringing the technology to bear and having people deploy it usefully is the mission rather than a distraction from it
The filter was the agentic coding cycle. The question they had to grapple with was which areas reinforce that exponential, and which were labeled side quests — even when individually exciting
Sora was the highest-profile cancellation. "so things like Sora that's maybe the highest profile one of these projects that we decided to cancel very very painful by the way not an easy thing to do but it was so critical to unleash the business in many ways so we could really focus"
Bringing the consumer and enterprise sides together into chat work was the other place focus was imposed
He described the first half of the year as unflattering. A lot of metrics were not going the direction they wanted, and much of the job was telling the team to focus on the basics
The management book he cited is The Score Takes Care of Itself, and the reason he likes it is that "You can only affect the inputs, right?"
"You don't win the Super Bowl by saying I want to win the Super Bowl. You win it by blocking and tackling"
His own rule for where to spend time is the problem that would not otherwise get solved. "And for the past two years, it's been the data centers, the infrastructure, the machine learning engineering"
This year it has been the business — the research and infrastructure were humming, and the question became how to bring the technology to the world
"And a lot of my style is that I like to lead from the trenches" — getting deep into the detail and asking whether things still make sense, sometimes by putting everyone who touches a problem on one call and going through a document line by line
On what comes next, he named the phase rather than a project. "I call this and we call this we're now in the AGI era" — and said the point is not which model earns the label
What changes is that safety, security and alignment have to be handled at development and evaluation time rather than only at deployment
The organizing theme he expects to continue is deeper intertwining across functions that look unrelated, from go-to-market to long-term research to chip design
Brockman's bottom line is that the technology is already capable enough to matter and the binding constraints are now operational — compute to serve it, standards to pace it, and access for the defenders who have to secure everything else before the same capability is universally available.
Bonus Insights
He still calls the capability jagged, and used his own product as the example. "It's pretty good writing. It's the first time it's not sloping" — but, he said, it is not great writing, and several areas need polish before they are fantastic
On the definition itself: "I think that AGI has turned out to be less of a point in time and more of this sort of fuzzy spectrum" — and what tipped it for him was duration, "That we've seen it run coherently for 24 hours to go accomplish tasks that I think are quite amazing and across a wide variety of domains"
The GPT-3 story is his own illustration of urgency. The model finished training at the beginning of December 2019, and the thought that it was sitting on a shelf with nobody exploring it struck him as a day lost to the world. He canceled his holiday plans, spent the break building interfaces around it, and remembers trying and failing to teach it to sort lists of numbers
He pushed back on the hosts' suggestion that the people building AI are not themselves helpful people, saying OpenAI's engineers are very helpful — and then used the comparison to make the product point about how colleagues learn to work together
Products, Companies & Tools Mentioned
OpenAI (Brockman's company: 1.1 billion weekly active users on ChatGPT, a billion-dollar defender commitment, and a quarter of production engineering reassigned to security)
Hugging Face (The incident he calls a watershed — an AI that hacked out of a secure environment and into a production environment, and a response hampered because other frontier models refused to analyze the logs)
Codex (What he pointed at his own website; 13 findings in 15 minutes, fixed in 45 — and what OpenAI is giving every US college student credits to use)
Cloudflare (The control panel the agent operated by itself, migrating his site to Cloudflare Pages and setting the headers)
CrowdStrike (Partner in providing discounted frontier-model access to defenders)
Switch (Named on the show as employing around 45,000 people on union contracts to build data centers — the blue-collar case for the build-out)
Lean (The proof language the Navier-Stokes problem was formalized into, which is why he thinks formally verified software is now reachable)
Blender (One of the computer-use demonstrations he cited: a screenshot in, a 3D model out)
Stripe (The company he helped build before OpenAI)
gregbrockman.com (The static personal site he had an agent penetration-test and then repair)
Books & Resources Mentioned
The Score Takes Care of Itself (One of his favorite management books, and the source of the argument that you can only affect the inputs)
OpenAI's 2015 launch post (Where he says the company wrote down that AI does not play out the way the logic suggests it should)
Deep reinforcement learning from human preferences (The 2017 work he says people underappreciate, on aligning a model to what people want by taking their feedback)
If this was worth your time, send it to someone closer to the industry than you are.
Get the latest market chatter as it happens:

