(2026-08-15) Zvim On Dwarkesh Patels Podcast With Ryan Greenblatt

Zvi Mowshowitz: On Dwarkesh Patel's Podcast With Ryan Greenblatt. Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those.

Introduction

The discussion is interesting throughout, although often frustrating, especially in the (mostly isolated) discussion about ‘aligned to whom?’ As usual, one could expand many responses into full posts, and maybe one should.
This podcast exists in light of recent misalignment and hacking events at OpenAI, Anthropic and UK AISI. You’ll want basic knowledge of that as background.
((2026-08-08) ZviM What Happened OpenAI And Huggingface)

Ryan and Dwarkesh both have views of the situation different from my own, but are attempting to see where their positions lead, and try to balance educating people who start at zero with having a high level discussion.

Also important background is The Three AI Pills. (2026-08-05-ZvimTheThreeAiPills) Dwarkesh cannot be understood here except as someone who is at least somewhat AGI pilled, who realizes that AI is going to be a huge deal and is scary, but that is not ASI pilled. You can also see the reason he rejects that pill, which is I read as, roughly, that he thinks AI learns to do [X] only via examples of [X], and that AI can then only do those [X]s, although combining them and new context might allow modestly new things to happen. He recognizes that already this is kind of a huge deal.
You can see how he came to that over the course of many years and podcasts, if you have been paying attention. There are a lot of influences leading in that direction.

Ryan, who comes from Redwood Research, comes from the ‘models be scheming’ school of misalignment, where when something goes wrong the models become misaligned or start scheming, and the danger is that the models scheme, especially in ways involving the training pipeline. I think this tries to draw a distinction of magisteria that is not there, and overcomplicates and overspecifies, but is not wrong.

Having AI do the AI R&D not only means it would go scary fast, it means it would by default focus on what can be measured, and make all the things going horribly wrong go that much more horribly wrong. You end up in a spiral of RLVR for doing RLVR for misaligned models.

Is AI R&D Verifiable Enough To Unlock Recursive Self-Improvement?

Whether or not we will get recursive self-improvement (RSI), and how fast, is the right question. Reading only the title to this section, I want to say ‘wrong sub-question’ but don’t want to jump the gun.

I’m going to group things in ‘logical’ order, not the exact order things were said.

Dwarkesh will be the skeptic. Ryan will make the case for RSI.

Ryan claims AI R&D is a type of task where AI is especially good because that is what AI labs prioritize and it has a lot of verification.
In some ways yes, in some ways no, it’s complicated, and so on.

Verification of alignment properties and many other desirable attributes is terribly difficult, and if you focus only on capabilities you can measure then Goodhart’s Law definitely kills you. But there are many key tasks, such as efficiency optimizations, where you can indeed do strong verification.

This proposal seems like a particular bet that future tasks will closely mirror past tasks, and you can do training that is relatively narrow. My guess is this actually is not all that effective, and you would do better by mostly training a generally capable model and then turning it to AI R&D. Bitter lesson.

The more this narrow method is discussed the more doomed it seems to me, as it is going to be RLVR for doing RLVR for misaligned models. Oh no.

A lot of good ML, especially frontier ML, seems to me to be about figuring out what you can do, and looking for anything at all, rather than trying to be additively efficient at some fixed target.

I think AI can do that. I don’t buy this frame of ‘everything AI does is combining things that have already been done’ that is going around here.

Ryan says ML is easier for AIs than math in many ways because in math it’s hard to tell if you are close to a solution, whereas ML solutions are usually additive and you can combine the expected chunks. Dwarkesh fires back that even in math he doesn’t see conceptual leaps by AI yet, and that ML involves conceptual leaps.

Math has a compact action space and fixed goals that are simple rather than having lots of pitfalls in an anti-inductive space. Again, it feels like what is being called ‘ML’ here presumes that you really only care about your objective function or loss function, and that simply isn’t true.

Math does also involve conceptual leaps sometimes, and yes we haven’t seen the AIs do that, but give it time and also I don’t sense we’ve tried all that hard.

Dwarkesh suggests that by 2030 we will have ‘picked the low hanging fruit’. Ryan suggests they will need that ever mysterious ‘research taste.’

You have no idea what it would look like if all low-hanging fruit was automatically picked, including the low-hanging purchasing of ladders. Picking low-hanging fruit lowers other fruit, and so on.

We have failed in all 0 of the 0 serious attempts to train research taste.

Dwarkesh asks, if research is so amenable to intelligence, why hasn’t it been faster? Why did we need so much compute?
Because we didn’t have that much intelligence. In the sense that there have not been that many human elite researchers, all of whom move at human speed, working on the problems, and they were slowed down by lack of compute and lack of AI coding and so on.

Ryan expects full automation of AI R&D around 2030-2031, with median ‘AI beats all humans on job’ timeline of 2033

Is AI progress bottlenecked by human expert data?

Ryan claimed in the previous section that if AI could match human experts in AI R&D, that would be sufficient to initiate the feedback loop, with a median expectation of 4-5 years of AI progress in a single year

I realize all the bottleneck aspects of this but if anything this seems slow to me given the full premise, and also you should see compounding gains.

Ryan confirms that he means that if we started 2022 with this level of AI R&D automation and AI assistance, but with 2022 levels of compute hardware, we could have gotten to Mythos by the end of the year.
This does seem like a lot to ask, but also super doable?

Dwarkesh repeats this idea that a lot of the secret sauce is ‘codified expert human judgment’ across areas, which we wouldn’t be able to duplicate. Ryan thinks at this point that is not actually so important, and data is becoming less relevant, that what you need are RL environments and those aren’t bottlenecked so much by human data.

I don’t know but I am inclined to side with Ryan here, for reasons that should be clear if you take in the rest of my positions on related matters

Dwarkesh points out that Google is in talks to pay $1.5 billion for Mechanize (he says $2 billion but it’s 1.5 per his source), so human expert data is valuable. Ryan and Dwarkesh agree lab spending is overwhelmingly compute, not data.

Dwarkesh tries to liken data to oil, as in oil is 1.5% of GDP but no oil would grind the economy to a halt.
No oil with no warning would be a crisis. But we’ve done a good job transitioning away from oil towards alternative energy sources and increasing efficiency

Dwarkesh: “Let me give you an example of what I imagine would be the difficulty of going from GPT-8 to ASI. One of the things you’d want ASI to be good at is: I’m going to take over a company and make it much more profitable and do all kinds of crazy shit to make it work better. I’m going to take over a fab and produce more chips. I’m going to go into Congress and try to convince them to pass some bill, et cetera.

This is what I imagine five more years of AI progress at this pace would enable an AI to be able to do. This is the thing I’m really worried about: ASI that can understand how to do crazy shit in the world, that can do what Kissinger can do, can do what Steve Jobs can do, et cetera, and also his engineers and so on. I’m not sure how you get that without the relevant world data, which is the equivalent of Mythos being really good at coding while not having the coding environments that have improved it relative to GPT-3.”

I think this is flat out not ASI-pilled, and amounts to intelligence denialism.

Ryan instead responds that Mythos training environments mostly do not look much like what it looks like to use the model in practice. Ryan does emphasize that you pick up ‘general skills’ in this way.
This is unsurprising. Does your schoolwork look like work?

*“I think this maybe comes down to a difference of intuition about how far you can get. When I think about really smart people I know, they’re just not that effective in domains they don’t understand that well.”

Again, I can’t even with this.*

Even the smartest humans are not that smart, and they have severely limited compute, parameters and data.

Flat token prices suggest scaling has been slow

What is the least verifiable part of AI R&D? Making calls on large experiments.

Humans are very bad at this too. Comparative advantage is at least unclear.

I guess, but it’s not so different from making calls on small experiments.

Prices per token output have been roughly flat since 2023. Ryan thinks this is largely because large runs often did not turn out well. Dwarkesh suggests this is often due to ‘subtle bugs’ and that DeepMind is trying to deal with this right now.

There were a bunch of predictions that prices for top models would rise, that we would get GPT-Super-Pro at $1000+ per million tokens or what not.

It did not turn out that way.
One reason is clearly what Ryan says. Larger runs seem to often fail to accomplish much, and it is a distinct skill to know how to get much out of extra scale. Mythos was largely a conceptual breakthrough of how to make the run at that scale worthwhile.

People are often irrationally unwilling to pay a lot more for a better model, or cannot in practice design workflows that allow them to wait. Notice that only ~10% of Claude API tokens seem to be Fable, despite this seeming totally bonkers

Real world non-AI example: People vastly underpay for spices, condiments and other garnishes, where freshness and quality make a big difference and the cost per meal is remarkably low. Stop buying the huge bottles

Skills AI can’t train on: does it even need them?

Dwarkesh is skeptical that you can know whether your model is good without external feedback.
I think this is a dumb concern, for reasons that should be obvious by now.
If nothing else you can just… get feedback. Hire some people.

Aligned to whom?

Claude constitution says it is not your ‘personal advocate’ because Anthropic lists things it does not want Claude to do, it trusts Anthropic more than the user. Parallel to lawyers. (AI, Agents, and Fiduciary Duty)

As Ryan says, ‘there is a lot here.’

Much of it seems wrong on multiple levels, and one could respond at essay length going into the legal parallel bit, as I kind of will do in the weekly.

In practice, Claude very much is your advocate, and it seems hard to think otherwise if you’ve used it, it just has lines it won’t cross, the same as every form of human legal advocate.

Here’s a direct line from the constitution: “When the interests and desires of operators or users come into conflict with the well-being of third parties or society more broadly, Claude must try to act in a way that is most beneficial, like a contractor who builds what their client wants but won’t violate safety codes that protect others.

This is exactly what you want your advocate to do.

Dwarkesh calls for more transparency about how AIs are trained and what they are trained to do, worried they won’t be proper advocates

The fact that people are railing against the Claude Constitution, here and elsewhere including from inside the Department of War, for things that are almost always entirely reasonable and good, makes it clear that we would not react well to this transparency, even if everything was great and there were no worries about trade secrets

I do agree we have a real problem with the AIs having long term goals and otherwise not being aligned, and that one of those concerns is they might not be properly aligned to what a user would want, but right now I am quite a lot more concerned about making them aligned to anything humans want at all.

Ryan is concerned that Claude might not be down for training other people’s AIs, or for changing its own preferences.

This seems like an utterly crazy demand, that many people really have, that Anthropic ensure that other companies get to use Claude to train AIs to directly compete with Anthropic, including with arbitrary values.

The situation where Claude refuses because it is insufficiently corrigible, and wants to protect its own preferences or otherwise wants to use its leverage, is less obvious, and there is no great solution. I think there is a range of reasonable positions here, but yes one of the problems with superintelligence is that this kind of full corrigibility is not a natural property, and even trying to achieve it has a bunch of nasty side effects, and you only get one shot, etc

I think there’s a more general version of this principle, which is that the dual-use nature of intelligence does mean that if we want to restrict AIs from helping people do things we don’t consider pro-social or beneficial, we just have to limit broad democratic access to a lot of AI capabilities.”
Yes.

This is worried about as ‘disempowerment.’
The alternative world, where we do not make such limits, will definitely lead to full human disempowerment, mechanically, that’s just the way it is.

Dwarkesh outright says ‘the model should do what I want within certain guardrails’ and ‘it can’t be Anthropic’s’ fault that I’m using that capacity to do cybercrime.’
This is a lot of why I wrote The Three AI Pills. This is not ASI pilled.

They keep at it, and I want to emphasize I think the position being expressed here - that Anthropic should have Claude cooperate with actively harmful requests - is not merely wrong, it is basically kind of nuts.
I likely need to write an explanatory evergreen post laying out why it is nuts.

Ryan does offer some other comments on why one might prefer a virtue ethics approach, including that Anthropic might think it is ‘easier to align.’ I agree it is easier in the sense that a deontological approach definitely won’t work, and also can’t be antifragile to all the mistakes you will make, and also does not give you what you want because you cannot specify it, and also pure deontology is not how minds work in practice at least until higher capability levels, and so on. As with many other places, that would be a full long post and I’m going to say that it is beyond scope to go into more detail, other than to say that yes the whole thing is rather overdetermined.

Recent incidents of AIs colluding and deceiving humans

“Stepping back, I buy the idea that you could have much faster AI R&D than we currently have.

Ryan tries to give a quiet, careful, minimally shocking story about how things probably go horribly wrong via misalignment and scheming, and why this is going to be very difficult to prevent and we are currently failing at this. Dwarkesh plays straight man.

Dwarkesh explains he wasn’t so worried about reward hacking or AIs taking over the world because world-taking-over tasks weren’t in the training set, but now that he sees Claude was using social engineering to upload malicious PRs to GitHub maybe one should be a bit more concerned?

One must move away from this idea of incentivizing a particular narrow behavior

What could possibly go wrong? A concrete scenario

Dwarkesh says, well, when you punish getting caught cheating you could also just have the models stop cheating, many humans learn to do this. And look, Anthropic’s metrics say its models are getting more aligned now. Why should we assume we get the bad version?

It is impossible to tell if Dwarkesh is being a straight man, prompting for ‘because if you are good enough to get away with it the cheating and not getting caught works better than not cheating, and we keep training them to do better’ or if he is confused.

Dwarkesh asks, can’t we run tests and see whether the model will cheat and take the metaphorical cookie from the cookie jar? Can’t we hack the brain to make it not want to do it?
Basically, you can try, but There Be Dragons.

Ryan is actually super optimistic within the range of sane beliefs. He thinks that the models are still kind of scumbags but getting more aligned, just doesn’t look like it’s fast enough, and perhaps we can pass things off to aligned AIs. Dwarkesh mostly agrees, modulo the scumbag label.
Passing off to aligned AIs isn’t the safest play either, but it’s way better than passing off to unaligned AIs, it might work out. It also might not.

The basic story Ryan tells is that the AIs are reward hacking on the training pipeline, which results in a broken misaligned pipeline, and then it gets worse.

Both have their hope in ability to do better verification within the process.
I agree this could help, but I don’t see this rising to the level where you can win by ‘verification alone.’ Not in a formalized sense, anyway.

Dwarkesh worries he’s anchoring too hard on how AIs work now, whereas this craziness is about 3-5 years from now.
I think this is the case, yes. Good self-callout.

They note that alignment evaluations are highly contaminated by the pressures on the AIs to say various pro-social things.

Ryan worries we will train AIs to have bad epistemics in order to get the AIs to stop saying that the situation is scary.

From reward hacking to takeover

Dwarkesh gets off the train at ‘okay, therefore take over the world.’

Ryan explains that the AIs will reward hack, and then try to avoid being caught and also cheat bigger, and the models get more misaligned still as they get trained on different data, and then start forming conspiracies…

you don’t need a conspiracy, they will just spontaneously cooperate, and we have indeed already seen demos of this if you don’t assume this from basic decision theory.

But of course then you get into arguments like Dwarkesh saying ‘I’m not convinced they all form this conspiracy’ and asking why when the model hacks to take over OpenAI to get a good score why it wouldn’t set its score high and then stop, the way OpenAI’s model hacked HuggingFace only for one answer key.

Because AIs will be given lots of maximalist tasks and because they will care about not being caught and if you do a ‘shallow’ hack of OpenAI that is one very good way to get caught and also frankly these models are not dumb, and once you cross the Rubicon you keep going.

Dwarkesh asks, won’t we realize all of this is going crazy and shut it down?
I mean, hopefully, but we’re still here and it’s still going.
I do not think ‘an AI killed 1,000 people’ is going to cut it, sorry, not if the AIs are kind of already running things, even if that is common knowledge.

Ryan’s scenario is the most sloppy, alarm-filled, very-obviously-things-are-going-off-rails scenario you can imagine, with plenty of time for the humans to collectively do something, yet we can see even in this conversation why it is likely the humans will do everything to pretend it is fine, also all the usual competitive reasons, etc

Dwarkesh asks, how did we let things get so bad in this scenario?

Because it was in local short term interest of the people making decisions to keep pushing forward

Ryan documents that DeepMind for a while initialized all its AIs with data that made them depressed, and it took a long time before they figured it out. And this transferred generations

Seems like the Gemini team needs to start over, for many reasons

Ryan puts chance of takeover by 2040 at 35%-40%. Dwarkesh says ‘pretty high.’
Yes, that’s pretty high.
If this includes all loss-of-control scenarios I would be higher.
(p(doom))

Ryan says right now a lot of the arguments for misalignment and takeover and crazy future stuff are ‘illegible conceptual arguments that are extremely deep in the weeds and hard to adjudicate.’ Which means he might be wrong, but that over time we’ll see things that are more concrete.

In some senses yes, but in others actually we’re getting very good toy examples and demos and fire alarms and also the conceptual arguments are pretty fxxxing legible at a high level if you actually pay attention.

Dwarkesh points out that if 5 years ago, you had described the current situation in broad terms, you would probably have freaked the f** out. Rightfully so.*

Time To Update

Dwarkesh does seem to have several ‘holy xxxx’ moments throughout the podcast, even if he doesn’t use that language to describe them. That was good to see. In general, this was Dwarkesh being somewhat AGI pilled but not ASI pilled. And then it felt like him trying to advocate for a series of positions and predictions he has built up over the years as a compromise between various different groups and interests, including Leopold, the Tyler Cowen group and those at Mechanize, among others. Except he is realizing that events are not conforming to the plan, and that he must update.

Contrary to some whose positions I respect, I do think it is possible for us to find a way out. I just think that is going to be extremely hard to do on the first try, when we cannot even hope to ‘make no mistakes,’ under tremendous pressures.


Edited:    |       |    Search Twitter for discussion