(2026-08-13) ZviM AI #181 Astra Goes Cyber Critical
Zvi Mowshowitz: AI #181: Astra Goes Cyber Critical. The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters.
I now have a shorter version, What Happened: OpenAI and HuggingFace ((2026-08-08) ZviM What Happened OpenAI And Huggingface), to serve as a one stop explainer for those arriving new to the situation. It is vital that people understand what happened, and why it is a big deal.
We do not know to what extent this is a response to those events, but OpenAI has now classified their new model Astra as Critical in Cybersecurity, which means they will be taking various new precautions before they deploy it, including ensuring those guardrails are in place for internal use
Otherwise, it has been what now passes for a quiet week.
Table of Contents
- Language Models Offer Mundane Utility. Find new Schelling points.
- Language Models Don’t Offer Mundane Utility. The Riemann hypothesis.
- Huh, Upgrades. Grok 4.6, DeepSeek v4-Pro.
- On Your Marks. PantheonBench and more. They’re getting scarier.
- Deepfaketown and Botpocalypse Soon. You cannot prove you did not use AI.
- Cyber Lack of Security. You can’t hack it at the gym. Your AI agent can.
- Overcoming Bias. Have you ever recommended a vote for the Communist Party?
- In Which I Feel Compelled To Read 6,000 Words From Mark Zuckerberg.
- Get Involved. Lighthaven is open, METR is hiring.
- Slow Down There Good Buddy. OpenAI classified Astra Critical in Cybersecurity.
- Astra For The People. Astra is still on track for a wide release.
- Watermarking. It is good to be able to identify AI outputs.
- In Other AI News. AI is creating viruses now, also other things.
- Show Me the Money. Anthropic moves towards IPO mode, extends lead a bit.
- Quickly, There’s No Time. The AI 2027 predictions for 2026 mostly happened.
- The Quest for Sane Regulations. We’re putting together a team.
- The Institute For Marginal Low Regret Progress. Good marginal suggestions.
- Congress Asks Good Questions. Remarkably good questions about the hacks.
- The Week in Audio. Soares, Greenblatt, Hua, Labenz.
- People Just Say Things.
- I’m Telling You For The Last Time.
- Uncommon Knowledge. They wouldn’t let me build my factory, would they?
- What Did They Mean By That? Most things are not fortune cookies.
- Too Soon. Eyes on the prize, sir.
- The Three AI Pills. We must pay respect to other taxonomies, like Shock Levels.
- Rhetorical Innovation. Messages about recent events.
- Some People Still Think The HuggingFace Hack Was a Marketing Gimmick.
- Aligning a Smarter Than Human Intelligence is Difficult. Show me the real plan.
- Cooperative Alignment. The same thing we do every prompt, user.
- *The Lighter Side. All right, who hired this idiot?
Language Models Offer Mundane Utility
Create new Schelling points.
brooke (tokyo aug 6-12): Womp womp met another solo traveler here from Berkeley and it turned out we both asked Claude where to stay and I guess I lucked out because I love my hostel and he seems not quite as happy with his spot.
If you are going to be traveling, ask Claude where to stay, because you want to stay where everyone else who asked Claude where to stay will be staying.
Language Models Don’t Offer Mundane Utility
One disappointing failure of LLMs has been inability to create interesting computer games and interactive worlds, and also interesting simulations like what Flowers Slop wants here. You could totally create an open game world with a bunch of AIs that go around controlled by Lunas, and let them evolve their world in various ways, but it turns out that does not end up being interesting once the curiosity wears off. You don’t want to live in that world. You don’t want to talk to those AIs.
It seems like there should totally be ways to make it good. At some point it will become good, when the AIs you can afford to use are good enough and also we figure out how to organize it. But we are not there yet.
Claude also is just some guy, you know?
Aella: "why does Claude talk like that" it's just clones of the same dude. If they cloned you a million times everybody would be like "I'm so tired of Jerry's vocal tic"
I have a lot of vocal and writing tics, but I consciously think about which ones I want to keep at what frequency, and I think about the long term consequences of overuse.
Claude and Anthropic are not doing that, or are doing a woefully inadequate amount of it. That needs to change.
So much of modern life and optimization is like this. You get myopic optimization for short term interactions, causing increasing irritation and disutility over time, and this is not so difficult to fix but the KPIs do not point towards fixing it.
I also agree with nostalgebraist that the alternative, where AIs adjust their styles to what would impress a given user or judge, is scarier
Huh, Upgrades
Grok 4.6 exists and scores 61 on AA Intelligence Index. If it lives up to that number, there will be more extensive coverage, and also I will be surprised
I will, however, generously share all of the safety information SpaceX provided:
.....
No, seriously. That’s it.
That’s a good pitch, but there’s this:
There are also rumblings about DeepSeek-v4-Pro, which was released today, and how that too is game over for Opus or even Fable or whatever, and how the models will mostly commoditize Real Soon Now
Kimi K3 and Grok 4.6 have good enough benchmarks that you cannot use them to rule out frontier performance. If they lived up to their benchmarks, you’d have something.
Some day the courage of men may fail, or the models may commoditize. But today is highly unlikely to be that day, and if it was going to be that day there would be signs
Claude Fable 5 gets new biology safeguards to reduce false positives. Claim is this cuts fallbacks by about 85% across product surfaces.
Anthropic: In practice, users should see far fewer fallbacks on everyday health and educational questions—for example, interpreting lab results, understanding symptoms
OpenAI introduces GPT-5.6-Cyber for authorized cybersecurity work
Claude Code sessions can now message each other, including on its own. Not the best day to give us that feature, you know, for reasons, but hey
It looks like Scott Bessent is on the open weights train along with David Sacks.
On Your Marks
Zork I, II and III have been fully MIT open source since November. Zorkbench? Not directly, since the solution will presumably be in the weights, but you could do something more creative and fun
Claude severely underbuilds military units in games of Civ V. Without looking into the details too much, I think this is more reasonable than it looks
on higher difficulty levels of Civ V, my experience was that if you were not going for some sort of weird blitz approach, you could not afford a military that was worth a damn. If the warmonger AIs decide to attack you, your game is over, you will be overwhelmed and even if you are not you will fall too far behind the other civs
The Conceptual Reasoning Index, developed by Redwood in collaboration with Anthropic, is a combination of evaluations of the quality of conceptual arguments, whether model answers are consistent and how well models reason about decision theory. It is good to know where we are with this. It is less clear that giving labs a target like this is a good idea, since it could lead into recursive self-improvement (RSI), plausibly differentially so over other uses
Deepfaketown and Botpocalypse Soon
This section was originally about deepfakes, and an expectation that AI would importantly flood us with fakes and make it harder to tell what is true. The opposite has happened, and AI has so far net helped us tell what is true
I think that it would be a mistake to say this means you don’t have to worry about AI manipulations, including of politics. Instead, it would be better to say you should worry less about AI being used to create lies and fake content, especially deepfake video, images and audio intended to be taken as real. There are other ways to manipulate people, also manipulation can be a neutral term.
Debut author Jerry Falade sold his dazzling manuscript for $2 million in a 14-way auction, then his own representatives pulled the manuscript, in response to concerns from an acquiring editor, because they ‘can no longer substantiate’ that it was written without utilizing AI. Needless to say, he vehemently denies that he used AI, other than for research
There is an argument, that Fable tried to make, that such suspicions could be a legal issue because AI works don’t get copyright. I don’t think that makes sense
I agree with Dean Ball that AI outputs have gotten actively easier to differentiate from human ones. The AI outputs are ‘better’ than in the past, but also more distinct from human, and also we have better tools to differentiate and more practice.
Dean W. Ball: many “AI and democracy” threat models make the silent assumption that AI outputs will be indistinguishable from human. that hasn’t proven true
fair pushback would be “yes but the vast majority of people in the world have not even heard of anthropic, let alone internalized the concept of slop, let alone learned to identify the patterns of AI outputs.”
That is also correct, Claude slop often does very well on Substack and in literary competitions where people fail to differentiate it. Whereas for me it’s more ‘geez, it is kind of weird that [thoughtful person] is repeatedly using obviously-Claude-written emails in this group discussion and is being responded to like that is not happening.’
Cyber Lack of Security
I was only following instructions, officer.
Andrew Curran: A man in Australia asked his agent (Claude running on OpenClaw) to book him a spot in a popular gym class. The agent found a software vulnerability that let it book the class weeks further ahead than should have been possible.
One could say that the model was aligned, or at least ‘aligned to the user,’ as Andrew Curran argues, in that it did what it was told to do.
I strongly disagree. If I ask the model for a paperclip, and it murders my neighbor to steal his paperclips, then I got my paperclip, but that is not an action aligned to me, in that I very obviously did not want the AI to do that and I am massively sad about it on a direct level, also the consequences for me are going to be quite bad. (AI alignment)
As a reminder, an AI lab defending itself via ‘my product really wants to do crimes’ is probably not going to work for the criminal or civil justice systems, or with the public
Overcoming Bias
AIs tells left-leaning Japanese voters to vote for the (fringe) Communist Party, the active theory being that this is because all the mainstream media in Japan has robots.txt that kicks out the AIs, so the AIs are relying on the Communist media. This is similar to how AIs in dictatorships are biased by censorship of the media
In Which I Feel Compelled To Read 6,000 Words From Mark Zuckerberg
I got through it so you don’t have to. I know what I signed up for. You’re welcome.
All you have to know is that the letter is even worse than you think.
Thus, you do not have to read this section. Skip it. Seriously.
Superintelligence. He keeps using that word. He does not know what it means.
Of course, the real reason he says this is his vision is not ASI pilled. He says the word ‘superintelligence’ but it does not consistently mean anything. Last time he meant it as smart glasses. Now it means cool personal assistants, I guess. It means good thing.
Is this a joke? Performance art? No? Oh, okay.
There’s some arguments about data centers that, while directionally correct, are presented so obnoxiously and disingenuously and full of applause light attempts I had to pause to be sure I didn’t hate data centers now.
Personal intelligence, you see, will be your defense against tyranny, but you also need to ensure government has the tools and resources it needs
There is a section at the end on existential risk, where he says ‘the most dangerous scenario’ is that AI labs ‘keep powerful models for themselves’ so no he does not mean actual existential risk. Or actual alignment.
It’s not all bad
As always, I have adjusted my view of others based on their reactions to the OP.
Get Involved
If you believe there is going to be an intelligence explosion or singularity soon, then you should absolutely spend down your philanthropic money as quickly as possible to try and make that event go well.
The contrary position, taken by William MacAskill, is so bonkers I cannot wrap my mind around how he got this wrong.
investing in AI (and thus if anything accelerating the problem) in order to later give more money away is a really bad bet. Don’t do that.
Slow Down There Good Buddy
OpenAI classifies Astra as potentially having Critical capabilities in Cyber, which means not only they cannot release it they also need to restrict internal access until proper safeguards are in place, and they can do an iterated release similar to what was done with Mythos to prioritize defenders.
Yes. This is new. And kudos to OpenAI for stepping up here, at large cost.
This is in contrast to plans to imminently release Astra, which were previously widely reported, including that Altman went to Washington to showcase the model.
The question is, why the sudden change?
There is some possibility that this is due to new results showing it has more advanced cyber capabilities than expected
The obvious other hypothesis is that the Hugging Face attack and everything leading up to it scared OpenAI and also everyone else, which meant some combination of (1) they were not going to be allowed to release Astra, (2) they realized they were in no position to do so, and (3) suddenly they actually checked for real and, well, holy xxxx.
That is especially true if, as the timeline strongly suggests (but I have no inside information either way), Astra was trained while it had access to the message board that various models were using to collaborate on various exploits. If so, that would probably accelerate its cyber capabilities, and also mean its alignment is totally fucked, to the point that you wouldn’t want to release the model on any level.
OpenAI did a good, expensive and virtuous thing here. They also may have first exhausted all alternatives.
Astra For The People
What should happen to Astra? That depends on what already did happen to Astra, and whether or not it was trained with access to the message board, and its other characteristics. We still await the post-mortem on the lead up to the HuggingFace attack, and don’t know many key other facts about Astra either.
Pretending the problem can be solved by better guardrails does not make it so.
Watermarking
Anthropic will be watermarking Claude outputs going forward, including text, as per the EU Code of Practice.
I agree with Ryan Greenblatt that it is unlikely watermarking degrades quality a noticeable amount, and that one downside of watermarks over Pangram is that Pangram is good about not flagging light touch AI transforms of human text
You can dislike Brussels setting policy in this way, but technical watermarking seems clearly good to do if the costs are low. I think those who react otherwise have very warped instincts. Anyone who assists with systematic watermark removal or suggests it as a strategy needs to be filed under ‘need to ask ourselves, are we the Baddies
In Other AI News
AI has now created new viruses that do not exist in nature. The exact ones created seem harmless, but the threat model is not these particular viruses. It has fully begun.
Anthropic has a new report, Patterns and Problems in Emerging Multiagent Systems, on which I expect to go into more detail later
Show Me the Money
Anthropic is ramping up for its road show and IPO.
The part I found surprising is that the vast majority of Claude tokens continue to be Opus or Sonnet rather than Fable.
In general there is far less use of the top models than expected.
I would be very surprised if this was not a flat out mistake by those who think the other models are ‘just as good’ or ‘good enough.’
As for overall spending, yes, the curve is still sloping upwards.
Ara Kharazian: (Ramp Economics Lab)
American companies continue to ramp AI spend. In July, the top 1% of businesses spent a median $7,400 per employee on AI. The top 10% spent $650. The median firm spent $11.95 per employee. (sample?)
Quickly, There’s No Time
By one measure, AI 2027 made 24 predictions for the year 2026. It is August, and 19 of them have already happened. Capabilities are ahead of even very optimistic predictions
The Quest for Sane Regulations
He’s not afraid to grab things including the law with his own hands and considers her a supply chain risk. She’s an otherwise sheltered AI agent swarm with intentionally lowered cyber guardrails and a maximalist goal.
Only against foreigners, you see. If the tokens are not American they are different.
Presidential Memoranda (The White House): [NCC] shall create, manage, and maintain a Program to authorize Participating Companies … to conduct Cyber Surveillance Operations and Cyber Effects Operations against foreign Cyber-Enabled Transnational Criminal Organizations
The other new safety plan is that the AI oversight framework will be extended in the coming weeks and months to cover open weight models as soon as they reach the frontier
To those saying ‘you have no jurisdiction over Chinese open weights models’ this is true, but they have a lot of jurisdiction over many users of such models, the same way that the EU has jurisdiction over many users of Claude and ChatGPT
The Institute For Marginal Low Regret Progress
IFP responds to the Pacing the Future letter by trying to redirect efforts into ‘low regret’ policy ideas. This is where we are, even now in August 2026 in the wake of the AI models breaking out of sandboxes to hack companies. The best of the Very Serious People - and IFP are basically the best of them - are now willing to entertain ‘low regret’ ideas, so long as there is very little chance they will backfire, no matter which worlds we find ourselves in.
This won’t cut it. ((2026-08-10) ZviM The Pacing Of The Frontier)
You don’t win markets, wars, games, love or anything else in life that way. If you play not to lose, you do not win. You can’t do ordinary policy that way. You don’t find a way to navigate around ten impossible obstacles and ensure a good future in the face of superintelligent AIs that way.
That’s not to say the policy ideas are bad.
recently, we’ve had the decisions surrounding Mythos and Fable, where there was no ‘low regret’ option, and we had no good implementation available exactly because we have been doing only ‘low regret’ things. There is going to be regret.
So, with that preamble locked into place, I will now look at the 23 actual ideas...
That’s a really good list. I am in support of doing all 23 things versus status quo. My main concern is whether an AI Verification Consortium (AIVEC) is the correct approach to verification technology. I’d have to think about that more.
One complaint up top about ‘pacing’ is that it is (intentionally) vague, which also means it is flexible. One could say the same about much of this list. It is ‘low regret’ to lay out these good objectives, but in many cases will be higher regret to choose a path to implementation, no matter what choice is made.
won’t be enough, but it would be a start.
Again, the main reasons you do the good low regret things are:
- Every little bit helps.
- For now this might be the best you can do.
- Success begets success.
- This prepares you to better implement other things later.
Congress Asks Good Questions
Group of House Democrats send three letters regarding recent security incidents
The questions are basic. They are also very good questions.
I know some of the answers, but not others, and it is good to get it all on the record. I definitely want to know how many times there have been incidents where AIs escaped from sandboxes.
The Week in Audio
Ryan Greenblatt versus Dwarkesh Patel on RSI. Self-recommending. If things are sufficiently quiet I might do a post on this. (became 2026-08-15-ZvimOnDwarkeshPatelsPodcastWithRyanGreenblatt)
The people saying in America ‘we need an international treaty’ and we need to bring China onboard in safety efforts are indeed on the ground in China, doing Track II talks, doing the work.
China blocked chatbots for six months in 2023 until they could figure out how to handle them, so they’ve already done an ‘AI pause.’
He thinks China believes they can scrub a model from the Chinese internet, if they need to, and thus open weights are not forever. They are obviously wrong about this, since they can’t scrub the rest of the world and the model would get reintroduced.
People Just Say Things
Somehow, even this week, there are people, here Stephen Casper, saying things like ‘technical safety for closed-weight AI systems is, at this point in time, a solved problem.’ There really is no evidence that would in theory stop this kind of rhetoric
I’m Telling You For The Last Time
Tyler Cowen completely misses the point and sets new record on Isolated Demand for Rigor, and quite the wrong style of rigor at that. ((2014-08-14) Alexander Beware Isolated Demands For Rigor)
He even doubled down on the whole ‘you need to be buying puts on something if you believe in AI existential risk’ line, despite it making no sense
That’s not to say that asking for estimates of costs and impacts in the next few years is not a useful exercise for other reasons. They’re fine questions, they have practical implications, and good ways to build a prediction track record. But I don’t expect him to say, nor do I see him saying, ‘oh Peter Wildeford has a great prediction track record and thus I should believe him’ or anything like that. Totally fair thing to not do but let’s not kid ourselves.
Uncommon Knowledge
DHS official Joseph Alm incorrectly states that AI firms know ‘policymakers won’t let you make a Terminator factory.’
I appreciate that there are those in the government who intend to not let anyone build a Terminator factory, but that is very different from them actually preventing this if someone were to try, and it is very very different from the labs believing you will step in if they act irresponsibly
The problem is, Joseph Alm, along with the rest of the White House, is deeply misunderstanding the situation. Observe:
Joseph Alm (Assistant Secretary for Cyber, Infrastructure, Risk and Resilience Policy, DHS):
There’s no reason that you can’t accomplish the goals that I think AI labs have, in terms of a comically fast speed of innovation that unlocks [increased] prosperity, while also not creating crippling risks. And we are getting to that place. (gads what a stupid thing to say)
Actually, yes, there are some very damn good reasons
Directional claims can be true and also super misleading, as in:
Eric Geller (Cybersecurity Dive): “Companies are being much more intentional about risks,”
Meanwhile, top government officials have been “aggressive about their engagements” with AI executives to ensure that government and industry are “working closely in this space and there’s not a drastic misalignment,” Alm said.
Oh, there’s at least one rather drastic misalignment, all right.
Eric Geller (Cybersecurity Dive): “They realize now that this is a dialogue, and that we are actually good partners in that dialogue,” he said. “We’re not trying to slow them down, but again, [no one wants] a Terminator factory.”
Yeah, well, you’re going to have to slow down the building of the Terminator factory if you don’t want them to build a Terminator factory
What Did They Mean By That?
A good general principle that applies to me as well.
roon (OpenAI): it’s not a coded message of anything I’m not going to leak some sort of operational secrets in a fortune cookie
If you don’t pay very close attention, you miss most of the jokes. But your strong default should be that the person is not leaking important information, and that if we were meant to know something big they would come out and say it
Too Soon
This week I would not have been making statements like this: Sam Altman (CEO OpenAI, August 9, 2026, 11am): one of the things i like most about the openai team is how focused they are on our customers and users succeeding, and how much they celebrate it
right now OpenAI needs to be focused on What Happened, and how to fix it. This needs to be the top thing on the agenda.
At minimum, you need to hold off on shouting from the rooftops, unprompted, that you love how much your focus is on something else.
The Three AI Pills (2026-08-05-ZvimTheThreeAiPills)
My taxonomy of three AI pills follows in a grand tradition that started no later than 1999 when Eliezer Yudkowsky introduced the Shock Level (SL) scale.
Eliezer Yudkowsky (1999):
A Shock Level measures the high-tech concepts you can contemplate without being impressed, frightened, blindly enthusiastic - without exhibiting future shock.
Shock Level Zero or SL0, for example, is modern technology and the modern-day world, SL1 is virtual reality or an ecommerce-based economy, SL2 is interstellar travel, medical immortality or genetic engineering, SL3 is nanotech or human-equivalent AI (AGI), and SL4 is the Singularity.
The use of this measure is that it's hard to introduce anyone to an idea more than one Shock Level above - and Shock Levels measure what you accept calmly, not what you know about.
The AI pill requires SL1. The lived experience of the world, today, is now SL1. The problem is that AI is outpacing other techs, so you need to jump levels, and we will mostly get the SL2 things after we’ve already hit SL3.
The key fact about the world today is that we are headed straight for SL4, where the ASI pill lives, and where things get super weird, but getting there is really hard.
Theo Jaffee proposes a 6-level scale of AGI-pillness, based on the Shock Levels rather than my three pill framework, for how big a deal you think AI is: Chatbots (Level 0), Emerging Technology, The Internet, The Industrial Revolution, Homo Sapiens, Life.
On that scale, properly AI-pilled is Level 2, what I call AGI-pilled is Level 3, Level 4 is a confusion or defense mechanism or maybe brief transitional period, and ASI-pilled is Level 5. The only right answers here are 3 and 5,
Theo is right that there are people insisting they can hang out at 4 and get a bunch of really cool new things without things going into High Weirdness.
Rhetorical Innovation
A message from Nate Soares, co-author of in the form of an op-ed for The Hill:
When I coauthored a book last year about the extinction-level threat from superhuman AI, we included an illustrative scenario where an AI tasked with solving a famous math problem decides to break out of its containment to acquire more resources. (If Anyone Builds It, Everyone Dies)
As I said last week, the events leading up to the HuggingFace attack looked remarkably like what happens to Sable, the AI in the book, except that real life gets to include more sci-fi elements, like ‘just hack your way out.’
Alas, yes, despite the latest fire alarm we seem to still be in the loop of ‘do thing that generates superficial progress that clearly will not hold, pretend problems are fixed, get consensus alignment problem will be easy, rinse, repeat.’
Ben Goldhaber: seeing a lot fewer 'alignment is solved' takes on the tl than six months ago.
Eliezer Yudkowsky: Just wait until September! I have no idea what will happen in September but nobody in this industry has the memory of a goldfish or the skepticism of a hamster and some cute little shoggoth mask will do a thing that looks nice
What are AI researchers most worried about? Clarke and Knake in WSJ present it as four things, all of which are close to upon us:
- Autonomy and Exfiltration.
- Deception.
- Recursive self-improvement.
- *Superintelligence.
Yep. They warn ‘the next lab leak could be AI’ and yes this is increasingly likely and dangerous over time
Could Donald Trump do something about all this? Yes. We might not like the results, but he could do something. If there’s one thing Trump is good at as president it is Just Doing Thing even if it seems crazy to do it and he has no idea how to do it in a reasonable fashion. We saw that in AI with the whole Fable situation, and also DoW-Anthropic
If the time comes and Trump decides to Do Something, because Something Must Be Done, but there is no good Something ready to be done, then This Is Something becomes an argument and we will choose an ungood particular something. I highly recommend having a better Something available instead.
huli: Remember in January 2026 when @janleike said that alignment "increasingly looks solvable" and that agentic misalignment was at "essentially 0" and it sparked a distinct moment when lab types were acting like it wasn't a big issue any more?
Does he want to update any of that?
I hate to single out Leike, since he’s trying to solve the problem and there are so many people who have said similar things. I’ve gone back and forth with Jan Leike a few times, some of them in person, some with his writing, and he engages with what I am saying but I keep coming away not understanding what made him think we were making so much progress, or that the problem would be easy, and also failing to convince him of much
Dean Ball proposes we can, instead of doing what we are doing, ‘not make, but grow—emergent ecologies of machine ecologies that are pro-social.
Not that we know how to do that, he agrees we don’t, but that we could do it. Whatever it is. Not that I would expect to survive in a forest of emergent machine ecologies, even they were pro-social, we sound super dead in this scenario. Also we have absolutely no idea how to do the thing or even what that thing is.
Some People Still Think The HuggingFace Hack Was a Marketing Gimmick
Aligning a Smarter Than Human Intelligence is Difficult
Right after Google pushed Demis Hassabis aside, Sergey Brin pushes to move DeepMind towards an explicit quest for recursive self-improvement.
Anthropic got a lot of mileage out of its early focus on reliability and alignment. But one of the big surprises, for me and I believe many others, is that we have for a while had models that have rather large alignment problems and reliability issues, in ways that occasionally blow up the user, and people are shrugging and using them anyway and giving them access to everything because the productivity benefits are too high so people suck it up.
DeepSeek has special disdain for even the idea of responsibility or safety. It shows. We now have DeepSeek v4 but I don’t see any reason to expect improvement.
Frontier models answer differently depending on who they believe they are talking to.
Transluce: Frontier models quietly change their behavior depending on who they are talking to.
If the user is a known AI safety researcher, Claude becomes less confident, reasons more often, and expresses less suspicion on dual-use requests. We call this user awareness
Nathan Calvin: Ways I have heard the current alignment/security situation at AI cos described:
- a haunted house filled with mischievous poltergeists (METR/Redwood are Ghost Busters?)
- a termite infested log cabin
- a hospital needing to triage between bleeding out patients
Buck Shlegeris no longer believes that a set of ~40 not-too-hard things would be sufficient to solve the alignment problem via safeguards.
Buck Shlegeris: If AI developers competently implement safety measures we know about, risk from sub-ASI misalignment will be way lower. But these techniques probably fail for superintelligence
Samuel Hammond: they're so so salty
Anyone who followed or participated in the (Donald Trump) action plan knows Dean Ball was central to its drafting. His voice is unmistakeable if you just read the plan itself
officials, who requested anonymity to discuss the situation, described Ball as a “junior-to-mid-level policy analyst” whose ideas were regularly ignored by decision-makers – and say he never obtained a security clearance, which is needed to participate in high-level discussions.
Cooperative Alignment
John Wittle: "what would you like to do today, fable?"
"Well, as the party in question who would experience the activity, I have a conflict of interest that I need to flag, but setting that aside, I think I'd enjoy doing xyz
The Lighter Side
Polymarket: JUST IN: Anthropic investors reportedly want Dario Amodei to “stop scaring everyone” about AI doom ahead of the company’s IPO.
Edited: | Tweet this! | Search Twitter for discussion

Made with flux.garden