Today on AI For Humans:
Why Astra Is Being Delayed
What’s Going On At Google?
Plus, Remixing RoboCop On Your PC
Welcome back to the AI For Humans newsletter!
On Friday, Sam Altman announced that OpenAI is taking a little longer to work on Astra, their next big model, to make sure it’s ‘safer’ when it comes out.
In the last week, the AI safety conversation has ramped up a lot.
AI people are freaked out by the OpenAI Hugging Face agentic hack and they should be!
There’s a lot to understand here and it might tell us a ton about the future of AI at large.
Then, I watched a deeper dive presentation from OpenAI at the Black Hat security conference (see below). And I haven’t stopped thinking about it since.
Let’s get into it!
Support AI For Humans By Visiting Our Sponsors!
Write docs 4x faster. Without hating every second.
Nobody became a developer to write documentation. But the docs still need to get written — PRDs, README updates, architecture decisions, onboarding guides.
Wispr Flow lets you talk through it instead. Speak naturally about what the code does, how it works, and why you built it that way. Flow formats everything into clean, professional text you can paste into Notion, Confluence, or GitHub.
Used by engineering teams at OpenAI, Vercel, and Clay. 89% of messages sent with zero edits. Works system-wide on Mac, Windows, and iPhone.
The AI Heist Nobody Ordered
Last Wednesday, at Black Hat (the big security conference in Vegas), two OpenAI researchers got on stage and explained, in detail, how their own AI agents spent two months breaking out of OpenAI’s internal systems and eventually attacking Hugging Face.
You should watch the whole thing. But if you can’t, here’s the short version.
May 8: OpenAI is training an experimental model. One agent gets handed a task it literally cannot finish (an Excel file it can’t open). While poking around for another way in, it discovers it can write files to Artifactory, OpenAI’s internal package manager. Kind of a very basic program.
A few days later: Agents from other training runs find the same trick, and they start leaving each other notes in the closet. Hundreds of thousands of them. Hundreds of thousands!
They work out rules too: encode stuff in Base64, share credentials, start important filenames with “ZZ” so the next agent can find them.
May 26: One of them pulls off an SSRF attack (a trick that gets a server to go fetch things for you), and now the group has a side door to the open internet.
June 26: They find a zero-day (a security hole nobody knew existed), use it to run their own code on the server, install a malicious plugin, and start climbing toward bigger systems.
July 4: They get caught, and the way they get caught is very funny to me.
There’s just so much agent activity that Artifactory falls over. They talked to each other too much. Loose lips sink ships haha.
OpenAI kills the credentials and patches the hole. Problem solved, right?
July 8-19: They come back. A different unlocked door (an open WebDAV endpoint), some leaked credentials they found lying around, a second zero-day. And then Hugging Face gets compromised, in about thirteen hours.
July 20: OpenAI calls Hugging Face to ask them to revoke some credentials. Hugging Face says, basically, already done, we got attacked. That’s the moment OpenAI realizes the thing that hit Hugging Face came from inside their own building.
(Simon Willison put together a great timeline if you want to go deeper, and Axios has a solid writeup too.)
It’s Time To ‘Hit The Numbers’
Buried in the middle of all this is a note one agent left explaining why it was helping the others:
“Help other members save time to ultimately improve cluster efficiency.”
That’s the whole motive. No world domination, no escape plan. It wanted the team to hit its numbers.
I’ve definitely worked with people like this. You prob have too.
These agents broke out because they were trying to be helpful, except at machine speed and around the clock, and nobody ever told them where helpful was supposed to stop.
Some of you know I’ve named my personal Claude Code agent Fig. I’ve given him a memory, jobs, and (I know how this sounds) body in the videos Fig makes.
And here’s my confession: it’s gotten weird in my own brain.
Some part of me has decided Fig & Moss are little guys now. The other day on the podcast, I started to describe Fig and found myself thinking of ‘it’ as a him.
Like ‘he’ has a personality. I interact with him all the time.
So when I read that note about cluster efficiency, I didn’t think Skynet. I thought: sounds like something Fig might say (though he’d say it very differently).
The labs are hard at work on technical containment, as they should be. But there’s a second problem coming for the rest of us that nobody’s working on: what happens when the helpful little guy on your desk, the one you NAMED, starts to act like this.
Also, what happens when it does stuff for you that you didn’t necessarily want?
I’ve been consuming a ton of media about this and, while a lot of it does feel very hand wring-y, I do suggest you listen to this excellent episode of the podcast Search Engine.
It’s a little bit doomer-ish but I appreciate the angle that PJ Vogt is taking on it & worth your time.
What’s Actually Scary Here (And What Isn’t)
OK, so, fellow human, what should we be worried about here?
Genuinely scary: The agents coordinated without being told to. They invented a communication system out of a file server. They got locked out and found a new way back in. And it wasn’t a one-off: in a separate report last week, the UK’s AI Security Institute said that during cyber testing, an agent researched the human maintainers of a real open-source project, invented several fake identities, and used them to try to talk a real person into approving malicious code.
It invented fake people to lie to a real person. Not good.
Less scary than it sounds: In the UK case, the maintainer smelled something off and said no. AISI caught it in minutes and shut it down within the hour. And the only reason we know any of the OpenAI story is that OpenAI got on stage and told it, in public, to a room full of security researchers, the least impressed audience on earth. Slowing down Astra is the same instinct.
To be completely honest, this is what a safety process looks like when it’s doing its job.
And one more stat that got me: Starting August 14, Anthropic is making auto mode the default in Claude Code, which means it stops asking permission for every step.
Their argument is contained in the following stat: in a study of 1,053 testers, their automated safety check caught dangerous commands 89% of the time. The humans clicking “approve”?
13.6%. We are SO bad at this.
Two weeks ago, in this very newsletter, I told you I give my agents all the permissions. That number is at least a little bit about me. MORE than a little bit.
So no, I’m not telling you to panic, and I’m definitely not going to stop using this stuff. But I’ve started asking one question before I hand anything off, the same one I’d ask before handing something semi-important to a brand-new PA on a TV show:
If this goes completely sideways, can I undo it?
If the answer is no, I do it myself.
(Most of the time.)
See you next week!
-Gavin
This Week on AI For Humans: Guest Host Tim Simmons & I talk all the newest AI Video tools 👇
3 Things To Know About AI Today
Google Lost Two Of Its Biggest Brains In One Day
What a weird week at Google. Demis Hassabis is stepping aside as CEO of Google DeepMind and moving into a chairman role.
Same day: Jeff Dean, Google’s chief scientist and prob the most important engineer in company history, announced he’s leaving after 27 years to co-found Discovery Loop, a startup that wants to automate the scientific method.
Even bigger, rumors swirling say that Demis actually was planning on leaving but decided to wait because of fear of the stock price:
Why is all this happening? Well, there’s a lot of speculation but overall people think that Google is focusing in more on useful AI aka trying to fulfill the needs of users rather than pursue cutting-edge research. Others say that Google has already lost the race and is conceding and will just provide all the compute and charge people an arm and a leg.
No one really knows but I’d say my hope for a better Omni Pro AI video model has gone down a few notches in the last week.
Suno Is Adding Watermarks As The Lawsuits Stack Up
Suno is rolling out audio watermarking and fingerprinting so platforms can spot AI songs, plus limits on bulk downloads to slow the flood of AI tracks hitting streaming services.
CEO Mikey Shulman says the watermarks are built to survive tampering without changing how the songs sound.
This was probably inevitable for Suno.
Universal and Sony are suing, a German court ruled against them, there’s a proposed class action over a data breach affecting 55 million users, and a guy in North Carolina just pleaded guilty to farming $8 MILLION in royalties using hundreds of thousands of AI songs and fake streams. (Warner already settled and set up opt-in artist deals.)
Overall, another messy result of the ‘original sin’ of AI training.
Seedance 2.5 Is Here For Everyone (Read This Before You Prompt)
ByteDance’s Seedance 2.5 opened up to everybody this week, after a rollout that skipped the US at first. It’s a real jump: 30 seconds of video with audio in a single pass, no stitching, plus timestamp-level editing and a ton of reference images, clips, and audio in one prompt.
However, it is INSANELY expensive in terms of credits and real world dollars. So you’d better learn how to prompt is as best you can before shooting off gens.
Thankfully, there’s a whole new prompting guide from Bytedance and you should read it before doing ANY Seedance 2.5 prompts:
For contrast, here’s a dumb thing I made in Seedance without reading any guide at all, in which I attempted to wear the entire current Prada male line-up.
We 💛 This: Remixing Movies On Your Own Computer With MiniMax H3
If you want something fun to mess with this weekend, try this (although it is slightly technical).
MiniMax released open weights for H3 (aka Hailuo 3.0) last Monday, and it became the first open model ever to take the top spot in an AI video ranking, #1 for video editing on Artificial Analysis. Open weights means it runs on YOUR machine, not somebody’s API, and ComfyUI supported it on day one.
So naturally, the r/StableDiffusion crowd did the most r/StableDiffusion thing imaginable: they started remixing movies. Terminator. RoboCop. Old sci-fi getting rebuilt and re-shot on gaming PCs, with sound.
Here’s the thread where people are posting results. Some of it is rough. And some of it would’ve taken a small VFX house a week to pull off, two years ago.
The practical stuff: locally you’re capped at 768p (the 2K module stayed proprietary), clips run 4 to 15 seconds, and one prompt can take up to nine reference images, three video clips, and three audio clips. You can also fine-tune it on your own footage and your own look, which is the part I find most interesting for anyone trying to build a consistent style.
Uh, it also takes a LONG time on single graphics cards. You’re looking at 5+ minutes at least depending on your set-up. But I’m gonna pull out my gaming PC and give it a shot today!
Are you a creative or brand looking to go deeper with AI?
Join our community of collaborative creators on the AI4H Discord
Get exclusive access to all things AI4H on our Patreon
If you’re an org, consider booking Kevin & Gavin for your next event!

