Today on AI For Humans:
Gemini Omni Cracks Video Editing
The Pope Pontificates All About AI
Plus, We Built An AI Focus Group?!

Welcome to the AI For Humans newsletter!

Last week at I/O 2026 (our coverage here), Google quietly dropped a video model that, if you squint, looks like the beginning of a pretty massive shift in how AI video actually works.

It's called Gemini Omni Flash. And unlike Veo or any of the AI video models you've used before, it wasn't built to be a video generator first.

It was built to be a world model first.

Which sounds like a small distinction. It is not.

The thing Ethan is showing off in that tweet, the seamless editing, the fact that you can drop something into a clip and have the model actually understand the space it's living in, is the part that has me thinking we're at a real turning point.

Let's get into it!

Please support AI For Humans by learning about our sponsors below:

We’ve partnered with HP & Intel to promote the ZBook Fury Pro workstation.

For more info & to support A4H, click here.

How Omni Flash Is Actually Differs From Veo

So I went down the rabbit hole on the actual architecture this week and was a little surprised.

Omni Flash isn't just a "better Veo."

Veo (technically a latent diffusion transformer, if you care about that stuff) is a generator first. It's gotten really good at predicting what frames should look like. It even picked up a kind of intuitive sense of physics along the way.

But at the end of the day, Veo is still generating, not simulating.

Same goes for basically every other AI video tool you've used. Runway. Kling. All of them.

Omni Flash is doing something else.

It's the first Google model that fuses world model reasoning with video output.

Before Omni generates anything, it actually tries to understand the scene first. The objects. Where they are. How they should move. How light should hit them. How a new element you drop in would behave in that physical space.

You can see the shadow of the Ogre in the background of this video as a good example.

The multimodal input thing is also kind of wild.

Veo takes text and video. Sora 2 took text and image (before OpenAI killed it). Runway and Kling, same.

Omni takes text, image, audio, and video. All in a single prompt.

Which is why working in Omni feels less like writing a prompt and more like talking to an editor.

That's the world model part.

The Best AI Editor I’ve Ever Used

The way I keep thinking about it:

A traditional video model is basically a really, really good guesser. It looks at frame one, then guesses what frame two should be, then guesses what frame three should be, and so on. The output looks like a video, but the model isn't really thinking about what's happening inside it.

A world model is doing something closer to what your brain does when you watch a movie. It's tracking what objects are, where they are, how they should behave, and what should happen next based on actual physics and logic.

Once you have that, you start to extrapolate.

This isn't just about linear video anymore. It's about interactive environments. It's about games that generate themselves around you. It's about training simulations that look and feel real. It's about, eventually, simulations of entire coherent worlds.

Kind of like how Google’s Genie 3 (mentioned here previously) does in this new real-world Google Street View upgrade this week.

And, yes, this might be the very early stage of the simulation hypothesis becoming a thing you can actually build, not just argue about at 2am.

Quick refresher: simulation theory is the idea (popularized by Nick Bostrom and absolutely beaten to death by Elon) that if any civilization eventually gets the ability to simulate reality at a high enough resolution, the math suggests we're probably already inside one of those simulations.

I'm not saying Gemini Omni Flash means we live in The Matrix.

I am saying that the gap between "AI that generates a 10-second clip of a dog" and "AI that can simulate an entire coherent reality" just got noticeably smaller this week.

Love this newsletter? Forward it to one curious friend. They can join in one click.

Not Perfect Yet, But These Are Baby Steps

Omni Flash is, of course, not perfect.

There are still plenty of things it fumbles. Lighting drifts. Edits sometimes feel slightly off. The model occasionally just decides the laws of physics are more of a suggestion.

But video editing is the obvious starting point for something much bigger. And remember, this is the Flash model. The smaller, faster, cheaper version.

Which means there's almost certainly a much bigger sibling already being trained somewhere inside Google.

The idea that we can now take a clip of something and reliably replace pieces of it is transformative right now for Hollywood and filmmakers. We're going to see this in commercials, in indie films, in fan edits, in YouTube videos, all within the next few months.

And further out? You might literally be using this to change your entire experience of the world. Drop yourself into a movie. Re-skin your morning walk. Build a custom training environment for whatever skill you're trying to learn.

But for now, we only get ten seconds at a time.

-Gavin

This week on AI For Humans: Spotify Goes Hard Into AI Music👇

3 Things To Know About AI Today

Claude Mythos Might Finally Be Coming

Rumors are swirling that we're going to finally see Claude Mythos at some point soon, and the first place it might show up is inside Claude Code.

A few possibilities here.

Maybe Anthropic gave themselves enough lead time to patch the security stuff that's been quietly worrying people.

Or maybe (and this is where my brain goes) Mythos lands inside Claude Code first specifically because that's where the real power users live, and Anthropic wants to stress-test it on the people most likely to push it to its limits.

Either way, if it drops this week, we'll be all over it.

DeepMind Just Solved Nine More Erdős Problems

Last week we talked about OpenAI's big advancement into solving one of those long-considered-unsolvable Erdős mathematics problems.

Well, DeepMind's AlphaProof Nexus just solved nine more. (Plus 44 open conjectures from the OEIS, for good measure.)

These are problems that have, in some cases, been open for decades.

Watching them fall one after another in the span of a few weeks is the kind of AI moment that doesn't get the same headline attention as a new model drop, but it's way more important.

This is the actual frontier of "AI doing things humans couldn't do alone."

Math today. Science very soon.

The Pope's First Encyclical Is About AI

Pope Leo XIV's very first encyclical, Magnifica Humanitas ("on the protection of the human person in the age of artificial intelligence"), got presented at the Vatican this morning.

It should be live here to watch but if not, you might just have to Google it.

A quick context note for non-Catholic readers: an encyclical is one of the most important things a Pope can write. It's the official teaching document. Not a speech, not an interview, the actual Vatican-level position.

This is the first one of Leo XIV's papacy, and he picked AI.

The kicker: he's co-presenting it with one of Anthropic's co-founders, which is a sentence I didn’t think I'd ever write.

If the Pope is making AI his Day One issue, and dragging an AI lab co-founder up to the Vatican to do it with him, that should tell you something about where this conversation is going.

We 💛 This: I Built An AI Focus Group (And You Can Use It)

Ok, I'm using this week's We 💛 This to show off something I've been quietly working on for a long time and is (kind of) now out in the world:

The tl;dr: it's a visual representation of an AI focus group for anything you want to test. A startup idea. A go-to-market plan. A creative concept. A personal decision you've been wrestling with.

A screenshot from the demo which you can play/watch here.

You drop in your idea, and you watch a room of AI personas actually talk to each other about it. They argue, push back, build on each other's points. Configurable participants, real conversation, you just watch and read. Then ask follow ups if you want.

I built the whole thing myself with Claude (you can read why I did it here), and I'll talk more about it on the show this week.

Short version: I think we need more ways beyond chatbots to help people understand the sorts of things AI can do. The Fishbowl is mostly a visualization of AI agentic conversations but I’m beginning to think that wrappers like this that help people better understand how to actually use AI are going to be a big deal.

A couple of caveats:

It's not free to run, because it's calling the Claude API under the hood. But I've loaded a few bucks into it so you can try it for free.

I’ve also thrown a Buy Me A Coffee Link on the page that if you wanna donate a few bucks to I can run it for a bit longer.

You can just watch the demo if you want and that costs me nothing.

Even better, I’m putting the whole thing up open source on GitHub with an MIT license if you want to run it yourself.

Let me know what you test in it. I'm super curious what people use this for.

PS, you’re seeing this before pretty much anyone else. I’ll post publicly about it later this week.

Are you a creative or brand looking to go deeper with AI?
Join our community of collaborative creators on the AI4H Discord
Get exclusive access to all things AI4H on our Patreon
If you’re an org, consider booking Kevin & Gavin for your next event!

Keep Reading