Google DeepMind launched Gemini Omni 1.1 Flash today, a video generation model that reads 10 seconds of prior footage to extend scenes, ten times more than the single second earlier versions could manage . That jump changes what developers can build: clips that hold visual coherence across cuts instead of resetting every second. But the total video still caps at 40 seconds, and every performance claim comes from Google alone.

My read: This is the first video generation model I've seen that treats scene extension as a real editing tool rather than a novelty. The jump from 1 second to 10 seconds of context is the difference between continuing a shot and continuing a scene. I don't buy the "production-ready" label yet, because there's no independent benchmarking and Google's own model card is the only verification we have. The 360p tier at one-third cost is smart: it lets developers iterate cheaply before paying for 4K. What I'd watch is whether the 40-second cap holds, and whether keyframe interpolation produces coherent transitions or just smooth morphs.

From one second to ten

The previous generation of Gemini Omni could extend a video, but only by looking at the last second of footage . One second is roughly 24 to 30 frames. The model had to guess what came next from a sliver of motion. Omni 1.1 Flash reads up to 10 seconds, giving it 10 times more visual history to work with . It can track a character walking through a doorway, maintain a lighting shift from a passing cloud, or keep a camera pan consistent across the extension point.

Scene extension works in 10-second increments, building up to a cumulative cap of 40 seconds . For a developer building a short product demo or a social ad, 40 seconds is workable. For anyone thinking about longer narrative work, it is a hard wall.

Prior context for scene extension

The 360p trick

Google added a 360p preview tier that generates up to 60% faster than the standard 720p resolution, based on system throughput . It costs one-third the price of 720p . The idea is straightforward: iterate cheaply at low resolution, then render the final cut at 1080p or 4K .

Video generation is expensive. Every prompt tweak, every keyframe adjustment, every scene extension costs compute. A developer who burns through 20 iterations before landing on the right shot pays for all 20. Dropping the draft cost to a third of the final price makes the iteration loop less painful.

Pinning the start and end

Omni 1.1 can generate continuous video between two user-specified keyframes . You give it a start frame and an end frame, and it fills in the motion between them. This is the kind of control that turns video generation from a slot machine into something closer to a tool. Instead of re-rolling a prompt until the output looks right, a developer can pin the beginning and end points and let the model handle the transition.

The broader Gemini Omni product line positions itself as conversational video editing. Google's own product page describes it as "like Nano Banana, but for video," where each edit builds on the previous one to maintain a coherent scene P⁴. The API reference in Google's skills repository shows the developer interface for media generation, including image generation with the gemini-3.1-flash-image model .

Omni 1.1 Flash follows a clear pattern: more knobs, more tiers, more ways to spend less while experimenting.

What to do about it

A small marketing agency producing product videos for e-commerce clients could use the 360p tier to test five or six scene variations for a 15-second ad spot at a fraction of the 720p cost, then render the winner at 4K. The keyframe feature lets the art director specify the opening product shot and the closing lifestyle shot, letting the model generate the transition. The 40-second cap fits a social ad or a product demo but stops at anything resembling a long-form piece.

If you build with the Gemini API, the practical move this week is to test scene extension on an existing clip you already own. Feed it 10 seconds of footage, ask for a 10-second extension, and check whether the lighting, subject position, and camera motion stay consistent across the seam. That seam is where every video extension model either earns trust or loses it.

What we don't know yet

Every number in this story comes from Google. The 60% speed claim is a system throughput comparison, not a guarantee for every generation task . The "production-ready" label is Google's own assessment, with no independent benchmarking yet . The model card published the same day, 27 August 2026, provides Google's summary of known limitations , but no third-party evaluations have appeared for output quality, coherence across the 40-second cap, or how keyframe interpolation handles complex transitions.

The 40-second cumulative limit is a real constraint that the "longer storytelling" framing in Google's blog does not fully address. Whether Google raises that cap, and how soon, will tell us whether Omni 1.1 is a stepping stone or the model Google expects developers to ship on.

The next signal: Google's model card for Omni Flash, which the card itself says "may be updated from time to time to include updated evaluations" . When the first revision lands, we'll check whether independent evaluations confirm the throughput and cost claims. If you want that follow-up in your inbox, subscribe and we'll send it.


Sources: S1 — Gemini Omni 1.1 Flash lets you build with more control · P2 — Gemini Omni Flash - Model Card — Google DeepMind · P3 — skills/cloud/gemini-api/references/media_generation.md · P4 — Gemini Omni — Google DeepMind · P5 — google-deepmind/tips

More from Not A Tech Guy


Generated from an audited evidence pack with primary-source research. Social-media items are discussion signals, not verified facts. Nothing here is financial, legal or medical advice.