Google DeepMind launched Gemini Omni 1.1 Flash today, a video generation model that reads 10 seconds of prior footage to extend scenes, ten times more than the single second earlier versions could manage S¹. That jump changes what developers can build: clips that hold visual coherence across cuts instead of resetting every second. But the total video still caps at 40 seconds, and every performance claim comes from Google alone.
My read: This is the first video generation model I've seen that treats scene extension as a real editing tool rather than a novelty. The jump from 1 second to 10 seconds of context is the difference between continuing a shot and continuing a scene. I don't buy the "production-ready" label yet, because there's no independent benchmarking and Google's own model card P² is the only verification we have. The 360p tier at one-third cost is smart: it lets developers iterate cheaply before paying for 4K. What I'd watch is whether the 40-second cap holds, and whether keyframe interpolation produces coherent transitions or just smooth morphs.
From one second to ten
The previous generation of Gemini Omni could extend a video, but only by looking at the last second of footage S¹. One second is roughly 24 to 30 frames. The model had to guess what came next from a sliver of motion. Omni 1.1 Flash reads up to 10 seconds, giving it 10 times more visual history to work with S¹. It can track a character walking through a doorway, maintain a lighting shift from a passing cloud, or keep a camera pan consistent across the extension point.
Scene extension works in 10-second increments, building up to a cumulative cap of 40 seconds S¹. For a developer building a short product demo or a social ad, 40 seconds is workable. For anyone thinking about longer narrative work, it is a hard wall.

The 360p trick
Google added a 360p preview tier that generates up to 60% faster than the standard 720p resolution, based on system throughput S¹. It costs one-third the price of 720p S¹. The idea is straightforward: iterate cheaply at low resolution, then render the final cut at 1080p or 4K S¹.
Video generation is expensive. Every prompt tweak, every keyframe adjustment, every scene extension costs compute. A developer who burns through 20 iterations before landing on the right shot pays for all 20. Dropping the draft cost to a third of the final price makes the iteration loop less painful.
Pinning the start and end
Omni 1.1 can generate continuous video between two user-specified keyframes S¹. You give it a start frame and an end frame, and it fills in the motion between them. This is the kind of control that turns video generation from a slot machine into something closer to a tool. Instead of re-rolling a prompt until the output looks right, a developer can pin the beginning and end points and let the model handle the transition.
The broader Gemini Omni product line positions itself as conversational video editing. Google's own product page describes it as "like Nano Banana, but for video," where each edit builds on the previous one to maintain a coherent scene P⁴. The API reference in Google's skills repository shows the developer interface for media generation, including image generation with the gemini-3.1-flash-image model P³.
Omni 1.1 Flash follows a clear pattern: more knobs, more tiers, more ways to spend less while experimenting.
What to do about it
A small marketing agency producing product videos for e-commerce clients could use the 360p tier to test five or six scene variations for a 15-second ad spot at a fraction of the 720p cost, then render the winner at 4K. The keyframe feature lets the art director specify the opening product shot and the closing lifestyle shot, letting the model generate the transition. The 40-second cap fits a social ad or a product demo but stops at anything resembling a long-form piece.
If you build with the Gemini API, the practical move this week is to test scene extension on an existing clip you already own. Feed it 10 seconds of footage, ask for a 10-second extension, and check whether the lighting, subject position, and camera motion stay consistent across the seam. That seam is where every video extension model either earns trust or loses it.
What we don't know yet
Every number in this story comes from Google. The 60% speed claim is a system throughput comparison, not a guarantee for every generation task S¹. The "production-ready" label is Google's own assessment, with no independent benchmarking yet S¹. The model card published the same day, 27 August 2026, provides Google's summary of known limitations P², but no third-party evaluations have appeared for output quality, coherence across the 40-second cap, or how keyframe interpolation handles complex transitions.
The 40-second cumulative limit is a real constraint that the "longer storytelling" framing in Google's blog does not fully address. Whether Google raises that cap, and how soon, will tell us whether Omni 1.1 is a stepping stone or the model Google expects developers to ship on.
The next signal: Google's model card for Omni Flash, which the card itself says "may be updated from time to time to include updated evaluations" P². When the first revision lands, we'll check whether independent evaluations confirm the throughput and cost claims. If you want that follow-up in your inbox, subscribe and we'll send it.
Sources: S1 — Gemini Omni 1.1 Flash lets you build with more control · P2 — Gemini Omni Flash - Model Card — Google DeepMind · P3 — skills/cloud/gemini-api/references/media_generation.md · P4 — Gemini Omni — Google DeepMind · P5 — google-deepmind/tips
More from Not A Tech Guy
- PyTorch hits GitHub trending at 102,613 stars
- DeepMind's double-blind AI test locks benchmarks in crypto box
- AI coding agent defense cuts malware severity 83%
Generated from an audited evidence pack with primary-source research. Social-media items are discussion signals, not verified facts. Nothing here is financial, legal or medical advice.