Gemini Omni 1.1 Flash Launches: 4K Upscaling and Keyframe Control Explained

Gemini Omni 1.1 Flash goes GA with keyframe control, draft mode, and 4K upscaling for production-ready AI video.
Google has announced the general availability of Gemini Omni 1.1 Flash, bringing four major upgrades: a 10-second scene extension context window to reduce shot drift; first/last frame interpolation for precise keyframe-based control; a 360p draft mode at one-third the standard cost for rapid iteration; and 4K upscaling to complete the full output pipeline. At $0.10 per second for 720p video with tiered pricing, all features are accessible via the Gemini API — signaling Google's developer-first strategy and the industry's shift from "can generate" to "controllable generation."
A Major Step Forward for Google's Video Generation
Google today announced the general availability (GA) of Gemini Omni 1.1 Flash, marking a formal transition of its video generation capabilities from preview to production-ready. Compared to previous versions, this update delivers meaningful improvements across four dimensions: context length, frame-level control, cost optimization, and resolution.
For developers and content creators, these improvements go beyond spec-sheet gains — they represent a critical turning point where AI video generation moves from "impressive demos" to "controllable production." Let's break down each core capability in detail.

Four Core Capability Upgrades, Explained
Scene Extension: A 10-Second Context Window
The new version introduces a 10-second Scene Extension context window. When generating new clips, the model can reference the previous 10 seconds of footage to maintain consistency in camera movement, lighting, and subject appearance.
Previously, one of the biggest pain points in AI video generation was "drift" in longer shots — subjects would gradually distort or scenes would jump abruptly over time. While 10 seconds isn't an especially long window, it's practically useful for seamlessly stitching together short video segments and provides a solid technical foundation for multi-shot storytelling.
First and Last Frame Interpolation: Keyframe Control
First and Last Frame Interpolation is arguably the most creatively valuable feature in this update. Developers can specify both the opening and closing frames of a video, and the model automatically generates the transition animation in between.
This "keyframe control" approach draws inspiration from traditional animation workflows: creators define the key moments, and AI handles the in-betweening. Compared to generating video purely from text prompts, this method returns significantly more creative control to the user and dramatically reduces unpredictability in the output. It's especially well-suited for commercial ads and product demos where precise start and end frames are required.
360p Draft Mode: Low-Cost Rapid Iteration
Google has also introduced a 360p draft mode that costs just one-third of the standard generation price. This feature directly addresses a real need in video production workflows — before finalizing a piece, creators typically need to iterate repeatedly on composition and camera movement.
The idea is to validate creative direction quickly with low-resolution, low-cost drafts, then commit to high-quality rendering only after confirming the result. This "draft-first, finalize-later" workflow can significantly reduce the cost of experimentation, making large-scale creative exploration economically viable.
4K Upscaling
The new version supports 4K upscaling. Generated videos can be enhanced to 4K resolution, meeting the demanding quality requirements of professional-grade content distribution. This completes the full output pipeline — from 360p draft all the way to 4K finished footage.
Pricing Analysis: What Does $0.10 Per Second Actually Mean?
According to official pricing, generating 720p video costs approximately $0.10 per second. This figure deserves a closer look.
A 10-second 720p clip would cost roughly $1. For individual creators, this requires careful budget management at scale; but for commercial teams, AI generation already offers a clear economic advantage over traditional video production in terms of time and labor costs.
When combined with the 360p draft mode (at one-third the cost), the overall cost of a real production workflow can be reduced further. This tiered pricing strategy — low cost for drafts, standard pricing for finished clips, and premium pricing for 4K — reflects Google's deep understanding of actual creative workflows, and gives users across different budget levels a viable entry point.
Industry Significance and What's Next
The general availability of Gemini Omni 1.1 Flash reflects a broader trend in AI video generation: the shift from "can generate" to "controllable generation." Keyframe control and scene extension both fundamentally address the same core challenge: giving creators stronger, more predictable control over generated output.
Notably, all capabilities in this update are available through the Gemini API, allowing developers to integrate video generation directly into their own products and workflows — rather than being confined to Google's first-party applications. This API-first strategy has the potential to spark a new wave of downstream tools and creative applications built on AI video generation.
That said, a 10-second context window and a 720p baseline resolution still lag behind some competing offerings. But given the cost advantages and the thoughtfully designed tiered workflow, Gemini Omni 1.1 Flash strikes a pragmatic balance between usability and value. As video generation models continue to iterate, there's good reason to anticipate a next generation with longer context, higher resolution, and finer-grained control.
Related articles

DeepSeek V4 Pro Burning Through Credits Too Fast? The Hidden Logic Behind AI Model Pricing
Why does DeepSeek V4 Pro drain credits so fast while Flash barely moves? A deep dive into AI token billing, Pro vs. Flash pricing differences, and cost optimization tips.

RealPDE Competition Breakdown: The Frontier Challenge of AI-Powered Real-World Fluid Dynamics PDE Solving
A deep dive into the NeurIPS 2026 RealPDE Competition, covering the Sim2Real and LTTTA tracks, and how neural operators tackle real-world PIV and CFD fluid PDE challenges.

Building a Production-Grade 3DGS Training Library from Scratch: A Deep Dive into Full-GPU Residency and the Vulkan Stack
A veteran graphics engineer builds a production-grade 3DGS training library from scratch using C++23, CUDA, and Vulkan, achieving 60fps with 5M splats. Deep dive into its architecture and design.