7ART

Gemini Omni: What Google Actually Shipped, and How It Works on 7ART

Ilyas IIlyas I·August 13, 2026·10 min read·
Gemini Omni: What Google Actually Shipped, and How It Works on 7ART

Google shipped its next-generation video model at I/O 2026. It is not called Veo 4, and Veo 4 does not exist.

On 19 May 2026 Google announced Gemini Omni — "our new model that can create anything from any input — starting with video" (blog.google, checked 13 Aug 2026). The first model in the family, Gemini Omni Flash, went live the same day. It is a Gemini model, not a Veo model, and that distinction is the whole story: Google moved its flagship video effort onto the model line that reasons, rather than the one that renders.

This is a working breakdown. What Google documents, what Google says it can't do yet, where Veo actually stands, and exactly what you get when you pick it inside the 7ART video studio.

Note

This article replaces our earlier "Veo 4 prediction" post, which speculated about a model Google never announced. Rather than patch a piece built on a premise that turned out to be wrong, we retired it and rewrote the subject from primary sources. Every vendor claim below carries a link and the date we checked it.

What Google actually announced, and when

| Date | What shipped | |---|---| | 19 May 2026 | Gemini Omni announced at I/O. Omni Flash live in the Gemini app, Google Flow and Flow Music for Google AI Plus, Pro and Ultra subscribers globally — and free in YouTube Shorts Remix and the YouTube Create app for users 18+ | | 19 May 2026 | Model card published | | 30 Jun 2026 | Developer public preview — gemini-omni-flash-preview in the Gemini API, AI Studio and the Gemini Enterprise Agent Platform | | 16 Jul 2026 | Gemini Omni lands in Google Vids for AI Pro/Ultra and Workspace business tiers |

All dates from Google's own posts, checked 13 Aug 2026.

Two naming points worth getting right, because most coverage gets them wrong:

  • Gemini Omni is the family. Gemini Omni Flash is the model — it is the name on the model card and the root of the API id.
  • Gemini Omni Pro has been referred to in press coverage as a premium tier with details to come. There is no Google primary source, no date and no specification for it. Treat anything you read about its capabilities as unsourced.

The documented specification

Straight from Google's API docs (checked 13 Aug 2026):

  • Model id: gemini-omni-flash-preview — still a preview model, not GA
  • Output: 3–10 second video at 720p, 24 fps, with audio included by default
  • Aspect ratios: 16:9 (default) and 9:16. No others documented
  • Inputs: text, images (JPEG/PNG) and video (MP4 via the Files API). Multiple reference images accepted
  • Tasks: text_to_video, image_to_video, reference_to_video, edit
  • Context window: 1,048,576 tokens
  • Watermark: every generated video carries an invisible SynthID watermark
  • Price: input $1.50 per million tokens; video output $17.50 per million tokens, billed at 5,792 tokens per second of 720p video — Google's own worked figure is roughly $0.10 per second. No free tier

720p is the only output resolution Google documents. Hold onto that; it matters again in the 7ART section below.

The headline isn't generation. It's editing.

Every video model generates. What separates Omni Flash is that you can talk to a clip after it exists.

Google's model card frames it as "a more natural way to edit videos through conversation." Mechanically, it runs through the Interactions API, where each turn chains to the last via previous_interaction_id. Swap a background. Relight the scene for dusk. Change the camera angle. Restyle the animation. Each instruction is a sentence, not a re-render from scratch with a longer prompt and crossed fingers.

The second thing worth understanding is how it treats mixed references. Give it several images plus a clip and it does not stitch them — Google's framing, echoed in 9to5Google's launch coverage (19 May 2026), is that Omni reasons across all of the inputs to produce one consistent output. Google's announcement puts it as physics intuition combined with Gemini's knowledge of history, science and cultural context: the model "doesn't just build scenes that look real, it reasons about what should happen next."

In Google Flow, the practical payoff Google names is identity: character consistency, with "identity and voice preserved across every scene."

Google's builder showcase (7 Aug 2026) is the most concrete look at what that buys you — twenty camera perspectives of a single subject, day-to-night and season changes driven by voice, doodles that direct how elements move, and one walk cycle rendered as live action, then anime, then claymation.

What it still can't do

Google is unusually direct about the limits, and you should plan around them rather than discover them.

From the model card: maintaining complete consistency throughout edits, generating scenes with complex motion, and rendering perfectly accurate text all remain challenges. Speech editing is restricted.

Listed as unsupported in the API docs: audio reference uploading, video extension and interpolation, voice editing, system instructions, temperature, top_p and stop sequences, YouTube videos as a source, and provisioned throughput. English only. Video references are accepted up to three seconds but Google notes they are "not reliably processed."

Regional: uploading or editing images of minors, and editing uploaded videos, are unavailable in the EEA, Switzerland and the UK. Uploading or editing recognisable people is restricted.

And the 10-second ceiling: Google has been explicit that this is a product decision rather than a model limitation (TechCrunch, 19 May 2026), and the 30 June developer post says longer durations are coming. Don't architect a workflow on the assumption that ten seconds is permanent — but don't plan on it changing this quarter either.

Where this leaves Veo

Veo is not dead, and Omni did not replace it.

Veo 3.1 is still Google's current Veo model, released 15 October 2025 with richer native audio, greater narrative control, enhanced image-to-video, reference-image guidance, scene extension and first/last-frame transitions. Google's Veo page lists output "in 1080p and 4K." Google has never announced a Veo 4.

What did get retired are the older API models: veo-2.0-generate-001, veo-3.0-generate-001 and veo-3.0-fast-generate-001 were deprecated and shut down on 30 June 2026, with Veo 3.1 as the migration target.

Google's own model-selection guidance now reads as a straight division of labour: Omni Flash is the default video model; Veo 3.1 is for scene extension or existing pipeline integration. If you need a clip to run past ten seconds inside Google's stack today, that is a Veo 3.1 job, not an Omni job.

On 7ART both are in the same picker, on the same credit balance — VEO 3.1 Quality, Fast and Lite sit directly above Gemini Omni Video in the Google group. You are not choosing a subscription; you are choosing a shot.

Try it on 7ART
AI Video Generator

Animate your AI artist into Reels, TikToks, ads and cinematic clips. Kling, VEO 3.1, Seedance, Wan, Grok Imagine, MiniMax H3 and our own 7Video – every major video model in one workspace.

Open the tool

What you actually get on 7ART

7ART runs Gemini Omni Flash through kie.ai, with fal.ai as a fallback route. In the picker it is labelled Gemini Omni Video — that is our label, inherited from the provider's endpoint name, not a Google product name.

Output settings

| Control | Options on 7ART | |---|---| | Duration | 4, 6, 8 or 10 seconds (8 by default) | | Aspect ratio | 16:9 or 9:16 | | Output tier | 720p (default), 1080p, 4K — 4K requires a Pro or Unlimited plan |

One honest caveat on that last row. Those are 7ART output tiers offered through our provider. Google documents 720p only, and our own fal adapter notes that fal's Omni Flash endpoint has no resolution field at all. Anyone telling you Gemini Omni generates native 4K is stating something Google has never published.

Reference inputs. This is the model's most distinctive surface on 7ART, and it is unusual in that the reference panel is the only input surface — there are no start/end frame cards on this model at all.

  • Up to 7 reference images, each ≤ 20 MB
  • 1 video clip — file ≤ 30 s, with a trim window of ≤ 10 s
  • Up to 3 saved Characters
  • Up to 3 saved Voices

These share one budget, enforced live in the panel: images + 2 × videos + characters ≤ 7. A reference video costs two slots, so a clip plus three characters leaves you two image slots, not five.

Voices deserve a footnote. Google lists audio reference uploading as unsupported; the Voices control is a kie.ai capability layered on top, and it is speech, not a way to hand the model a music bed.

Credits

| Duration | 720p / 1080p | 4K | |---|---|---| | 4s | 90 | 210 | | 6s | 120 | 240 | | 8s | 150 | 270 | | 10s | 180 | 300 |

720p and 1080p bill identically. Attach a reference video and pricing switches to a flat rate per generation — 240 credits at 720p or 1080p, 360 at 4K — because the model determines the output duration itself in that mode.

Gemini Omni Video is also one of the five video engines available in Short Dramas, alongside Seedance 2.0, Seedance 2.5, MiniMax H3 and HappyHorse 1.1.

How to get good output from it

Four things that follow directly from the specification rather than from vibes.

Spend the reference budget on identity, not scenery. Seven slots is generous by the standards of video models, but every slot you burn on a location plate is a slot not spent pinning a face. Backgrounds are the thing the prompt describes well. Faces are not. If you already keep characters as saved objects for image and artist work, those are the references to reach for.

Write the shot, then edit the shot. The instinct carried over from older models is to front-load one enormous prompt because you only get one shot at it. Omni's design assumes the opposite — generate something close, then change one thing per turn. Prompts that try to specify ten attributes at once are where "complete consistency throughout edits" starts to slip, by Google's own admission.

Keep text out of the frame. Google names accurate text rendering as an open challenge. Titles, lower-thirds and packaging copy belong in your editor, not in the generation.

Treat 10 seconds as the unit, not the limit. Plan in beats that fit inside ten seconds and cut between them. If a single continuous shot genuinely has to run longer, that is a different model's job today — Veo 3.1 for scene extension inside Google's stack, or Seedance 2.5 (up to 30 seconds) and Grok Imagine (6–30 seconds) in the 7ART picker. See the full model line-up for what each one is actually for, and our Kling 3 prompt patterns for cinematography vocabulary that transfers across all of them.

The bottom line

The interesting fact about 2026's biggest video release is not a spec. It is that Google's next-generation video model shipped as a Gemini model.

Veo was a renderer that got progressively better at listening. Omni is a reasoner that renders — one model taking text, images and clips together, holding an identity across them, and accepting edits as conversation. It is a preview, it is capped at ten seconds and 720p by Google's own documentation, and Google will tell you itself that complex motion and legible text are still hard. It is still the clearest signal yet of where this category is going.

On 7ART it is one option in a picker, on one balance, next to VEO 3.1, Seedance, Kling, MiniMax H3, Wan, HappyHorse and Grok. Pick it when you need mixed references reasoned into one consistent shot. Pick something else when you need thirty seconds or 4K you can actually verify.

Run Gemini Omni Flash on 7ART

Up to 7 reference images, a video clip, saved characters and voices — 4 to 10 seconds, 16:9 or 9:16, from 90 credits. Same balance as every other model in the studio.

Try it on 7ART
AI Short Drama Generator

Take one idea to a finished vertical series. An AI showrunner writes the episodes and the cliffhangers, you cast the characters and lock the sets once, and every scene is filmed with the same faces.

Open the tool

Try the tools mentioned

Frequently asked questions

  • Gemini Omni is a Google model family announced at Google I/O on 19 May 2026, described by Google as 'our new model that can create anything from any input — starting with video.' The first and so far only shipped model in the family is Gemini Omni Flash. It takes text, images and video as input and returns video with audio. Source: blog.google, checked 13 August 2026.

  • No. Google has never announced a model called Veo 4. Google's current Veo model is Veo 3.1, released 15 October 2025. Gemini Omni is a Gemini model, not a Veo model, and Google's own API documentation now recommends Gemini Omni Flash as the default video model, with Veo 3.1 kept for scene extension and existing pipelines.

  • Google's API documentation specifies 3-10 second video at 720p, 24 fps, in 16:9 or 9:16, with audio. That is the only output resolution Google documents. On 7ART the model is offered with 720p, 1080p and 4K output tiers through our provider — that is a 7ART product option, not a documented Google specification, and 720p is our default.

  • Google documents 3-10 seconds. Google has said the cap is a product decision rather than a model limit, and its 30 June 2026 developer post says longer durations are coming. On 7ART you pick 4, 6, 8 or 10 seconds, with 8 as the default.

  • No. Both are live and Google routes them to different jobs — Omni Flash as the default video model, Veo 3.1 for scene extension or existing pipelines. What did get retired are the older Veo API models: veo-2.0-generate-001, veo-3.0-generate-001 and veo-3.0-fast-generate-001 were deprecated and shut down on 30 June 2026, with Veo 3.1 as the migration target.

  • At 720p or 1080p a clip costs 90 / 120 / 150 / 180 credits for 4 / 6 / 8 / 10 seconds. At 4K it is 210 / 240 / 270 / 300. If you attach a reference video, billing switches to a flat 240 credits at 720p or 1080p, and 360 at 4K. The 4K tier requires a Pro or Unlimited plan.

  • Google's API documentation lists audio reference uploading as unsupported. 7ART's Voices control is a provider feature from kie.ai, not a documented Google API capability: you can attach up to three saved Voices, and it is a speech capability, not a way to feed the model a music track.

Ilyas I
Written by

Ilyas I

Covers AI model releases, head-to-head comparisons, and deep technical breakdowns of image, video, and music generators. Part of the 7ART team.

Stay updated

Newsletter signup coming soon.

Continue reading