APIXO
Video GenerationMultimodal InputPreview

Gemini Omni Flash is Google’s preview multimodal model for video generation and editing. APIXO exposes video generation alongside reusable audio-asset and character-asset workflows, with text, image, video, and saved asset references. Create or transform videos with coordinated visuals and sound, world-aware scene logic, and APIXO output options up to 4K.

Inputs

Text / Image / Video / Audio Assets

APIXO Price

$0.10 / secondSource-video edits use per-operation pricing

APIXO Resolution

720p / 1080p / 4K

Duration

4 / 6 / 8 / 10 SecondsWithout source-video input

Released

May 30, 2026

Create with Gemini Omni

Loading workspace...

Available on APIXO

Build Multimodal Video Workflows with Gemini Omni

Create short videos from prompts and mixed references, or transform one source video through natural-language direction. APIXO adds reusable audio and character asset modes around its Gemini Omni video route, letting applications prepare voices and subjects before combining them with images, video, framing, duration, and resolution controls.

Capabilities

Combine Creative References into Coherent Video Output

Mixed-Reference Generation

APIXO’s video mode accepts prompts alongside images, one source video, reusable character IDs, and audio IDs. These references can guide appearance, movement, scene structure, and sound within a single video-generation request.

Source-Video Transformation

Supply one video and describe the transformation, such as changing the environment, visual treatment, action, or selected scene elements. APIXO also exposes start and end positions for choosing the relevant portion of the input.

Reusable Character Assets

Create a character asset from exactly one image, optionally associated with one prepared audio asset. The returned character ID can then be reused in video requests, with APIXO supporting as many as three character IDs per generation.

Reusable Audio Assets

Build an audio asset from an APIXO-supported base voice, an optional voice description, and preview dialogue. Save the resulting audio ID for later character or video requests; video mode accepts up to three audio IDs.

World-Aware Scene Logic

Google positions Gemini Omni Flash around world knowledge and an understanding of physical behavior. This helps it construct scenes informed by history, science, cultural context, gravity, motion, materials, and other real-world relationships.

Coordinated Visuals and Sound

Gemini Omni Flash generates video with an accompanying audio track and can respond to timing instructions for dialogue, music, effects, cuts, and on-screen events. Prompt precision remains important when exact synchronization or text timing is required.

API Details

Gemini Omni Configuration on APIXO

Verified asset modes, reference allowances, output settings, and quotas published for the APIXO route.

Audio Asset / Character Asset / Video

Modes

16:9 / 9:16

Aspect Ratios

Up to 3 IDs

Audio Assets

Up to 3 IDs

Character Assets

Up to 1 URL

Source Video

7 Units

Reference Quota

Production Uses

Build Reference-Driven Video Production Systems

Brand Campaigns

Consistent Presenter Videos

Prepare a reusable character from an approved portrait and pair it with a selected audio asset. Applications can then generate short presenter-led campaign videos across different scenes while reusing the same character and voice references.

Product Creative

Multimodal Product Stories

Combine product images, character assets, audio direction, and a written scene brief to create launch clips or advertising concepts. Use landscape or portrait framing to prepare variations for storefronts, campaign pages, and mobile channels.

Video Adaptation

Source-Footage Restyling

Submit an existing clip and request a new environment, visual style, object treatment, or narrative effect. Start and end controls let developers select the relevant source segment before applying the transformation through APIXO.

Knowledge Media

Visual Explainers and Storyboards

Use Gemini Omni’s world knowledge and physical reasoning to develop short science, history, process, or concept explainers with coordinated narration and motion. Treat factual output as draft material and verify important details before publication.

Asset Workflow

Prepare Reusable Inputs Before Generating Video

APIXO separates optional audio and character preparation from asynchronous Gemini Omni video generation.

Create an Audio Asset

Select an APIXO-supported base voice, name the asset, and optionally provide voice direction and preview dialogue. Save the returned audio ID for later use with a character or directly within a video request.

Prepare a Character Asset

Submit exactly one public character image with a descriptive prompt and, when useful, one saved audio ID. Store the resulting character ID so it can be referenced in future Gemini Omni video tasks.

Generate the Final Video

Combine the prompt with eligible images, audio IDs, character IDs, or one source video. Choose an APIXO-supported resolution and framing option, then retrieve the completed video through asynchronous polling or webhook delivery.

Notes & FAQ

Gemini Omni Availability and Constraints

Operational Notes

01

This page covers APIXO’s current Gemini Omni video, audio-asset, and character-asset modes. It does not represent Gemini chat, image-generation, document-analysis, or standalone speech APIs.

02

APIXO’s browser Playground is fixed to video mode; audio and character asset preparation is documented for API use.

03

APIXO does not document direct audio-file uploads for this route; video requests use audio IDs created through its audio-asset mode.

04

The combined image, source-video, and character quota is seven units; one video consumes two units, while each image or character consumes one.

05

APIXO does not expose Gemini’s stateful previous_interaction_id workflow, real-time sessions, document analysis, or tool use on this endpoint.

06

English is fully supported upstream; other languages have not been evaluated and may produce less predictable results.

Frequently Asked Questions

No. This APIXO page covers Google’s Gemini Omni video workflow, not Gemini chat, coding, document-analysis, or agent models. Its published outputs are reusable audio or character asset IDs and generated video URLs, depending on the selected APIXO mode.

APIXO exposes three modes: audio-asset creation, character-asset creation, and asynchronous video generation. The asset modes prepare reusable IDs for video requests; they are not standalone music, speech-transcription, image-generation, or conversational-response endpoints.

Without source-video input, 720p and 1080p cost $0.10 per generated second, while 4K costs $0.20 per second. With one source video, APIXO charges $1.20 per request at 720p or 1080p and $1.80 at 4K.

Prompt- and reference-generated videos without a source video support 4, 6, 8, or 10 seconds. When a source-video URL is supplied, APIXO ignores the duration setting and instead uses the selected input segment and fixed per-request billing.

APIXO documents prompt-based transformation of one uploaded source video, but it does not publish Google’s stateful multi-turn interaction parameter on this endpoint. Treat each APIXO task as an independent request unless the live schema later documents conversational state.

Explore Other Models

Discover more AI models for your next creative workflow