No. This APIXO page covers Google’s Gemini Omni video workflow, not Gemini chat, coding, document-analysis, or agent models. Its published outputs are reusable audio or character asset IDs and generated video URLs, depending on the selected APIXO mode.
Gemini Omni
Gemini Omni Flash is Google’s preview multimodal model for video generation and editing. APIXO exposes video generation alongside reusable audio-asset and character-asset workflows, with text, image, video, and saved asset references. Create or transform videos with coordinated visuals and sound, world-aware scene logic, and APIXO output options up to 4K.
Inputs
Text / Image / Video / Audio Assets
APIXO Price
$0.10 / secondSource-video edits use per-operation pricing
APIXO Resolution
720p / 1080p / 4K
Duration
4 / 6 / 8 / 10 SecondsWithout source-video input
Released
May 30, 2026
Create with Gemini Omni
Loading workspace...
Combine Creative References into Coherent Video Output
Mixed-Reference Generation
APIXO’s video mode accepts prompts alongside images, one source video, reusable character IDs, and audio IDs. These references can guide appearance, movement, scene structure, and sound within a single video-generation request.
Source-Video Transformation
Supply one video and describe the transformation, such as changing the environment, visual treatment, action, or selected scene elements. APIXO also exposes start and end positions for choosing the relevant portion of the input.
Reusable Character Assets
Create a character asset from exactly one image, optionally associated with one prepared audio asset. The returned character ID can then be reused in video requests, with APIXO supporting as many as three character IDs per generation.
Reusable Audio Assets
Build an audio asset from an APIXO-supported base voice, an optional voice description, and preview dialogue. Save the resulting audio ID for later character or video requests; video mode accepts up to three audio IDs.
World-Aware Scene Logic
Google positions Gemini Omni Flash around world knowledge and an understanding of physical behavior. This helps it construct scenes informed by history, science, cultural context, gravity, motion, materials, and other real-world relationships.
Coordinated Visuals and Sound
Gemini Omni Flash generates video with an accompanying audio track and can respond to timing instructions for dialogue, music, effects, cuts, and on-screen events. Prompt precision remains important when exact synchronization or text timing is required.
Gemini Omni Configuration on APIXO
Verified asset modes, reference allowances, output settings, and quotas published for the APIXO route.
Audio Asset / Character Asset / Video
Modes
16:9 / 9:16
Aspect Ratios
Up to 3 IDs
Audio Assets
Up to 3 IDs
Character Assets
Up to 1 URL
Source Video
7 Units
Reference Quota
Build Reference-Driven Video Production Systems
Consistent Presenter Videos
Prepare a reusable character from an approved portrait and pair it with a selected audio asset. Applications can then generate short presenter-led campaign videos across different scenes while reusing the same character and voice references.
Multimodal Product Stories
Combine product images, character assets, audio direction, and a written scene brief to create launch clips or advertising concepts. Use landscape or portrait framing to prepare variations for storefronts, campaign pages, and mobile channels.
Source-Footage Restyling
Submit an existing clip and request a new environment, visual style, object treatment, or narrative effect. Start and end controls let developers select the relevant source segment before applying the transformation through APIXO.
Visual Explainers and Storyboards
Use Gemini Omni’s world knowledge and physical reasoning to develop short science, history, process, or concept explainers with coordinated narration and motion. Treat factual output as draft material and verify important details before publication.
Prepare Reusable Inputs Before Generating Video
APIXO separates optional audio and character preparation from asynchronous Gemini Omni video generation.
Create an Audio Asset
Select an APIXO-supported base voice, name the asset, and optionally provide voice direction and preview dialogue. Save the returned audio ID for later use with a character or directly within a video request.
Prepare a Character Asset
Submit exactly one public character image with a descriptive prompt and, when useful, one saved audio ID. Store the resulting character ID so it can be referenced in future Gemini Omni video tasks.
Generate the Final Video
Combine the prompt with eligible images, audio IDs, character IDs, or one source video. Choose an APIXO-supported resolution and framing option, then retrieve the completed video through asynchronous polling or webhook delivery.
Gemini Omni Availability and Constraints
Operational Notes
This page covers APIXO’s current Gemini Omni video, audio-asset, and character-asset modes. It does not represent Gemini chat, image-generation, document-analysis, or standalone speech APIs.
APIXO’s browser Playground is fixed to video mode; audio and character asset preparation is documented for API use.
APIXO does not document direct audio-file uploads for this route; video requests use audio IDs created through its audio-asset mode.
The combined image, source-video, and character quota is seven units; one video consumes two units, while each image or character consumes one.
APIXO does not expose Gemini’s stateful previous_interaction_id workflow, real-time sessions, document analysis, or tool use on this endpoint.
English is fully supported upstream; other languages have not been evaluated and may produce less predictable results.
Frequently Asked Questions
APIXO exposes three modes: audio-asset creation, character-asset creation, and asynchronous video generation. The asset modes prepare reusable IDs for video requests; they are not standalone music, speech-transcription, image-generation, or conversational-response endpoints.
Without source-video input, 720p and 1080p cost $0.10 per generated second, while 4K costs $0.20 per second. With one source video, APIXO charges $1.20 per request at 720p or 1080p and $1.80 at 4K.
Prompt- and reference-generated videos without a source video support 4, 6, 8, or 10 seconds. When a source-video URL is supplied, APIXO ignores the duration setting and instead uses the selected input segment and fixed per-request billing.
APIXO documents prompt-based transformation of one uploaded source video, but it does not publish Google’s stateful multi-turn interaction parameter on this endpoint. Treat each APIXO task as an independent request unless the live schema later documents conversational state.
Explore Other Models
Discover more AI models for your next creative workflow