This page covers the original InfiniteTalk research release from MeiGen-AI, published with its technical report, code, and model weights on August 19, 2025. It does not describe LongCat Video Avatar, community checkpoints, or unrelated avatar systems.
InfiniteTalk is MeiGen-AI’s audio-driven human video model for synchronizing speech with facial expression, head movement, and body posture—not only the mouth. APIXO provides a focused portrait-and-audio workflow with an optional mask, up to ten minutes of driving audio, 480p or 720p output, and asynchronous delivery.
APIXO Inputs
Portrait + audio
APIXO Price
From $0.03 / second480p, five-second minimum
APIXO Resolution
480p / 720p
Audio Limit
Up to 600 seconds
Released
Aug 19, 2025
Create with InfiniteTalk
Loading workspace...
Audio drives the wider human performance
Speech-Synchronized Motion
InfiniteTalk uses the driving audio to coordinate mouth movement with expression, head motion, and body posture. Lip synchronization is a core objective, although accuracy can still vary with speech speed, pose, and source quality.
Long-Sequence Continuity
InfiniteTalk’s upstream framework uses temporal context between generated segments to support longer audio-driven sequences. APIXO accepts driving audio up to 600 seconds, although long outputs should still be reviewed for identity drift and transition consistency.
Portrait-to-Performance
APIXO’s implementation starts from one portrait image rather than a text-only character description. The source establishes the subject’s visible identity, clothing, composition, and background before audio-driven movement is synthesized.
Spoken and Sung Performance
Drive the portrait with finished speech, dialogue, narration, or singing audio. Vocal pacing, emotion, pronunciation, and delivery come from the supplied recording rather than a text-to-speech system inside this endpoint.
Long Audio Input
APIXO accepts one audio file lasting up to 600 seconds. This makes the route suitable for material longer than a conventional short avatar clip, while processing time and consistency demands increase with length.
Reproducible Iteration
Choose 480p or 720p output and optionally supply an APIXO seed. A fixed seed supports controlled experimentation, but it does not guarantee identical results across changing inputs or provider revisions.
InfiniteTalk controls through APIXO
Public limits for the dedicated portrait-and-audio workflow.
1 portrait + optional mask
Image Inputs
1 audio URL
Audio Count
MP3 / WAV / M4A
Audio Formats
128 MB
Audio Size Limit
5 seconds
Minimum Billing
Async / Callback
Delivery
Create spoken content from prepared audio
Presenter and Spokesperson Videos
Pair a prepared portrait with recorded or synthesized speech to create presenter-style content. The endpoint animates the supplied performance audio; it does not write a script or generate speech from text.
Localized Marketing Versions
Reuse an approved portrait with separately produced voice tracks for regional campaign versions. InfiniteTalk responds to the acoustic performance rather than a declared language, but pronunciation and synchronization should be reviewed for every recording.
Lessons and Training Modules
Convert narration into instructor, guide, or course-avatar videos at 480p or 720p. APIXO’s ten-minute audio limit supports longer explanations, while complex modules can still be divided into shorter segments for easier quality control.
Character Speech and Singing
Animate a suitable character portrait from dialogue, narration, or singing audio. Strongly stylized, non-human, full-body, or multi-person sources are not documented APIXO guarantees and should be tested before committing to a production format.
Prepare the portrait, audio, and delivery path
Clean source assets reduce avoidable synchronization and identity problems.
Prepare a Clear Portrait
Choose one visible subject with a clear face, limited occlusion, and sufficient image quality. Keep the intended framing and background in the source because APIXO does not expose aspect-ratio or camera controls.
Upload Finished Speech Audio
Provide one public MP3, WAV, or M4A URL within the 128 MB and 600-second limits. Complete transcription, translation, voice generation, cleanup, and timing before submitting the track.
Generate and Review Motion
Select 480p or 720p, submit the asynchronous task, and retrieve the result through polling or callback delivery. Review the mouth, teeth, eyes, hands, pose, identity, and segment transitions before publishing.
Understand what APIXO exposes
Important constraints
APIXO exposes one image-to-video mode requiring a portrait and a driving audio file.
APIXO’s published request schema marks the prompt as required; submit a clear motion instruction with every production request.
Billing uses the detected audio length or five seconds, whichever is greater.
Longer audio increases processing time; callback delivery is preferable for long production tasks.
APIXO does not publicly expose video input, aspect ratio, frame rate, or output-format controls.
Use portraits and voices only when you have the necessary rights, consent, and authorization.
InfiniteTalk questions
Not through the publicly documented infinitetalk schema. The upstream research supports audio-driven video-to-video dubbing, but APIXO exposes a portrait-and-audio image-to-video workflow without a source-video field.
No. APIXO requires an uploaded audio file, and that recording drives the resulting performance. Use a separate recording or text-to-speech system to create the spoken or sung audio before calling InfiniteTalk.
The uploaded audio drives the performance, and APIXO probes its duration for billing, up to 600 seconds. There is no separate duration control. Allow for possible processing or encoding differences when validating the final video length.
Occluded faces, extreme poses, rapid speech, unclear audio, visible hands near the face, multiple people, and low-quality portraits can increase synchronization errors or visual drift. Long outputs should be reviewed for identity and temporal consistency throughout.
Explore Other Models
Discover more AI models for your next creative workflow