APIXO exposes two MiniMax Speech 2.8 tiers: speech-2.8-turbo through Speech Turbo and speech-2.8-hd through Speech HD. It does not group Speech 2.6, Speech-02, Speech-01, or the separate MiniMax Voice workflow into this route.
MiniMax Speech 2.8
MiniMax Speech 2.8 is MiniMax’s expressive text-to-speech generation released with Speech 2.8 HD and Turbo tiers. Through APIXO, developers can synthesize multilingual audio with preset or previously created custom voice IDs, native sound tags, emotional delivery, pronunciation overrides, and production controls for pacing and output encoding.
APIXO Modes
Speech Turbo / HD
Price
From $0.06 / 1K Chars
Languages
40
Formats
MP3 / WAV / PCM / FLAC
Released
Jan 23, 2026
Create with MiniMax Speech 2.8
Loading workspace...
Speech Controls Built for Expressive Delivery
Native Sound Tags
Speech 2.8 introduces modeled interjections and vocal events such as breaths, chuckles, coughs, sighs, and throat clearing. These cues help written dialogue include conversational imperfections instead of uniformly polished delivery.
Emotional Performance
APIXO exposes seven emotional directions: happy, sad, angry, fearful, disgusted, surprised, and neutral. Emotion works with the selected voice and script, so teams should evaluate consistency across complete passages.
Cleaner Vocal Output
MiniMax positions Speech 2.8 around reduced background noise and fewer synthetic audio artifacts. The HD tier is intended for higher-fidelity delivery when narration clarity matters more than the lower Turbo price.
Cross-Lingual Synthesis
The model family supports speech synthesis across 40 documented languages. APIXO’s language-boost setting provides an explicit pronunciation and prosody hint, including Cantonese and a broad range of European and Asian languages.
Flexible Voice Selection
Use a documented MiniMax preset voice ID or pass a reusable custom voice ID created beforehand through MiniMax Voice. The synthesis route consumes that identity but does not create or clone a voice itself.
Pronunciation Overrides
APIXO exposes a pronunciation dictionary for terms that need a specific spoken form. This is useful for product names, abbreviations, uncommon names, and domain vocabulary that should not rely solely on automatic pronunciation.
Choose Faster Iteration or Higher-Fidelity Speech
Speech 2.8 Turbo for Efficient Iteration
APIXO maps Speech Turbo to MiniMax’s speech-2.8-turbo tier. It is the default lower-cost option for previews, frequently changing scripts, application messages, and workflows where teams need expressive synthesis without the HD rate.
Speech 2.8 HD for Final Delivery
h1Xyew
Speech 2.8 Output Controls on APIXO
Verified synthesis and encoding options for the current public route.
10,000 Characters
Text Limit
Preset / Custom ID
Voice Source
0.5–2.0
Speed
8 / 16 / 22.05 / 24 / 32 / 44.1 KHZ
Sample Rate
32 / 64 / 128 / 256 KBPS
Bitrate
Mono / Stereo
Channels
Production Workflows for MiniMax Speech 2.8
Guided App Experiences
Generate onboarding narration, assistant responses, help prompts, and spoken notifications using a tested preset or custom voice ID. Turbo supports rapid script iteration, while pronunciation overrides help keep product terminology consistent.
Narration and Audiobooks
Convert chapters, articles, explainers, or podcast scripts into controlled narration. Break longer works into reviewed sections so pacing, pronunciation, emotion, and voice character remain appropriate across the complete production.
Multilingual Campaign Audio
Synthesize localized ads, lessons, product tours, and digital-presenter tracks with an explicit language hint. Review names, code-switched passages, regional pronunciation, and emotional tone with native speakers before release.
Character and Game Dialogue
Apply sound tags and emotional controls to character lines, branching dialogue, educational scenarios, or accessibility narration. Custom voice IDs can support recurring characters when they were created separately with authorized source material.
Speech 2.8 Requirements and Voice Boundaries
Important Notes
APIXO requires a quality mode, valid voice ID, and non-empty synthesis text. Prompt length is capped at 10,000 characters after whitespace handling.
Custom voice creation belongs to the separate MiniMax Voice route; this endpoint only accepts an existing custom voice ID.
APIXO documents typical processing of 5–30 seconds, but latency varies with text length, tier, queue load, and route health.
The public route does not expose reference audio, speech-to-speech conversion, real-time streaming, subtitles, or timestamp alignment.
Pronunciation, code switching, emotion, and speaker character should be reviewed across full passages rather than isolated sample lines.
A technically valid custom voice ID does not establish consent or commercial rights to the underlying voice.
Expressive sound tags such as breaths, chuckles, and sighs can be embedded in the prompt text for more natural dialogue.
A technically valid custom voice_id does not establish consent or commercial rights to the underlying voice.
Frequently Asked Questions
No. It can synthesize speech with a compatible custom voice ID, but it does not accept reference audio or create that identity. Use the separate MiniMax Voice workflow first, then pass the resulting authorized voice ID to Speech 2.8.
APIXO charges for submitted text rather than final audio duration. Speech Turbo costs $0.06 per 1,000 prompt characters, while Speech HD costs $0.10 per 1,000 prompt characters.
APIXO exposes happy, sad, angry, fearful, disgusted, surprised, and neutral. Results depend on the voice, wording, punctuation, sound tags, and passage length, so a selected emotion should not be treated as a guaranteed performance.
The public APIXO route documents generated audio delivery but does not expose sentence timestamps, word timestamps, phoneme alignment, or subtitle controls. Applications requiring synchronized captions should add a separate alignment or transcription stage.
Explore Other Models
Discover more AI models for your next creative workflow
