APIXO exposes two MiniMax Speech 2.8 tiers: speech-2.8-turbo through Speech Turbo and speech-2.8-hd through Speech HD. It does not group Speech 2.6, Speech-02, Speech-01, or the separate MiniMax Voice workflow into this route.
MiniMax Speech 2.8
MiniMax Speech 2.8 is MiniMax’s expressive text-to-speech generation released with Speech 2.8 HD and Turbo tiers. Through APIXO, developers can synthesize multilingual audio with preset or previously created custom voice IDs, native sound tags, emotional delivery, pronunciation overrides, and production controls for pacing and output encoding.
APIXO Modes
Speech Turbo / HD
Price
From $0.06 / 1K Chars
Languages
40
Formats
MP3 / WAV / PCM / FLAC
Released
Jan 23, 2026
Create with MiniMax Speech 2.8
Loading workspace...
Speech Controls Built for Expressive Delivery
Native Sound Tags
Speech 2.8 introduces modeled interjections and vocal events such as breaths, chuckles, coughs, sighs, and throat clearing. These cues help written dialogue include conversational imperfections instead of uniformly polished delivery.
Emotional Performance
APIXO exposes seven emotional directions: happy, sad, angry, fearful, disgusted, surprised, and neutral. Emotion works with the selected voice and script, so teams should evaluate consistency across complete passages.
Cleaner Vocal Output
MiniMax positions Speech 2.8 around reduced background noise and fewer synthetic audio artifacts. The HD tier is intended for higher-fidelity delivery when narration clarity matters more than the lower Turbo price.
Cross-Lingual Synthesis
The model family supports speech synthesis across 40 documented languages. APIXO’s language-boost setting provides an explicit pronunciation and prosody hint, including Cantonese and a broad range of European and Asian languages.
Flexible Voice Selection
Use a documented MiniMax preset voice ID or pass a reusable custom voice ID created beforehand through MiniMax Voice. The synthesis route consumes that identity but does not create or clone a voice itself.
Pronunciation Overrides
APIXO exposes a pronunciation dictionary for terms that need a specific spoken form. This is useful for product names, abbreviations, uncommon names, and domain vocabulary that should not rely solely on automatic pronunciation.
Choose Faster Iteration or Higher-Fidelity Speech
Speech 2.8 Turbo for Efficient Iteration
APIXO maps Speech Turbo to MiniMax’s speech-2.8-turbo tier. It is the default lower-cost option for previews, frequently changing scripts, application messages, and workflows where teams need expressive synthesis without the HD rate.
Speech 2.8 HD for Final Delivery
APIXO maps Speech HD to `speech-2.8-hd`, MiniMax’s higher-fidelity tier for selected narration, character performances, learning content, and customer-facing audio where vocal clarity and finish receive closer review. Use it for selected narration, character performances, learning content, and customer-facing audio where vocal finish receives closer review.
Speech 2.8 Output Controls on APIXO
Verified synthesis and encoding options for the current public route.
10,000 Characters
Text Limit
Preset / Custom ID
Voice Source
0.5–2.0
Speed
8 / 16 / 22.05 / 24 / 32 / 44.1 KHZ
Sample Rate
32 / 64 / 128 / 256 KBPS
Bitrate
Mono / Stereo
Channels
Speech 2.8 Requirements and Voice Boundaries
Important Notes
APIXO requires a quality mode, valid voice ID, and non-empty synthesis text. Prompt length is capped at 10,000 characters after whitespace handling.
Custom voice creation belongs to the separate MiniMax Voice route; this endpoint only accepts an existing custom voice ID.
APIXO documents typical processing of 5–30 seconds, but latency varies with text length, tier, queue load, and route health.
The public route does not expose reference audio, speech-to-speech conversion, real-time streaming, subtitles, or timestamp alignment.
Pronunciation, code switching, emotion, and speaker character should be reviewed across full passages rather than isolated sample lines.
A technically valid custom voice ID does not establish consent or commercial rights to the underlying voice.
Expressive sound tags such as breaths, chuckles, and sighs can be embedded in the prompt text for more natural dialogue.
A technically valid custom voice_id does not establish consent or commercial rights to the underlying voice.
Frequently Asked Questions
No. It can synthesize speech with a compatible custom voice ID, but it does not accept reference audio or create that identity. Use the separate MiniMax Voice workflow first, then pass the resulting authorized voice ID to Speech 2.8.
APIXO charges for submitted text rather than final audio duration. Speech Turbo costs $0.06 per 1,000 prompt characters, while Speech HD costs $0.10 per 1,000 prompt characters.
APIXO exposes happy, sad, angry, fearful, disgusted, surprised, and neutral. Results depend on the voice, wording, punctuation, sound tags, and passage length, so a selected emotion should not be treated as a guaranteed performance.
The public APIXO route documents generated audio delivery but does not expose sentence timestamps, word timestamps, phoneme alignment, or subtitle controls. Applications requiring synchronized captions should add a separate alignment or transcription stage.
Explore Other Models
Discover more AI models for your next creative workflow
