APIXO
Text to SpeechMultilingual AudioCustom Voice IDs

MiniMax Speech 2.8 is MiniMax’s expressive text-to-speech generation released with Speech 2.8 HD and Turbo tiers. Through APIXO, developers can synthesize multilingual audio with preset or previously created custom voice IDs, native sound tags, emotional delivery, pronunciation overrides, and production controls for pacing and output encoding.

APIXO Modes

Speech Turbo / HD

Price

From $0.06 / 1K Chars

Languages

40

Formats

MP3 / WAV / PCM / FLAC

Released

Jan 23, 2026

Create with MiniMax Speech 2.8

Loading workspace...

Powered on APIXO

Add Expressive MiniMax Speech 2.8 to Voice Products

Convert scripts of up to 10,000 characters into speech while controlling voice, emotion, pronunciation, pace, pitch, volume, language handling, and audio encoding. APIXO combines the faster Speech 2.8 Turbo tier and higher-fidelity HD tier within one asynchronous text-to-speech route.

Capabilities

Speech Controls Built for Expressive Delivery

Native Sound Tags

Speech 2.8 introduces modeled interjections and vocal events such as breaths, chuckles, coughs, sighs, and throat clearing. These cues help written dialogue include conversational imperfections instead of uniformly polished delivery.

Emotional Performance

APIXO exposes seven emotional directions: happy, sad, angry, fearful, disgusted, surprised, and neutral. Emotion works with the selected voice and script, so teams should evaluate consistency across complete passages.

Cleaner Vocal Output

MiniMax positions Speech 2.8 around reduced background noise and fewer synthetic audio artifacts. The HD tier is intended for higher-fidelity delivery when narration clarity matters more than the lower Turbo price.

Cross-Lingual Synthesis

The model family supports speech synthesis across 40 documented languages. APIXO’s language-boost setting provides an explicit pronunciation and prosody hint, including Cantonese and a broad range of European and Asian languages.

Flexible Voice Selection

Use a documented MiniMax preset voice ID or pass a reusable custom voice ID created beforehand through MiniMax Voice. The synthesis route consumes that identity but does not create or clone a voice itself.

Pronunciation Overrides

APIXO exposes a pronunciation dictionary for terms that need a specific spoken form. This is useful for product names, abbreviations, uncommon names, and domain vocabulary that should not rely solely on automatic pronunciation.

Quality Tiers

Choose Faster Iteration or Higher-Fidelity Speech

Speech 2.8 Turbo for Efficient Iteration

APIXO maps Speech Turbo to MiniMax’s speech-2.8-turbo tier. It is the default lower-cost option for previews, frequently changing scripts, application messages, and workflows where teams need expressive synthesis without the HD rate.

$0.06 / 1K Chars

Speech 2.8 HD for Final Delivery

APIXO maps Speech HD to `speech-2.8-hd`, MiniMax’s higher-fidelity tier for selected narration, character performances, learning content, and customer-facing audio where vocal clarity and finish receive closer review. Use it for selected narration, character performances, learning content, and customer-facing audio where vocal finish receives closer review.

$0.10 / 1K Chars
Specifications

Speech 2.8 Output Controls on APIXO

Verified synthesis and encoding options for the current public route.

10,000 Characters

Text Limit

Preset / Custom ID

Voice Source

0.5–2.0

Speed

8 / 16 / 22.05 / 24 / 32 / 44.1 KHZ

Sample Rate

32 / 64 / 128 / 256 KBPS

Bitrate

Mono / Stereo

Channels

Notes & FAQ

Speech 2.8 Requirements and Voice Boundaries

Important Notes

01

APIXO requires a quality mode, valid voice ID, and non-empty synthesis text. Prompt length is capped at 10,000 characters after whitespace handling.

02

Custom voice creation belongs to the separate MiniMax Voice route; this endpoint only accepts an existing custom voice ID.

03

APIXO documents typical processing of 5–30 seconds, but latency varies with text length, tier, queue load, and route health.

04

The public route does not expose reference audio, speech-to-speech conversion, real-time streaming, subtitles, or timestamp alignment.

05

Pronunciation, code switching, emotion, and speaker character should be reviewed across full passages rather than isolated sample lines.

06

A technically valid custom voice ID does not establish consent or commercial rights to the underlying voice.

07

Expressive sound tags such as breaths, chuckles, and sighs can be embedded in the prompt text for more natural dialogue.

08

A technically valid custom voice_id does not establish consent or commercial rights to the underlying voice.

Frequently Asked Questions

APIXO exposes two MiniMax Speech 2.8 tiers: speech-2.8-turbo through Speech Turbo and speech-2.8-hd through Speech HD. It does not group Speech 2.6, Speech-02, Speech-01, or the separate MiniMax Voice workflow into this route.

No. It can synthesize speech with a compatible custom voice ID, but it does not accept reference audio or create that identity. Use the separate MiniMax Voice workflow first, then pass the resulting authorized voice ID to Speech 2.8.

APIXO charges for submitted text rather than final audio duration. Speech Turbo costs $0.06 per 1,000 prompt characters, while Speech HD costs $0.10 per 1,000 prompt characters.

APIXO exposes happy, sad, angry, fearful, disgusted, surprised, and neutral. Results depend on the voice, wording, punctuation, sound tags, and passage length, so a selected emotion should not be treated as a guaranteed performance.

The public APIXO route documents generated audio delivery but does not expose sentence timestamps, word timestamps, phoneme alignment, or subtitle controls. Applications requiring synchronized captions should add a separate alignment or transcription stage.

Explore Other Models

Discover more AI models for your next creative workflow