APIXO
Text to SpeechMultilingual AudioCustom Voice IDs

MiniMax Speech 2.8 is MiniMax’s expressive text-to-speech generation released with Speech 2.8 HD and Turbo tiers. Through APIXO, developers can synthesize multilingual audio with preset or previously created custom voice IDs, native sound tags, emotional delivery, pronunciation overrides, and production controls for pacing and output encoding.

APIXO Modes

Speech Turbo / HD

Price

From $0.06 / 1K Chars

Languages

40

Formats

MP3 / WAV / PCM / FLAC

Released

Jan 23, 2026

Create with MiniMax Speech 2.8

Loading workspace...

Powered on APIXO

Add Expressive MiniMax Speech 2.8 to Voice Products

Convert scripts of up to 10,000 characters into speech while controlling voice, emotion, pronunciation, pace, pitch, volume, language handling, and audio encoding. APIXO combines the faster Speech 2.8 Turbo tier and higher-fidelity HD tier within one asynchronous text-to-speech route.

Capabilities

Speech Controls Built for Expressive Delivery

Native Sound Tags

Speech 2.8 introduces modeled interjections and vocal events such as breaths, chuckles, coughs, sighs, and throat clearing. These cues help written dialogue include conversational imperfections instead of uniformly polished delivery.

Emotional Performance

APIXO exposes seven emotional directions: happy, sad, angry, fearful, disgusted, surprised, and neutral. Emotion works with the selected voice and script, so teams should evaluate consistency across complete passages.

Cleaner Vocal Output

MiniMax positions Speech 2.8 around reduced background noise and fewer synthetic audio artifacts. The HD tier is intended for higher-fidelity delivery when narration clarity matters more than the lower Turbo price.

Cross-Lingual Synthesis

The model family supports speech synthesis across 40 documented languages. APIXO’s language-boost setting provides an explicit pronunciation and prosody hint, including Cantonese and a broad range of European and Asian languages.

Flexible Voice Selection

Use a documented MiniMax preset voice ID or pass a reusable custom voice ID created beforehand through MiniMax Voice. The synthesis route consumes that identity but does not create or clone a voice itself.

Pronunciation Overrides

APIXO exposes a pronunciation dictionary for terms that need a specific spoken form. This is useful for product names, abbreviations, uncommon names, and domain vocabulary that should not rely solely on automatic pronunciation.

Quality Tiers

Choose Faster Iteration or Higher-Fidelity Speech

Speech 2.8 Turbo for Efficient Iteration

APIXO maps Speech Turbo to MiniMax’s speech-2.8-turbo tier. It is the default lower-cost option for previews, frequently changing scripts, application messages, and workflows where teams need expressive synthesis without the HD rate.

$0.06 / 1K Chars

Speech 2.8 HD for Final Delivery

h1Xyew

$0.10 / 1K Chars
Specifications

Speech 2.8 Output Controls on APIXO

Verified synthesis and encoding options for the current public route.

10,000 Characters

Text Limit

Preset / Custom ID

Voice Source

0.5–2.0

Speed

8 / 16 / 22.05 / 24 / 32 / 44.1 KHZ

Sample Rate

32 / 64 / 128 / 256 KBPS

Bitrate

Mono / Stereo

Channels

Use Cases

Production Workflows for MiniMax Speech 2.8

Product Voice

Guided App Experiences

Generate onboarding narration, assistant responses, help prompts, and spoken notifications using a tested preset or custom voice ID. Turbo supports rapid script iteration, while pronunciation overrides help keep product terminology consistent.

Long-Form Media

Narration and Audiobooks

Convert chapters, articles, explainers, or podcast scripts into controlled narration. Break longer works into reviewed sections so pacing, pronunciation, emotion, and voice character remain appropriate across the complete production.

Localization

Multilingual Campaign Audio

Synthesize localized ads, lessons, product tours, and digital-presenter tracks with an explicit language hint. Review names, code-switched passages, regional pronunciation, and emotional tone with native speakers before release.

Interactive Media

Character and Game Dialogue

Apply sound tags and emotional controls to character lines, branching dialogue, educational scenarios, or accessibility narration. Custom voice IDs can support recurring characters when they were created separately with authorized source material.

Notes & FAQ

Speech 2.8 Requirements and Voice Boundaries

Important Notes

01

APIXO requires a quality mode, valid voice ID, and non-empty synthesis text. Prompt length is capped at 10,000 characters after whitespace handling.

02

Custom voice creation belongs to the separate MiniMax Voice route; this endpoint only accepts an existing custom voice ID.

03

APIXO documents typical processing of 5–30 seconds, but latency varies with text length, tier, queue load, and route health.

04

The public route does not expose reference audio, speech-to-speech conversion, real-time streaming, subtitles, or timestamp alignment.

05

Pronunciation, code switching, emotion, and speaker character should be reviewed across full passages rather than isolated sample lines.

06

A technically valid custom voice ID does not establish consent or commercial rights to the underlying voice.

07

Expressive sound tags such as breaths, chuckles, and sighs can be embedded in the prompt text for more natural dialogue.

08

A technically valid custom voice_id does not establish consent or commercial rights to the underlying voice.

Frequently Asked Questions

APIXO exposes two MiniMax Speech 2.8 tiers: speech-2.8-turbo through Speech Turbo and speech-2.8-hd through Speech HD. It does not group Speech 2.6, Speech-02, Speech-01, or the separate MiniMax Voice workflow into this route.

No. It can synthesize speech with a compatible custom voice ID, but it does not accept reference audio or create that identity. Use the separate MiniMax Voice workflow first, then pass the resulting authorized voice ID to Speech 2.8.

APIXO charges for submitted text rather than final audio duration. Speech Turbo costs $0.06 per 1,000 prompt characters, while Speech HD costs $0.10 per 1,000 prompt characters.

APIXO exposes happy, sad, angry, fearful, disgusted, surprised, and neutral. Results depend on the voice, wording, punctuation, sound tags, and passage length, so a selected emotion should not be treated as a guaranteed performance.

The public APIXO route documents generated audio delivery but does not expose sentence timestamps, word timestamps, phoneme alignment, or subtitle controls. Applications requiring synchronized captions should add a separate alignment or transcription stage.

Explore Other Models

Discover more AI models for your next creative workflow