MiMo-V2.5-TTS

Premium voice synthesis with fine-grained emotion control.

Style instructions from one-liner to screenplay, and zero-prompt character inference.

Model Specs

Modality

InputText
OutputAudio

Capabilities

Speech Synthesis
Streaming Output

Performance

Context Length8K tokens
Max Output8K tokens
RPM100
TPM10M

Model Pricing

Price
Free (limited time)

Model Advantages

Premium Voice Library

A curated library of high-quality voices spanning genders, ages, and styles — ready to use out of the box for audiobooks, podcasts, dubbing, and more.

Style Instruction Following

From a single-line style cue to a full screenplay-style layered script, the model delivers faithfully. Fine-grained control over pacing, emotion, tone, and breath; complex character delivery lands consistently.

Inline Emotion Tag Control

Embed audio tags directly in text to control emotion word by word. Transitions between emotional states are natural and reliable, ideal for audiobook production and voice-acting workflows.

Zero-Prompt Text Comprehension

No style instructions needed — the model infers emotional arc and character voice directly from the content. In multi-character dialogue, it identifies and switches between distinct voices naturally.

Real-World Performance

Style Instruction: Weathered Elder

Instruct

声音低沉沙哑一点,像个历经沧桑的老前辈在讲述传奇人物。语气里带点由衷的敬佩,娓娓道来。

Text

街口那个老周啊,媳妇走得早,一个人拉扯俩娃,白天蹬三轮,晚上还去夜市摆摊修鞋。现在俩孩子都有出息喽,想接他去城里享福——他不去,就守着那间小铺子。哎,人哪,骨头硬,心里头就踏实。

Audio
0:00

Screenplay Script: Divine Monologue

Text

你们求我垂怜,求我降下甘霖洗净这浊世。可这世间的沉疴,唯有烈火能剔骨刮毒。闭上眼吧。这业火烧起来的时候,一点也不疼。

Audio
0:00

Inline Tags: Funeral Eulogy

Text

[crying] She's gone... she's really gone...[pause] but you know what's funny? [sniffles] She always said she'd outlive us all. [crying] God, I miss her so much.

Audio
0:00

Text Comprehension: Multi-Character Dialogue

Text

The five-year-old squealed, "Look, Grandpa! A PUPPY!" The old man squinted and grumbled, "That ain't a puppy, that's a raccoon." The teenager rolled her eyes: "It's OBVIOUSLY a cat, you're both blind." The police officer stepped forward: "Ma'am, sir, I'm going to need everyone to step back slowly." The little boy whimpered, "Is it gonna bite me?"

Audio
0:00

Choose Your Access Method

Pay-As-You-Go API

1

Get Your API Key

Create an account in the console and generate your API key (TTS series currently free).

2

Sample Code

Pass style instructions and text content via messages, specify audio format and voice to call the API.

import os
from openai import OpenAI
import base64

client = OpenAI(
    api_key=os.environ.get("MIMO_API_KEY"),
    base_url="https://api.xiaomimimo.com/v1"
)

completion = client.chat.completions.create(
    model="mimo-v2.5-tts",
    messages=[
        {
            "role": "user",
            "content": "Bright, bouncy, slightly sing-song tone — like you're bursting with good news you can barely hold in. Fast pace, rising pitch at the end."
        },
        {
            "role": "assistant",
            "content": "Hey boss — guess what, guess what? I just got the results back and I actually passed! Not just passed, I got a distinction! I know, I know — you told me I was cutting it close, but hey, here we are. Drinks are on me tonight, okay?"
        }
    ],
    audio={
        "format": "wav",
        "voice": "Chloe"
    }
)

message = completion.choices[0].message
audio_bytes = base64.b64decode(message.audio.data)
with open("audio_file.wav", "wb") as f:
    f.write(audio_bytes)

Token Plan Subscription

1

Subscribe

Monthly or annual plans covering mimo-v2.6-pro, mimo-v2.6-flash and the full MiMo V2.5 model suite — significantly better value than pay-as-you-go at high usage volumes.

2

Get Token Plan API Key

After subscribing, generate a dedicated API key from the Token Plan page — separate from your pay-as-you-go key.

3

Quick Integration

Same API format as pay-as-you-go — just swap the key and base URL. Works directly with Claude Code, Cline, and other popular AI tools.

Use in MiMo Claw

MiMo Claw is equipped with MiMo's latest flagship model, and supports both the latest multimodal understanding and voice series large language models. It can be added with one click for subscription across all tiers of TokenPlan, and is currently available for a limited-time free trial.

Copyright©2026 Xiaomi. All Rights Reserved | Cookie Policy | Cookie Preferences

We use cookies and similar technologies of our own to ensure the proper functioning of the website, customize content according to user preferences and analyze users' interactions on the website, as well as their browsing habits. You can find more information in our Cookie Policy. Select an option or go to Cookie Settings to manage your preferences. Learn More.