MiMo-V2.5-TTS-VoiceDesign

Generate any voice from a text description alone — no reference audio needed.

Multi-dimensional customization across accent, age, timbre, and recording texture.

Model Specs

Modality

InputText
OutputAudio

Capabilities

Voice Design
Speech Synthesis
Streaming Output

Performance

Context Length8K tokens
Max Output8K tokens
RPM100
TPM10M

Model Pricing

Price
Free (limited time)

Model Advantages

Voice Generation from Text

Describe a voice in natural language and the model generates it from scratch — no reference audio required.

Multi-Dimensional Customization

Describe from any angle: age, gender, accent, timbre, pacing, temperament, or recording texture — the full spectrum without specialized audio knowledge.

Complex Description Handling

Strong comprehension of complex, ambiguous, or layered descriptions. Handles both unique unheard voices and precise reproductions of classic character archetypes.

Multi-Language Coverage

Voice design works across Chinese, English, and varied accents — Mandarin or dialect voices, different national accents, and more — with flexible range across languages and styles.

Real-World Performance

Russian Accent

Instruct

Heavy Russian accent, gruff middle-aged male, blunt and matter-of-fact.

Text

You want my opinion? Fine. This plan will not work. I have seen many plan like this before, all fail.

Audio
0:00

Binaural ASMR

Instruct

Young female, extreme close-up with a binaural, ear-to-ear ASMR feel.

Text

[Whispering] Shhh... just relax, come a little closer. I'm right here beside you now.

Audio
0:00

Documentary Narrator

Instruct

一位中年男性,说标准普通话,嗓音低沉有磁性,带有轻微的沙哑质感,像纪录片旁白解说员。

Text

当最后一缕阳光消失在地平线之下,这片沉睡了亿万年的大地开始显露它真正的面貌。

Audio
0:00

Elderly Narrator

Instruct

一位年迈的老先生,说带北方口音的普通话,语速缓慢而沉稳,嗓音略带沙哑和沧桑感。

Text

我这辈子啊,走南闯北六十多年。到头来才明白一个道理——这人哪,不在走了多远的路,在于记住了多少风景。

Audio
0:00

Choose Your Access Method

Pay-As-You-Go API

1

Get Your API Key

Create an account and generate your API key (TTS series currently free).

2

Sample Code

Pass a voice description and text content via messages.

import os, base64
from openai import OpenAI

client = OpenAI(
    api_key=os.environ.get("MIMO_API_KEY"),
    base_url="https://api.xiaomimimo.com/v1",
)

completion = client.chat.completions.create(
    model="mimo-v2.5-tts-voicedesign",
    messages=[
        {"role": "user", "content": "Give me a young male tone."},
        {"role": "assistant", "content": "Yes, I had a sandwich."}
    ],
    audio={"format": "wav", "optimize_text_preview": True}
)

message = completion.choices[0].message
audio_bytes = base64.b64decode(message.audio.data)
with open("audio_file.wav", "wb") as f:
    f.write(audio_bytes)

Token Plan Subscription

1

Subscribe

Monthly or annual plans covering mimo-v2.6-pro, mimo-v2.6-flash and the full MiMo V2.5 model suite — significantly better value than pay-as-you-go at high usage volumes.

2

Get Token Plan API Key

After subscribing, generate a dedicated API key from the Token Plan page — separate from your pay-as-you-go key.

3

Quick Integration

Same API format as pay-as-you-go — just swap the key and base URL. Works directly with Claude Code, Cline, and other popular AI tools.

Use in MiMo Claw

MiMo Claw is equipped with MiMo's latest flagship model, and supports both the latest multimodal understanding and voice series large language models. It can be added with one click for subscription across all tiers of TokenPlan, and is currently available for a limited-time free trial.

Copyright©2026 Xiaomi. All Rights Reserved | Cookie Policy | Cookie Preferences

We use cookies and similar technologies of our own to ensure the proper functioning of the website, customize content according to user preferences and analyze users' interactions on the website, as well as their browsing habits. You can find more information in our Cookie Policy. Select an option or go to Cookie Settings to manage your preferences. Learn More.