MiMo-V2.5-TTS-VoiceClone
Clone any voice from a short audio sample with high fidelity — no training, no labeling.
Style instructions and emotion tags still work after cloning.
Model Specs
Modality
Capabilities
Performance
Model Pricing
Model Advantages
High-Fidelity Voice Cloning
A few seconds of reference audio — faithfully reproduces the target voice, capturing timbre, breath patterns, pause habits, and prosodic rhythm.
Zero Training, Instant Results
No fine-tuning or data annotation needed. Provide the reference audio and get the cloned voice immediately.
Full Control Stack Inherited
The cloned voice inherits full control capabilities — style instructions and inline audio tags both work on top of the clone.
Cross-Language Cloning
Chinese and English reference audio both supported. Language switching does not compromise timbre reproduction accuracy.
Real-World Performance
Chinese Clone: Direct Transfer
Text
风轻轻地吹过,带来了远方的花香,和记忆里那个夏天的味道。
Clone + Style Override
Instruct
用尖锐刻薄的嗓音,带着狐假虎威的得意感说话,营造压迫感。
Text
你以为我是谁,也敢在这儿跟我耍横?我告诉你,站在我身后的那个人,说出来吓死你。
English Clone: Cyberpunk Narration
Text
Ignore the sirens, ignore the neon bleeding through your eyelids, and just breathe with me.
English Clone + Style: Post-Match Pundit
Instruct
Broadcast this like a blistering post-match pundit tearing into a disastrous performance.
Text
No shape, no urgency, no clue what they're trying to do out there.
Choose Your Access Method
Pay-As-You-Go API
Sample Code
Upload a reference audio clip and pass text content to clone the target voice.
import os
import base64
import urllib.request
from openai import OpenAI
client = OpenAI(
api_key=os.environ.get("MIMO_API_KEY"),
base_url="https://api.xiaomimimo.com/v1",
)
# Example audio URL
audio_url = "https://example-files.cnbj1.mi-fds.com/example-files/audio/audio_example.wav"
audio_file = "audio_example.wav"
# Download the audio file from URL (skip if already exists)
if not os.path.exists(audio_file):
urllib.request.urlretrieve(audio_url, audio_file)
# To use a local file directly, replace the above with:
# audio_file = "your_local_audio.wav"
with open(audio_file, "rb") as f:
audio_bytes = f.read()
voice_base64 = base64.b64encode(audio_bytes).decode("utf-8")
completion = client.chat.completions.create(
model="mimo-v2.5-tts-voiceclone",
messages=[
{
"role": "user",
"content": ""
},
{
"role": "assistant",
"content": "Yes, I had a sandwich."
}
],
audio={
"format": "wav",
"voice": f"data:audio/wav;base64,{voice_base64}"
}
)
message = completion.choices[0].message
audio_bytes = base64.b64decode(message.audio.data)
with open("audio_file.wav", "wb") as f:
f.write(audio_bytes)Token Plan Subscription
Subscribe
Monthly or annual plans covering mimo-v2.6-pro, mimo-v2.6-flash and the full MiMo V2.5 model suite — significantly better value than pay-as-you-go at high usage volumes.
Get Token Plan API Key
After subscribing, generate a dedicated API key from the Token Plan page — separate from your pay-as-you-go key.
Quick Integration
Same API format as pay-as-you-go — just swap the key and base URL. Works directly with Claude Code, Cline, and other popular AI tools.
Use in MiMo Claw
MiMo Claw is equipped with MiMo's latest flagship model, and supports both the latest multimodal understanding and voice series large language models. It can be added with one click for subscription across all tiers of TokenPlan, and is currently available for a limited-time free trial.