# Xiaomi MiMo API Open Platform > Xiaomi MiMo API Open Platform provides high-performance inference services for Xiaomi's AI models, compatible with OpenAI and Anthropic API formats. This platform offers comprehensive API documentation, integration guides, and detailed update logs for Xiaomi-related AI models. It is designed to empower developers to build and deploy next-generation intelligent applications and agents with ease. --- DOCUMENT: First API Call --- URL: https://mimo.mi.com/static/docs/quick-start/summary/first-api-call.md # First API Call ## Supported API Types {/* feishu-style:text-align:left */} Xiaomi MiMo API Open Platform is compatible with OpenAI API and Anthropic API formats. You can use existing SDKs to access model inference services. ## Preparation Before Calling ### Log in to Xiaomi MiMo API Open Platform {/* feishu-style:text-align:left */} Currently, the platform only provides personal account login. You need to use a Xiaomi account to log in. If you already have a Xiaomi account, you can log in directly. If you don't have a Xiaomi account, you can visit the [Console](https://platform.xiaomimimo.com/#/console/usage) to register, or register in advance at [id.mi.com](https://id.mi.com/). ### Obtain Credentials {/* feishu-style:text-align:left */} Supports two usage methods, but the corresponding credential acquisition methods are different:
Usage Method Description Acquisition Method (BASE_URL and API Key below are examples)
Pay-as-you-go MiMo API Charged based on actual usage, suitable for light use Real-time Inference
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://api.xiaomimimo.com/v1`
    • Anthropic Compatibility Protocol: `https://api.xiaomimimo.com/anthropic`
  • API Key
    • Format: `sk-xxxxx`

Go to [API Keys](https://platform.xiaomimimo.com/#/console/api-keys) to create an API Key
Batch Inference
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://batch-api-cn.xiaomimimo.com/v1`
  • API Key
    • Format: `sk-xxxxx`

Go to [**Batch Inference**](https://platform.xiaomimimo.com/console/batch) to get your dedicated Base URL.
Token Plan Fixed subscription fee, with limited calls based on the package
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/v1`
    • Anthropic Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/anthropic`
  • API Key
    • Format (Individual): `tp-xxxxx`
    • Format (Team): `ttp-xxxxx`

After successful subscription, go to [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) to obtain the exclusive Base URL and API Key
> Please keep your API Key safe to avoid leakage that may result in quota theft. It is recommended to configure the API Key in environment variables. ## Quick Integration Examples {/* feishu-style:text-align:left */} You can copy the following API example code and replace the API Key value to quickly make calls. If you use the Token Plan, you need to replace the BASE_URL and use the dedicated API Key. {/* feishu-style:text-align:left */} The following system prompts are highly recommended, please choose from English and Chinese version. > Chinese version > > ```json > 你是MiMo(中文名称也是MiMo),是小米公司研发的AI智能助手。 > 今天的日期:{date} {week},你的知识截止日期是2024年12月。 > ``` > English version > > ```json > You are MiMo, an AI assistant developed by Xiaomi. > Today's date: {date} {week}. Your knowledge cutoff date is December 2024. > ``` ### OpenAI Chat Completions API Compatibility {/* feishu-style:text-align:left */} MiMo models are compatible with the OpenAI Chat Completions API and support invocation via the Python SDK and Curl. Integration examples are provided below. {/* feishu-style:text-align:left */} 1. Install the OpenAI Python SDK by running the following command: ```shell # If the run fails, you can replace pip with pip3 and run again pip install -U openai ``` {/* feishu-style:text-align:left */} 2. Call the API: ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": "please introduce yourself" } ], max_completion_tokens=1024, temperature=1.0, top_p=0.95, stream=False, stop=None, frequency_penalty=0, presence_penalty=0 ) print(completion.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": "please introduce yourself" } ], "max_completion_tokens": 1024, "temperature": 1.0, "top_p": 0.95, "stream": false, "stop": null, "frequency_penalty": 0, "presence_penalty": 0 }' ``` ### Anthropic Messages API Compatibility {/* feishu-style:text-align:left */} MiMo models are compatible with the Anthropic Messages API and support invocation via the Python SDK and Curl. Integration examples are provided below. {/* feishu-style:text-align:left */} 1. Install the Anthropic Python SDK by running the following command: ```shell # If the run fails, you can replace pip with pip3 and run again pip install -U anthropic ``` {/* feishu-style:text-align:left */} 2. Call the API: ```python import os from anthropic import Anthropic client = Anthropic( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/anthropic" ) message = client.messages.create( model="mimo-v2.6-pro", max_tokens=1024, system="You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.", messages=[ { "role": "user", "content": [ { "type": "text", "text": "please introduce yourself" } ] } ], top_p=0.95, stream=False, temperature=1.0, stop_sequences=None ) print(message.content) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/anthropic/v1/messages' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "max_tokens": 1024, "system": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "please introduce yourself" } ] } ], "top_p": 0.95, "stream": false, "temperature": 1.0, "stop_sequences": null }' ``` ### Make Multi-turn Tool Calls in Thinking Mode {/* feishu-style:text-align:left */} During the multi-turn tool calls process in thinking mode, the model returns a `reasoning_content` field alongside `tool_calls`. To continue the conversation, it is recommended to keep all previous `reasoning_content` in the `messages` array for each subsequent request to achieve the best performance. {/* feishu-style:text-align:left */} The requested example is as follows: ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "messages": [ { "role": "assistant", "content": "Hello! I am MiMo.", "reasoning_content": "Okay, the user just asked me to introduce myself. That is a pretty straightforward request, but I should think about why they are asking this." }, { "role": "user", "content": "What is the weather like in Hebei?" } ], "model": "mimo-v2.6-pro", "max_completion_tokens": 1024, "temperature": 1.0, "stream": false, "tools": [ { "type": "function", "function": { "name": "get_current_weather", "description": "Get the current weather in a given location", "parameters": { "type": "object", "properties": { "location": { "type": "string", "description": "The city and state, e.g. San Francisco, CA" }, "unit": { "type": "string", "enum": [ "celsius", "fahrenheit" ] } }, "required": [ "location" ] } } } ], "tool_choice": "auto" }' ``` ## Check Usage Information {/* feishu-style:text-align:left */} On the [Usage Information](https://platform.xiaomimimo.com/#/console/usage) page, you can view and export detailed data of your account's model Token usage and request counts by date. --- DOCUMENT: Models --- URL: https://mimo.mi.com/static/docs/quick-start/summary/model.md # Models {/* feishu-style:text-align:left */} This page lists all the models currently supported by the Xiaomi MiMo API Open Platform, including model capabilities, length limits, and rate-limiting quotas, to help you select the appropriate model based on your usage scenario. ### Rate Limiting Instructions {/* feishu-style:text-align:left */} The platform sets a model concurrency limit for each account. When the server load is high, response delays or ` 429 ` error may occur. We recommend that you reasonably plan your request frequency and implement request retry and backoff strategies in high-concurrency scenarios to avoid triggering rate limits.
- **RPM (Requests Per Minute)** : The maximum number of requests initiated per minute. The calculation scope is the sum of the total number of requests from all API Keys under a single account when calling the same model. - **TPM (Tokens Per Minute)** : The maximum number of Tokens that can be interacted with per minute. The calculation scope is the sum of the total number of requested Tokens for all API Keys under a single account when calling the same model.
{/* feishu-style:text-align:left */}

### Text Generation Model
- `mimo-v2.5-pro` and `mimo-v2.5` **will be officially deprecated at 10:00 (Beijing time) on October 21,2026. It is recommended to switch to the new version of the models as soon as possible. For details, please refer to** [**Model Deprecation**](https://mimo.mi.com/docs/zh-CN/updates/deprecate)**.**
**模型 ID (Model ID)** **能力支持** **长度限制(token)** **限流**
`mimo-v2.6-pro`
`mimo-v2.6-flash`
`mimo-v2.5`(to be deprecated)
  • **Full-modal Understanding**
  • Text Generation
  • Deep Thinking
  • Streaming Output
  • Function Call
  • Structured Output
  • Web Search
Context Window: 1M
Maximum Output: 128K
Maximum RPM: 100
Maximum TPM: 10M
`mimo-v2.6-pro-ultraspeed` Context Window: 1M
Maximum Output: 128K
Customized services available, please [contact us](https://platform.xiaomimimo.com/contact?userId=2451887661)
`mimo-v2.5-pro`(to be deprecated)
  • Text Generation
  • Full-modal Understanding
  • Deep Thinking
  • Streaming Output
  • Function Call
  • Structured Output
  • Web Search
Context Window: 1M
Maximum Output: 128K
Maximum RPM: 100
Maximum TPM: 10M
{/* feishu-style:text-align:left */}

### Automatic Speech Recognition (ASR) Model
**Model ID** **Capability Support** **Length Limit (token)** **Rate Limiting**
`mimo-v2.5-asr` Speech Recognition Context Window: 8k
Maximum Output: 2k
Maximum RPM: 100
Maximum TPM: 10k
{/* feishu-style:text-align:left */}

### Text-to-Speech (TTS) Model
**Model ID (Model ID)** **Capability Support** **Length Limit (token)** **Rate Limiting**
`mimo-v2.5-tts` Speech Synthesis Context Window: 8K
Maximum Output: 8K
Maximum RPM: 100
Maximum TPM: 10M
`mimo-v2.5-tts-voiceclone` Speech Synthesis
Timbre Cloning
`mimo-v2.5-tts-voicedesign` Speech Synthesis
Timbre Design
{/* feishu-style:text-align:left */}

### Quick Selection Guide
Requirement Scenario Recommendation Model
Complex projects, long-term tasks, high-value work, cybersecurity and scientific research requirements `mimo-v2.6-pro`
Frequent Calls and Large-Scale Tasks in Professional Office Scenarios `mimo-v2.6-flash`
Production scenarios with strong real-time interaction and sensitivity to response speed `mimo-v2.6-pro-ultraspeed`
Speech-to-Text (supports Chinese and English) `mimo-v2.5-asr`
Text-to-Speech (Standard Preset Voice) `mimo-v2.5-tts`
Voice Cloning (Upload Audio Sample) `mimo-v2.5-tts-voiceclone`
Custom Timbre Design `mimo-v2.5-tts-voicedesign`
--- DOCUMENT: Web Search --- URL: https://mimo.mi.com/static/docs/quick-start/usage-guide/text-generation/tool-calling/web-search.md # Web Search {/* feishu-style:text-align:left */} Web Search is a basic online search tool that helps your large model obtain real-time public online information (such as news, products, weather, etc.). {/* feishu-style:text-align:left */} **Core Capabilities** - **Flexible search modes**: Supports forced search and intent recognition. With intent recognition enabled, the model will autonomously decide whether to perform an online search without manual triggering. - **Early search source return**: In the streaming response, the first packet will return all search sources. - **Hybrid multi-tool invocation**: Can work with custom functions and tools; the model will automatically determine invocation priority and necessity. - **Flexible response modes**: Supports both streaming and non-streaming responses, and both methods will return search and summary content. ## Quick Start
The [Web Search Plugin](https://platform.xiaomimimo.com/#/console/plugin) must be activated prior to use.
### Enable the Service 1. Go to [Console → Plugin Management](https://platform.xiaomimimo.com/#/console/plugin), and activate the Web Search Plugin. 1. Web Search Plugin fee, refer to the [Pricing](https://platform.xiaomimimo.com/#/docs/pricing). Note: **Search invocation is determined by the model. A single search round (if required by the model) may initiate multiple keywords concurrently, resulting in multiple invocations of the Internet Content Plugin. You may use the** `max_keyword` **parameter to limit the maximum number of keywords per search round, thereby further controlling invocation frequency and costs**. {/* feishu-style:text-align:left */} **Note**:For preparations such as obtaining an API Key, please refer to [First API Call](https://platform.xiaomimimo.com/#/docs/quick-start/first-api-call). ### Sample Code ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": "武汉明天天气怎么样?" } ], max_completion_tokens=1024, temperature=1.0, top_p=0.95, stream=False, stop=None, frequency_penalty=0, presence_penalty=0, extra_body={ "thinking": {"type": "disabled"} }, tools=[ { "type": "web_search", "max_keyword": 3, "force_search": True, "limit": 1, "user_location": { "type": "approximate", "country": "China", "region": "Hubei", "city": "Wuhan" } } ], tool_choice="auto" ) print(completion.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "user", "content": "武汉明天天气怎么样?" } ], "tools": [ { "type": "web_search", "max_keyword": 3, "force_search": true, "limit": 1, "user_location": { "type": "approximate", "country": "China", "region": "Hubei", "city": "Wuhan" } } ], "max_completion_tokens": 1024, "temperature": 1.0, "top_p": 0.95, "stream": false, "stop": null, "frequency_penalty": 0, "presence_penalty": 0, "thinking": { "type": "disabled" } }' ``` {/* feishu-style:text-align:left */} **Response** ```json { "id": "9ffcbd3a-d3a2-4788-9edd-d8a79e47ae15_33e773cc83af482aa0b41b9fd120fdf1", "choices": [ { "finish_reason": "stop", "index": 0, "message": { "content": "武汉明天(9月23日)将以阴天为主,局部有阵雨,气温在22~25℃之间。受冷空气影响,近期气温有所回落,体感偏凉,出门建议携带雨具并适当添衣。", "role": "assistant", "annotations": [ { "type": "url_citation", "url": "https://www.weather.com.cn/weather/101200101.shtml", "title": "武汉天气预报,武汉7天天气预报,武汉15天天气预报,武汉天气查询", "summary": " 武汉天气预报,及时准确发布中央气象台天气信息,便捷查询武汉今日天气,武汉周末天气,武汉一周天气预报,武汉 …", "site_name": "中国天气网", "publish_time": "2026-09-21T08:32:25.0000000", "logo_url": "https://weather.com.cn/favicon.ico" }, { "type": "url_citation", "url": "https://www.qweather.com/weather/wuhan-101200101.html", "title": "武汉市天气 空气质量 降水 预警以及武汉市历史天气 | 和风天气", "summary": "2级 东北风56%相对湿度弱紫外线24°体感温度25km能见度0.0mm降水量1015hPa大气压24小时预报 5am 7am 8am 9am 11am 1pm 2pm 3pm 5pm 7pm 8pm 9pm 11pm 1am 2am 3am 温度 查看完整天气地图 未来预报今天09月15日 29° 22° 周三09月16日 26° 18° 周四09月17日 22° 18° 周五09月18日 24° 21° 周六09月19日 28° 23° 周日09月20日 31° 24° 周一09月21日 32° 22° 30天预报 空气质量优 34 PM2.5 39 PM10 92 O3 0.7 CO 11 SO2 17 NO2 预报和更多 太阳现在04:45现在04:45高度角-18°方位75°ENE月亮 蛾眉月 现在04:45现在04:45高度角-64°方位68°ENE生活指数舒适度指数较舒适洗车指数较适宜穿衣指数热感冒指数少发运动指数较不宜旅游指数适宜紫外线指数弱空气污染扩散指数良武汉市的历史天气 更多历史天气 武汉市气象概述武汉属亚热带季风气候区武汉稳定≥10℃活动积温在5000℃~5300℃之间,年无霜期240天。温度 查看完整天气地图 卫星云图天气排行 今日最高温 空气质量...", "site_name": "和风天气", "publish_time": "2026-09-15T00:00:00.0000000", "logo_url": "https://www.qweather.com/favicon.ico" }, { "type": "url_citation", "url": "https://www.sina.cn/news/article/comos_nishzsi2828295.html", "title": "武汉未来一周天气预报:晴好升温至33℃,下周初阵雨来袭+FAQ|武汉天气|一周天气预报|武汉未来7天|武汉温度|武汉降雨_新浪新闻", "summary": "和风天气30天预报指出,未来一周最高温出现在9月25日(33℃),最低温出现在9月... 🌡️ 最高33℃❄️ 最低21℃☀️ 晴好为主🌧️ 阵雨2天🍃 风力1-2级 | 日期 | 天气 | 气温(℃) | 风力 | 空气质量 | |-----|-----|-----|-----|-----| | 9月19日(周六) | 晴转多云 | 21~27 | 东北风2级 | 优 | | 9月20日(周日) | 多云 | 22~30 | 北风1级 | 良 | | 9月21日(周一) | 阴转阵雨 | 22~29 | 东北风2级 | 优 | | 9月22日(周二) | 阴 | 22~25 | 东北风2级 | 优 | | 9月23日(周三) | 多云 | 21~27 | 东北风1级 | 优 | | 9月24日(周四) | 阴转多云 | 23~30 | 东北风1级 | 良 | | 9月25日(周五) | 多云转中雨 | 24~33 | 北风1级 | 优|   数据来源:综合天气网、和风天气、全国天气网、2345天气预报(2026年9月19日发布) 02降雨分布:仅21日、25日有阵雨 ⏰ 出行提醒 9月21日(周一)阵雨伴雷电,建议错峰出行或备雨具 9月25日(周...", "site_name": "新浪新闻", "publish_time": "2026-09-19T06:13:50.0000000", "logo_url": "https://sina.cn/favicon.ico" } ], "tool_calls": null } } ], "created": 1790008349, "model": "mimo-v2.6-pro", "object": "chat.completion", "usage": { "completion_tokens": 60, "prompt_tokens": 1476, "total_tokens": 1536, "completion_tokens_details": { "reasoning_tokens": 0 }, "prompt_tokens_details": { "cached_tokens": 0 }, "web_search_usage": { "tool_usage": 3, "page_usage": 3 } } } ``` {/* feishu-style:text-align:left */} For detailed parameters and invocation instructions, please refer to the [OpenAI Chat Completions API](https://platform.xiaomimimo.com/#/docs/api/text-generation/openai-api). Other API protocols are not supported for the time being. ## Supported models {/* feishu-style:text-align:left */} Currently, the `mimo-v2.6-flash`, `mimo-v2.6-pro`, `mimo-v2.6-pro-ultraspeed`, `mimo-v2.5-pro` and `mimo-v2.5` models are supported. ## Price {/* feishu-style:text-align:left */} The billing for the Web Search Plugin consists of the following two parts: - **Usage of Web Search**: The number of times internet resources appear in one response of the Internet Service Plugin. - Cost per 1,000 calls from the Web Search tool: China ¥16 / 1K requests、Overseas $5 / 1K requests.
When invoking Web Search via API, one search round will initiate concurrent keyword searches according to the `max_keyword` value, resulting in multiple uses of this plugin.
- **Model token fee**: The webpage content from internet search will be appended to the prompt, increasing the model’s input tokens. Billing is based on the model’s standard price. For price details, please refer to [Pricing and Rate Limits](https://platform.xiaomimimo.com/#/docs/pricing). ## FAQ {/* feishu-style:text-align:left */} **Why doesn’t the model perform a web search after enabling online search?** {/* feishu-style:text-align:left */} There may be three reasons: - **Cache**: There is a 5-minute cache period after enabling / disabling online search. The online search switch will not take effect immediately within 5 minutes. - **Model determines no need for search**: The model judges that the current query does not involve real-time information and can be answered directly with its own knowledge. To force a search, set `force_search: true`. - **Model not supported**: Currently, `mimo-v2.6-flash`, `mimo-v2.6-pro`, `mimo-v2.6-pro-ultraspeed`, `mimo-v2.5-pro` and `mimo-v2.5` support web search. --- DOCUMENT: Deep Thinking --- URL: https://mimo.mi.com/static/docs/quick-start/usage-guide/text-generation/deep-thinking.md # Deep Thinking {/* feishu-style:text-align:left */} Deep Thinking enables the model to perform deep reasoning before generating the final answer, analyzing problems step-by-step through an internal Chain of Thought (CoT), significantly improving accuracy on complex tasks. It is suitable for scenarios requiring deep analysis, such as complex reasoning, code generation, mathematical computation, and multi-step analysis. {/* feishu-style:text-align:left */} **Core Capabilities** - **Deep Reasoning**: Breaks down complex problems into multiple steps for analysis, improving reasoning accuracy - **Transparent Thinking Process**: Returns the complete thinking process, enhancing interpretability - **Flexible Control**: Supports enabling/disabling, allowing on-demand use based on task complexity ## Supported Models {/* feishu-style:text-align:left */} Currently, the `mimo-v2.6-flash`, `mimo-v2.6-pro`, `mimo-v2.6-pro-ultraspeed`, `mimo-v2.5-pro` and `mimo-v2.5` models are supported. ## Request Parameters {/* feishu-style:text-align:left */} Set the `thinking.type` parameter in the request to control deep thinking: `enabled` to enable deep thinking or `disabled` to disable deep thinking. {/* feishu-style:text-align:left */} **Default Status:** - Enabled by default: `mimo-v2.6-flash`, `mimo-v2.6-pro`, `mimo-v2.6-pro-ultraspeed`, `mimo-v2.5-pro` and `mimo-v2.5` ## Important Notes ### Parameter Limitations {/* feishu-style:text-align:left */} In deep thinking, `mimo-v2.6-flash`, `mimo-v2.6-pro`, `mimo-v2.6-pro-ultraspeed`, `mimo-v2.5-pro` and `mimo-v2.5` models do not support custom `temperature` and `top_p` parameters. Even if these parameters are provided, the actual effective values will be forced to use the recommended defaults of `1.0` and `0.95`. ### Multi-turn Conversation Pass-through Requirements {/* feishu-style:text-align:left */} When deep thinking is enabled in Agent product multi-turn conversations, and historical conversations contain tool calls, the assistant responses passed back in all subsequent user interaction rounds that contain tool calls must completely pass back the `reasoning_content` field, otherwise the API will return a 400 error. For the correct pass-back method, please refer to the "Multi-turn Tool Calls in Thinking Mode" section in the Call Examples.
If historical `reasoning_content` is missing, the model's context will be incomplete, which may result in decreased instruction following and increased hallucinations.
{/* feishu-style:text-align:left */} **Affected Agent Products:**
Protocol Affected Agent Products
OpenAI Compatible Protocol TRAE, Cursor, Roo Code, Codex, GitHub Copilot CLI, Zed, AutoGen, Goose
Anthropic Compatible Protocol TRAE, GitHub Copilot CLI, AutoGen, Goose, OpenClaw, OpenCode, Kilo Code
### Other Notes 1. **Output Length Limit**: `max_completion_tokens` limits the total length of thinking content and the final answer. If the thinking process is long, the token space available for the final answer will be reduced accordingly. It is recommended to set a sufficient `max_completion_tokens` to avoid answer truncation. 1. **Response Time**: Enabling deep thinking will increase response latency, especially for complex tasks. It is recommended to use `stream: true` to view the thinking process in real-time. ## Call Examples
The `thinking` field is not a standard OpenAI parameter. When passing thinking-related parameters via the OpenAI Python SDK, they must be included in `extra_body`.
### Thinking Enabled ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": "Introduce machine learning in three sentences." } ], max_completion_tokens=1024, extra_body={ "thinking": {"type": "enabled"} } ) print(completion.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": "Introduce machine learning in three sentences." } ], "max_completion_tokens": 1024, "thinking": { "type": "enabled" } }' ``` {/* feishu-style:text-align:left */} **Response Example** ```json { "id": "045bd1a4-688a-4e53-9de2-7631205e0182_5f467b0704184928936bafb270896022", "choices": [ { "finish_reason": "stop", "index": 0, "message": { "content": "Machine learning is a branch of artificial intelligence that enables computers to learn from data and improve their performance over time without being explicitly programmed for every task. It works by feeding algorithms large datasets, allowing them to identify patterns, make predictions, or take actions based on what they've learned. Common applications include spam filters, recommendation systems, image recognition, and language translation, making it a powerful tool that is reshaping industries from healthcare to finance.", "role": "assistant", "tool_calls": null, "reasoning_content": "The user wants a concise three-sentence introduction to machine learning." } } ], "created": 1790008551, "model": "mimo-v2.6-pro", "object": "chat.completion", "usage": { "completion_tokens": 103, "prompt_tokens": 60, "total_tokens": 163, "completion_tokens_details": { "reasoning_tokens": 14 }, "prompt_tokens_details": { "cached_tokens": 0 } } } ``` ### Thinking Disabled ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": "Write a short paragraph about the beauty of nature." } ], max_completion_tokens=1024, extra_body={ "thinking": {"type": "disabled"} } ) print(completion.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": "Write a short paragraph about the beauty of nature." } ], "max_completion_tokens": 1024, "thinking": { "type": "disabled" } }' ``` {/* feishu-style:text-align:left */} **Response Example** ```json { "id": "ab231e1e-4d8c-4084-84bf-521959063597_b7c5e3eece0143ae98993a27e86e6df7", "choices": [ { "finish_reason": "stop", "index": 0, "message": { "content": "There is an enduring, quiet magic in the world outside our windows. It’s found in the delicate, painted canvas of a sunrise bleeding across the horizon, and in the intricate architecture of a single leaf, each vein a tiny river carrying life. The power of nature's beauty lies not just in its grandest vistas—the silent, snow-capped mountains and vast, churning oceans—but also in its most intimate details: the scent of rain on dry earth, the gentle rhythm of waves on a shore, and the resilience of a wildflower pushing through a crack in the stone. It is a constant, unfolding masterpiece that whispers to the soul, a powerful reminder of a force that is at once both magnificent and profoundly soothing.", "role": "assistant", "tool_calls": null } } ], "created": 1790008574, "model": "mimo-v2.6-pro", "object": "chat.completion", "usage": { "completion_tokens": 146, "prompt_tokens": 64, "total_tokens": 210, "completion_tokens_details": { "reasoning_tokens": 0 }, "prompt_tokens_details": { "cached_tokens": 0 } } } ``` ### Streaming Response (Thinking Enabled)
During streaming responses, thinking content and answer content are output sequentially: first, the thinking process is returned step-by-step via `reasoning_content`, and after thinking is complete, the final answer is output step-by-step via `content`.
```python import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": "Give me some tips for improving work efficiency." } ], max_completion_tokens=1024, stream=True, extra_body={ "thinking": {"type": "enabled"} } ) for chunk in completion: print(chunk.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": "Give me some tips for improving work efficiency." } ], "max_completion_tokens": 1024, "stream": true, "thinking": { "type": "enabled" } }' ``` {/* feishu-style:text-align:left */} **Response Example** ```json data: {"id":"ac06ba44-79d8-419a-830f-8e443a15b962_f3425553986841e7bd6d1799bbb89be8","choices":[{"delta":{"content":"","role":"assistant","tool_calls":null,"reasoning_content":null},"finish_reason":null,"index":0}],"created":1790008592,"model":"mimo-v2.6-pro","object":"chat.completion.chunk"} data: {"id":"ac06ba44-79d8-419a-830f-8e443a15b962_f3425553986841e7bd6d1799bbb89be8","choices":[{"delta":{"content":null,"role":null,"tool_calls":null,"reasoning_content":"The user is asking"},"finish_reason":null,"index":0}],"created":1790008592,"model":"mimo-v2.6-pro","object":"chat.completion.chunk"} data: {"id":"ac06ba44-79d8-419a-830f-8e443a15b962_f3425553986841e7bd6d1799bbb89be8","choices":[{"delta":{"content":null,"role":null,"tool_calls":null,"reasoning_content":" for tips to"},"finish_reason":null,"index":0}],"created":1790008592,"model":"mimo-v2.6-pro","object":"chat.completion.chunk"} data: {"id":"ac06ba44-79d8-419a-830f-8e443a15b962_f3425553986841e7bd6d1799bbb89be8","choices":[{"delta":{"content":null,"role":null,"tool_calls":null,"reasoning_content":" improve work efficiency."},"finish_reason":null,"index":0}],"created":1790008592,"model":"mimo-v2.6-pro","object":"chat.completion.chunk"} ... data: {"id":"ac06ba44-79d8-419a-830f-8e443a15b962_f3425553986841e7bd6d1799bbb89be8","choices":[{"delta":{"content":null,"role":null,"tool_calls":null,"reasoning_content":" tips"},"finish_reason":null,"index":0}],"created":1790008593,"model":"mimo-v2.6-pro","object":"chat.completion.chunk"} data: {"id":"ac06ba44-79d8-419a-830f-8e443a15b962_f3425553986841e7bd6d1799bbb89be8","choices":[{"delta":{"content":null,"role":null,"tool_calls":null,"reasoning_content":"."},"finish_reason":null,"index":0}],"created":1790008593,"model":"mimo-v2.6-pro","object":"chat.completion.chunk"} data: {"id":"ac06ba44-79d8-419a-830f-8e443a15b962_f3425553986841e7bd6d1799bbb89be8","choices":[{"delta":{"content":"# Tips","role":null,"tool_calls":null,"reasoning_content":null},"finish_reason":null,"index":0}],"created":1790008593,"model":"mimo-v2.6-pro","object":"chat.completion.chunk"} data: {"id":"ac06ba44-79d8-419a-830f-8e443a15b962_f3425553986841e7bd6d1799bbb89be8","choices":[{"delta":{"content":" for Improving Work","role":null,"tool_calls":null,"reasoning_content":null},"finish_reason":null,"index":0}],"created":1790008593,"model":"mimo-v2.6-pro","object":"chat.completion.chunk"} ... data: {"id":"ac06ba44-79d8-419a-830f-8e443a15b962_f3425553986841e7bd6d1799bbb89be8","choices":[{"delta":{"content":"? ","role":null,"tool_calls":null,"reasoning_content":null},"finish_reason":null,"index":0}],"created":1790008602,"model":"mimo-v2.6-pro","object":"chat.completion.chunk"} data: {"id":"ac06ba44-79d8-419a-830f-8e443a15b962_f3425553986841e7bd6d1799bbb89be8","choices":[{"delta":{"content":"😊","role":null,"tool_calls":null,"reasoning_content":null},"finish_reason":null,"index":0}],"created":1790008602,"model":"mimo-v2.6-pro","object":"chat.completion.chunk"} data: {"id":"ac06ba44-79d8-419a-830f-8e443a15b962_f3425553986841e7bd6d1799bbb89be8","choices":[{"delta":{"content":null,"role":null,"tool_calls":null,"reasoning_content":null},"finish_reason":"stop","index":0}],"created":1790008602,"model":"mimo-v2.6-pro","object":"chat.completion.chunk","usage":null} data: {"id":"ac06ba44-79d8-419a-830f-8e443a15b962_f3425553986841e7bd6d1799bbb89be8","choices":[],"created":1790008602,"model":"mimo-v2.6-pro","object":"chat.completion.chunk","usage":{"completion_tokens":467,"prompt_tokens":61,"total_tokens":528,"completion_tokens_details":{"reasoning_tokens":29},"prompt_tokens_details":{"cached_tokens":0}}} data: [DONE] ``` ### Multi-turn Tool Calls in Thinking Mode {/* feishu-style:text-align:left */} In multi-turn conversations with deep thinking enabled, passing back `reasoning_content` when tool calls are involved ensures thinking continuity and improves model output quality. ```python import os import json from openai import OpenAI # Initialize client client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) # Define tools tools = [ { "type": "function", "function": { "name": "get_current_weather", "description": "Get the current weather for a given city", "parameters": { "type": "object", "properties": { "location": {"type": "string", "description": "City name, e.g. Beijing"}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]} }, "required": ["location"] } } }, { "type": "function", "function": { "name": "get_time", "description": "Get the current time in a given timezone", "parameters": { "type": "object", "properties": { "timezone": {"type": "string", "description": "Timezone, e.g. Asia/Shanghai"} }, "required": ["timezone"] } } } ] # Tool execution functions (replace with real API calls in production) def get_current_weather(location: str, unit: str = "celsius") -> str: weather_data = {"Beijing": "Sunny 25°C", "Shanghai": "Cloudy 22°C", "Shenzhen": "Rainy 28°C"} return weather_data.get(location, f"Weather unknown for {location}") def get_time(timezone: str) -> str: from datetime import datetime return datetime.now().strftime(f"%Y-%m-%d %H:%M:%S ({timezone})") TOOL_MAP = { "get_current_weather": lambda **kw: get_current_weather(**kw), "get_time": lambda **kw: get_time(**kw) } def run_turn(messages, turn_num): """Execute a single user turn: call model, run tools in a loop until final answer.""" request_num = 0 while True: request_num += 1 print(f"\nRequest {turn_num}-{request_num}:") response = client.chat.completions.create( model="mimo-v2.6-pro", messages=messages, tools=tools, extra_body={"thinking": {"type": "enabled"}} ) assistant_message = response.choices[0].message messages.append(assistant_message) # Print full model response print(f"reasoning_content: {assistant_message.reasoning_content}") print(f"content: \"{assistant_message.content}\"") print(f"tool_calls: {assistant_message.tool_calls}") # If no tool calls, we have the final answer if not assistant_message.tool_calls: break # Execute each tool call and append results for tool_call in assistant_message.tool_calls: func_name = tool_call.function.name func_args = json.loads(tool_call.function.arguments) result = TOOL_MAP[func_name](**func_args) print(f"-> Tool result [{func_name}]: {result}") messages.append({ "role": "tool", "tool_call_id": tool_call.id, "content": result }) # --- Multi-turn conversation --- messages = [] # Turn 1 print("=== Turn 1 ===") messages.append({"role": "user", "content": "How is the weather in Beijing today? What time is it now?"}) run_turn(messages, turn_num=1) # Turn 2: reasoning_content from Turn 1 is already in messages via assistant_message print("\n=== Turn 2 ===") messages.append({"role": "user", "content": "How about Shanghai? And is it hotter or colder than Beijing?"}) run_turn(messages, turn_num=2) ``` {/* feishu-style:text-align:left */} **Example Output** {/* feishu-style:text-align:left */} **Turn 1**: The user asks about the weather in Beijing and the current time. After receiving the user message, the model thinks and decides to call both `get_current_weather` and `get_time` tools simultaneously (Request 1-1). The client executes the tools and appends the results as `role: "tool"` messages to `messages`, then requests the model again. The model generates the final answer based on the tool results (Request 1-2). ```bash === Turn 1 === Request 1-1: reasoning_content: The user wants both weather and time in Beijing. I can call both in parallel. Weather: get_current_weather(location: Beijing). Time: Beijing timezone is Asia/Shanghai. Call both together. content: "" tool_calls: [ChatCompletionMessageFunctionToolCall(id='call_01e402113df94ebfb85e3056', function=Function(arguments='{"location": "Beijing"}', name='get_current_weather'), type='function'), ChatCompletionMessageFunctionToolCall(id='call_f2bf1ab9f73842baa64e93d3', function=Function(arguments='{"timezone": "Asia/Shanghai"}', name='get_time'), type='function')] -> Tool result [get_current_weather]: Sunny 25°C -> Tool result [get_time]: 2026-09-22 00:38:45 (Asia/Shanghai) Request 1-2: reasoning_content: Present both results clearly. content: "Here's the current information for Beijing: - **Weather:** Sunny, 25°C (77°F) ☀️ - **Current time:** 00:38 AM, Tuesday, September 22, 2026 (Asia/Shanghai time) It's a warm, clear day in Beijing — though it's just past midnight there right now, so the 25°C reading is more reflective of the warm night. Is there anything else you'd like to know?" tool_calls: None ``` {/* feishu-style:text-align:left */} **Turn 2**: The user follows up asking about Shanghai's weather and comparing it with Beijing. Since the first turn's `assistant` message (containing `reasoning_content`, `content`, `tool_calls`) has been accumulated into the conversation history via `messages.append()`, the model can directly access Beijing's 25°C information from the context, and only needs to call `get_current_weather` for Shanghai (Request 2-1), then provides a comparison answer based on both cities' weather data (Request 2-2). ```bash === Turn 2 === Request 2-1: reasoning_content: Need weather and time for Shanghai, and compare to Beijing's 25°C. Call weather and time for Shanghai in parallel. The comparison only needs weather, but the previous pattern included time — the user asked "How about Shanghai?" which could mean both. I'll fetch both to be thorough, or just weather since the question focuses on comparison. I'll get both since it's ambiguous and follows the previous context. content: "" tool_calls: [ChatCompletionMessageFunctionToolCall(id='call_7e1e7118be8a49ef9de2e8f8', function=Function(arguments='{"location": "Shanghai"}', name='get_current_weather'), type='function'), ChatCompletionMessageFunctionToolCall(id='call_3ea278a5658443258c92b3de', function=Function(arguments='{"timezone": "Asia/Shanghai"}', name='get_time'), type='function')] -> Tool result [get_current_weather]: Cloudy 22°C -> Tool result [get_time]: 2026-09-22 00:38:51 (Asia/Shanghai) Request 2-2: reasoning_content: Shanghai: cloudy, 22°C vs Beijing sunny 25°C. Shanghai is 3°C colder. Both cities share the same timezone (Asia/Shanghai), so the time is the same. content: "Here's the current information for Shanghai: - **Weather:** Cloudy, 22°C (72°F) ☁️ - **Current time:** 00:38 AM, Tuesday, September 22, 2026 (Note: Shanghai and Beijing are in the same timezone, so the time is identical in both cities) **Comparison:** Shanghai is **colder** than Beijing right now — 22°C vs. 25°C, a difference of about 3°C (5°F). Beijing is also sunny while Shanghai is cloudy. Anything else you'd like to check?" tool_calls: None ``` --- DOCUMENT: Structured Output --- URL: https://mimo.mi.com/static/docs/quick-start/usage-guide/text-generation/structured-output.md # Structured Outputs {/* feishu-style:text-align:left */} Structured Output (JSON mode) enables models to generate responses in a specified JSON format, ensuring outputs are controllable and easy to parse. It is suitable for scenarios requiring structured data, such as data extraction, form filling, classification and tagging, and API response formatting. ## Supported Models {/* feishu-style:text-align:left */} Currently, the `mimo-v2.6-flash`, `mimo-v2.6-pro`, `mimo-v2.6-pro-ultraspeed`, `mimo-v2.5-pro` and `mimo-v2.5` models are supported. ## Request Parameters - `response_format`: Response format control parameter. Pass `{"type": "json_object"}` to enable JSON mode output. - `messages`: You must explicitly instruct the model in the system or user message to return only JSON, and fully define the fields, hierarchy, and data types of the expected JSON structure. Providing examples is recommended.
Plan the value of `max_completion_tokens` carefully. Setting it too low may cause the output to be truncated, resulting in incomplete JSON that cannot be parsed.
## Call Examples ### Basic Call ```python import os import json from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.\nDo not add any explanations outside JSON. Parse meeting information and return nested structured JSON: {\"meeting_meta\": {\"meeting_topic\": string | null, \"start_time\": string | null, \"participants\": array}, \"action_items\": [{\"task\": string, \"responsible_person\": string, \"deadline\": string}]}\nFill null for unknown fields." }, { "role": "user", "content": "Product iteration meeting at 14:00 tomorrow. Attendees: Li Ming, Wang Hua. Task: Complete API document revision by Friday, taken by Wang Hua." } ], response_format={ "type": "json_object" } ) result = completion.choices[0].message.content try: parsed_json_data = json.loads(result) print(json.dumps(parsed_json_data, indent=4, ensure_ascii=False)) except json.JSONDecodeError as e: print(f"JSON parsing failed: {str(e)}") print(f"Raw complete content: {result}") ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.\nDo not add any explanations outside JSON. Parse meeting information and return nested structured JSON: {\"meeting_meta\": {\"meeting_topic\": string | null, \"start_time\": string | null, \"participants\": array}, \"action_items\": [{\"task\": string, \"responsible_person\": string, \"deadline\": string}]}\nFill null for unknown fields." }, { "role": "user", "content": "Product iteration meeting at 14:00 tomorrow. Attendees: Li Ming, Wang Hua. Task: Complete API document revision by Friday, taken by Wang Hua." } ], "response_format": { "type": "json_object" } }' ``` {/* feishu-style:text-align:left */} **Response Example** ```json { "id": "d7bbc231-ed6a-4884-a9e4-d7caff96cf25_c37a9ce97e9542079b9d948b07c10209", "choices": [ { "finish_reason": "stop", "index": 0, "message": { "content": "{\"meeting_meta\": {\"meeting_topic\": \"Product iteration meeting\", \"start_time\": \"2025-12-17 14:00\", \"participants\": [\"Li Ming\", \"Wang Hua\"]}, \"action_items\": [{\"task\": \"Complete API document revision\", \"responsible_person\": \"Wang Hua\", \"deadline\": \"2025-12-19\"}]}", "role": "assistant", "tool_calls": null, "reasoning_content": "The user wants me to parse meeting information and return structured JSON.\n\nMeeting: Product iteration meeting\nTime: 14:00 tomorrow\nParticipants: Li Ming, Wang Hua\nTask: Complete API document revision by Friday, responsible: Wang Hua\n\nI need to fill unknown fields with null. Today's date is Tuesday, December 16, 2025, so tomorrow is December 17, 2025. But I should probably just put what was given: \"14:00 tomorrow\" — I could interpret it. The instruction says to fill null for unknown fields. \"Tomorrow\" relative to today (2025-12-16) is 2025-12-17. I'll provide \"2025-12-17 14:00\" as start_time. Actually, the user message just says \"at 14:00 tomorrow\" — it's reasonable to resolve it. Friday relative to 2025-12-17 would be 2025-12-19. I'll provide that.\n\nI'll return JSON only, no explanations outside JSON." } } ], "created": 1790008400, "model": "mimo-v2.6-pro", "object": "chat.completion", "usage": { "completion_tokens": 344, "prompt_tokens": 161, "total_tokens": 505, "completion_tokens_details": { "reasoning_tokens": 253 }, "prompt_tokens_details": { "cached_tokens": 0 } } } ``` ### Streaming Output > In streaming mode, the model outputs JSON content incrementally. You must concatenate the complete JSON string on the client side before parsing to avoid failures caused by truncation. ```python import os import json from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.\nReturn only compact JSON without any extra explanations. Use the specified format: {\"destination\": string, \"trip_days\": number, \"core_activity\": array, \"budget_level\": \"low\"|\"mid\"|\"high\"}" }, { "role": "user", "content": "Plan a 3-day city trip focusing on museums and local food with a moderate budget." } ], response_format={ "type": "json_object" }, stream=True ) full_json_content = "" for chunk in completion: if not chunk.choices: continue message_delta = chunk.choices[0].delta if message_delta.content: full_json_content += message_delta.content try: parsed_json_data = json.loads(full_json_content) print(json.dumps(parsed_json_data, indent=4, ensure_ascii=False)) except json.JSONDecodeError as e: print(f"JSON parsing failed: {str(e)}") print(f"Raw complete content: {full_json_content}") ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.\nReturn only compact JSON without any extra explanations. Use the specified format: {\"destination\": string, \"trip_days\": number, \"core_activity\": array, \"budget_level\": \"low\"|\"mid\"|\"high\"}" }, { "role": "user", "content": "Plan a 3-day city trip focusing on museums and local food with a moderate budget." } ], "response_format": { "type": "json_object" }, "stream": true }' ``` {/* feishu-style:text-align:left */} **Response Example (Concatenated Streaming Result)** ```json { "destination": "Lisbon, Portugal", "trip_days": 3, "core_activity": [ "Visit the National Museum of Ancient Art", "Explore the MAAT and Belém cultural district", "Wander the Calouste Gulbenkian Museum", "Taste pastéis de nata at Pastéis de Belém", "Sample petiscos and seafood at Time Out Market", "Try bacalhau and bifana in local tascas", "Walk through Alfama with fado dinner" ], "budget_level": "mid" } ``` ## How to Get Accurate JSON Output {/* feishu-style:text-align:left */} The `json_object` mode only guarantees syntactically valid JSON output. The actual data structure is entirely defined by your prompt. The clearer and more complete your prompt constraints are, the more closely the model's JSON output will match your expectations. {/* feishu-style:text-align:left */} **1. Enforce JSON-Only Output** {/* feishu-style:text-align:left */} Explicitly require the model to return JSON only, with no explanations, comments, or Markdown code blocks, to prevent parsing errors caused by extra text. {/* feishu-style:text-align:left */} **2. Provide a Complete JSON Structure Template** {/* feishu-style:text-align:left */} List all fields, data types, and nesting levels completely. Adding examples makes it easier to constrain the model: ```json { "name": string, "count": number, "tags": string[], "date": "YYYY-MM-DD" } ``` {/* feishu-style:text-align:left */} **3. Constrain Enum Values and Numeric Ranges** - Use enum annotations for fixed options: `"status": "active" | "inactive" | "pending"` - Constrain numeric ranges: `score: number (0-100)` {/* feishu-style:text-align:left */} **4. Define Null Value Handling Rules** {/* feishu-style:text-align:left */} Specify in advance that unknown fields should be filled with `null`, e.g., in your prompt: fill null for unknown fields. {/* feishu-style:text-align:left */} **5. Ensure Reliable JSON Output** {/* feishu-style:text-align:left */} Enabling `json_object` mode only guarantees syntactically valid JSON; it cannot enforce that returned fields, data types, and nesting match your predefined structure. In production environments, it is recommended to use the `jsonschema` library for strict structural validation of model output. If validation fails, you can retry with an enhanced prompt, apply business-level fallback logic, or take other contingency measures. {/* feishu-style:text-align:left */} Below is a Python example. The scenario: parse customer service conversation text, automatically extract core ticket information, and implement automatic ticket classification, urgency level assignment, and intelligent dispatch. ```python import json import os from jsonschema import validate, ValidationError from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) ticket_schema = { "type": "object", "properties": { "problem_type": { "type": "string", "enum": ["device_fault", "consult", "complaint", "other"] }, "device_model": {"type": ["string", "null"]}, "urgent_level": { "type": "string", "enum": ["low", "normal", "high"] }, "user_requirements": {"type": "array", "items": {"type": "string"}} }, "required": ["problem_type", "device_model", "urgent_level", "user_requirements"] } def extract_ticket(text: str): sys_prompt = ( "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.\n" "Return JSON only, no explanations, no extra text.\n" "Format: {\"problem_type\":\"device_fault|consult|complaint|other\",\"device_model\":string|null,\"urgent_level\":\"low|normal|high\",\"user_requirements\":array}\n" "Fill null if device model is unknown." ) messages = [ {"role": "system", "content": sys_prompt}, {"role": "user", "content": text} ] resp = client.chat.completions.create( model="mimo-v2.6-pro", messages=messages, response_format={"type": "json_object"}, stream=False ) json_text = resp.choices[0].message.content data = json.loads(json_text) try: validate(instance=data, schema=ticket_schema) return {"ok": True, "data": data} except ValidationError as e: return {"ok": False, "error": e.message, "raw": data} if __name__ == "__main__": # Test with two sample tickets input_list = [ "My speaker can't connect Wi-Fi, I restarted it but no use, hope to solve quickly, don't know the model.", "Want to ask how to set timing function, device model is Watch S1, no urgent demand." ] for idx, text in enumerate(input_list, 1): print(f"==== Ticket {idx} extraction result ====") res = extract_ticket(text) if res["ok"]: print(json.dumps(res["data"], indent=4)) else: print(f"Validation failed: {res['error']}") ``` --- DOCUMENT: Batch Inference (Batch API) --- URL: https://mimo.mi.com/static/docs/quick-start/usage-guide/text-generation/batch-api.md # Batch API {/* feishu-style:text-align:left */} Batch Inference (Batch API) is an offline, large-scale data processing solution for workloads that do not require real-time responses. The interface is OpenAI-compatible and is well suited for scenarios such as model evaluation, data labeling, and batch regression testing. You can create asynchronous jobs via the SDK or the console. Key advantages: - ⭐ **Lower cost**: Batch Inference (Batch API) is priced at 50% of the real-time API > Models supported for batch inference: `mimo-v2.6-pro`, `mimo-v2.6-flash`,when using the batch inference service for model calls, the model name must be in lowercase. - ⭐ **Off-peak smart scheduling**: After a job is submitted, the system automatically schedules it for execution during off-peak hours, making full use of idle compute - ⭐ **Flexible usage**: Two ways to use it — API and console — supporting full lifecycle job management through either the API or the Open Platform console - ⭐ **OpenAI compatible**: The API protocol is identical to OpenAI's, making migration from OpenAI Batch simple ## Applicable Scenarios
Scenario Description
**Data labeling** Label large volumes of text and images
**Model evaluation** Benchmark testing, regression validation
**Content moderation** Offline batch classification and filtering
**Batch generation** Summarization, translation, structured extraction
**Academic research** Large-scale data experiments, research data processing
## Batch Inference API Pricing {/* feishu-style:text-align:left */} Xiaomi MiMo API pay-as-you-go billing uses a regular Open Platform API Key and deducts from your account balance based on actual token usage. It is not interchangeable with Token Plan package quotas.
Billing notes - Batch Inference (Batch API) price = real-time API price × 50% - Billing unit: CNY per million tokens (Mainland China); USD per million tokens (Overseas) - Cache hits: When the request's prefix content hits the Prompt Cache, billing uses the cache-hit price - Batch inference currently supports the `mimo-v2.6-pro` and `mimo-v2.6-flash` models. Pricing for both models in Mainland China and overseas is listed below - API Key format for batch inference: sk-xxxxx. Go to [API Keys](https://platform.xiaomimimo.com/console/api-keys) to create an API Key.
{/* feishu-style:text-align:left */} **Mainland China pricing**
Inference type Real-time Inference API Batch Inference API
MiMo-V2.6 Series Input (cache hit) Input (cache miss) Output Input (cache hit) Input (cache miss) Output
`mimo-v2.6-pro` ¥0.025 ¥3.00 ¥6.00 ¥0.0125 ¥1.50 ¥3.00
`mimo-v2.6-flash` ¥0.02 ¥1.00 ¥2.00 ¥0.01 ¥0.50 ¥1.00
{/* feishu-style:text-align:left */} **Overseas pricing**
Inference type Real-time Inference API Batch Inference API
MiMo-V2.6 Series Input (cache hit) Input (cache miss) Output Input (cache hit) Input (cache miss) Output
`mimo-v2.6-pro` $0.0036 $0.435 $0.87 $0.0018 $0.2175 $0.435
`mimo-v2.6-flash` $0.0028 $0.14 $0.28 $0.0014 $0.07 $0.14
## Getting Started with Batch Inference ### Prerequisites {/* feishu-style:text-align:left */} Before using Batch Inference (Batch API), complete the following steps:
Step Description
1. Register an account Register a Xiaomi MiMo Open Platform account
2. Real-name verification Complete real-name verification
3. Top up balance Top up your account balance (Batch API deducts from your balance)
4. Get an API Key Create a usable API Key
- Batch task Created via API: use the API Key you created
- Batch task Created via the console: attached to the platform's default API Key by default; no need to use your own API Key
5. Get the Base URL Get the Base URL,Go to the [Batch Inference](https://platform.xiaomimimo.com/console/batch) page to get it
### Step 1: Prepare the File {/* feishu-style:text-align:left */} Prepare your data file in JSONL format (JSON Lines — one JSON object per line) following the file format requirements below. Each line contains the details of a single API request. #### File Examples {/* feishu-style:text-align:left */} Below is an example of an input file containing 2 requests (a `.jsonl` file), corresponding to the 3 supported endpoint formats. **Note: when using batch inference via the console, only the OpenAI | completions file format is currently supported for upload. Using the API is more flexible and supports all 3 file formats.**
Base URL: Go to the [Batch Inference](https://platform.xiaomimimo.com/console/batch) page to get it
{/* feishu-style:text-align:left */} **OpenAI | completions,** [**download sample file**](https://aistudio-cdn.xiaomimimo.com/xiaomimimo-static/mimo-docs-figures/others/input-completions.jsonl) ```json {"custom_id": "request-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "mimo-v2.6-flash", "messages": [{"role": "user", "content": "Hello"}]}} {"custom_id": "request-2", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "mimo-v2.6-pro", "messages": [{"role": "user", "content": "World"}]}} ``` {/* feishu-style:text-align:left */} **OpenAI | responses,** [**download sample file**](https://aistudio-cdn.xiaomimimo.com/xiaomimimo-static/mimo-docs-figures/others/input-responses-update.jsonl) ```json {"custom_id": "request-1", "method": "POST", "url": "/v1/responses", "body": {"model": "mimo-v2.6-pro", "input": "Hello"}} {"custom_id": "request-2", "method": "POST", "url": "/v1/responses", "body": {"model": "mimo-v2.6-flash", "input": "please introduce yourself"}} ``` {/* feishu-style:text-align:left */} **Anthropic | messages,** [**download sample file**](https://aistudio-cdn.xiaomimimo.com/xiaomimimo-static/mimo-docs-figures/others/input-claude-update.jsonl) ```json {"custom_id": "request-1", "method": "POST", "url": "/anthropic/v1/messages", "body": {"model": "mimo-v2.6-pro", "messages": [{"role": "user", "content": [{"type": "text", "text": "Hello"}]}]}} {"custom_id": "request-2", "method": "POST", "url": "/anthropic/v1/messages", "body": {"model": "mimo-v2.6-flash", "messages": [{"role": "user", "content": [{"type": "text", "text": "please introduce yourself"}]}]}} ``` #### File Specification Requirements - **Basic file requirements** {/* feishu-style:text-align:left */} Default maximum file size is 128 MB. A single file may only contain requests for one batch inference endpoint. - **Field rules** - Every request must include a `custom_id` field, **a string that is unique within the file**, used to map each request to its corresponding result. - Each request is sent and returns its result independently. If multiple requests share the same prompt, include the same prompt in each request. - The `body` field of each request must be consistent with the request body of the underlying model invocation API and must be a valid JSON Object. ### Step 2: Create a Batch Inference Job 1. On the [Batch Inference](https://platform.xiaomimimo.com/console/batch) page , click **Create Batch Inference Job**. 1. On the create job page, upload your JSONL file, fill in a **job description**, and set the **maximum wait time** (1–14 days). > **When using batch inference via the console, only the OpenAI | completions file format is currently supported for upload. Using the API is more flexible and supports all 3 file formats.** > > - OpenAI | completions > > - OpenAI | responses > > - Anthropic | messages 1. When finished, click **Create**. ### Step 3: Manage Jobs 1. **View**: - On the job list page, view the **progress** (processed requests / total requests) and **status** of each job. Or enter the job detail page for more information. - Search by job description or ID to quickly locate a target job. 1. **Manage**: - Cancel: Jobs in the "In Progress" state can be cancelled from the **Actions** column. - Troubleshoot errors: "Failed" jobs allow you to download the error file for details. File downloads are available from both the job list Actions column and the job detail page. ### Step 4: Download Results
The system retains your data for only 30 days. Please download and back up your data promptly. After expiry, files are automatically deleted and cannot be recovered.
{/* feishu-style:text-align:left */} After a job completes, files can be downloaded from the job list Actions column and the job detail page: - **Success file**: records all successful requests and their `response` results. - **Error file (if any)** : records all failed requests and their `error` details. {/* feishu-style:text-align:left */} Both files include the `custom_id` field, used to match against your original input data, associate results, or locate errors. ### Step 5: View Usage Statistics (Optional) {/* feishu-style:text-align:left */} On the [Billing Details](https://platform.xiaomimimo.com/console/usage) page, filter and view usage statistics for batch inference. {/* feishu-style:text-align:left */} **View data overview**: **Select a time range**, set **Inference Type** to **Batch Inference**, select an **API Key**, and view the batch inference model invocation overview. > - Jobs created via the API: use the API Key you created; select that Key to view usage > > - Jobs created via the console: attached to the platform's default API Key by default; no need to use your own API Key; select **Other** to view usage ## API Reference ### Step 1: Upload the Job File to the File Service {/* feishu-style:text-align:left */} You can use Curl to upload the job file to the file service's bucket. The gateway platform will subsequently read the request information from the file for batch inference. {/* feishu-style:text-align:left */} **Request example** {/* feishu-style:text-align:left */} API Key format for batch inference: sk-xxxxx. Go to [API Keys](https://platform.xiaomimimo.com/console/api-keys) to create an API Key. ```json curl https://batch-api-${region}.xiaomimimo.com/v1/files \ -H "Authorization: Bearer $ARK_API_KEY" \ -F 'purpose=batch' \ -F 'file=@/Users/doc/demo.jsonl' \ # file path ``` {/* feishu-style:text-align:left */} **Request parameters** - purpose — Upload file category - batch (batch inference file) - file=@ — Local path of the file to upload
The default file expiration time is currently 30 days
{/* feishu-style:text-align:left */} **Sample request response** ```json { "id": "file-8a761b15a195", "object": "file", "purpose": "batch", "filename": "input.jsonl", "bytes": 327, "status": "active", "error": null, "metadata": null, "mime_type": "application/jsonl", "created_at": 1780034477, "expire_at": 1780038877, "preprocess_configs": null } ```
Field Description
id Unique file identifier, formatted as `file-` + UUID prefix
object Object type; fixed value `"file"`, indicating this is a file resource
purpose File category. `"batch"` indicates use for batch processing
filename Original file name
bytes File size in **bytes**
status File status. `"active"` means the file is available; other possible values include `"pending"` (processing) and `"error"` (failed)
error Error message; null when there is no error; contains the specific error description on failure
metadata Custom metadata — user-supplied key-value pairs; not currently used
mime_type MIME type
created_at Creation time, Unix timestamp (seconds)
expire_at Expiration time, Unix timestamp. The file may be automatically cleaned up after expiry
### Step 2: Create a Batch Inference Job
Custom job timeout is supported, ranging from 1 to 14 days
{/* feishu-style:text-align:left */} **Request example** ```bash curl -X POST https://batch-api-${region}.xiaomimimo.com/v1/batches \ -H "Authorization: Bearer $ARK_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "input_file_id": "$FILE_ID", "endpoint": "/v1/chat/completions", "completion_window": "24h", "name": "test" }' ``` {/* feishu-style:text-align:left */} **Response example** ```json { "id": "batch_c089f12dde664f07aeadd8c7", "object": "batch", "endpoint": "/v1/chat/completions", "errors": null, "input_file_id": "file-0f4add6a5001", "completion_window": "24h", "status": "validating", "output_file_id": null, "error_file_id": null, "created_at": 1711402400, "in_progress_at": null, "expires_at": 1711488800, "finalizing_at": null, "completed_at": null, "failed_at": null, "expired_at": null, "cancelling_at": null, "cancelled_at": null, "request_counts": { "total": 0, "completed": 0, "failed": 0 } } ```
**Field** **Type** **Description**
id string Unique Batch Job identifier, formatted as `batch_` + random ID
object string Object type, fixed as `"batch"`
endpoint string API endpoint for this batch call, e.g. `/v1/chat/completions`
errors object/null Error message; null when there is no error
input_file_id string Input file ID, pointing to the uploaded `.jsonl` file
completion_window string Completion window; currently fixed at `"24h"`. Becomes `expired` on timeout
status string Current status (9 possible values — see the previous message)
output_file_id string/null Output file ID; populated only after completion; can be used to download results
error_file_id string/null Error file ID; populated only when there are failed requests; records the reason for each failure
created_at int Creation time (Unix timestamp, seconds)
in_progress_at int/null Time the job entered processing; null when processing has not started
expires_at int Expiration time, i.e. `created_at + completion_window`
finalizing_at int/null Time the job entered the finalizing stage
completed_at int/null Completion time; populated only in the `completed` state
failed_at int/null Failure time; populated only in the `failed` state
expired_at int/null Expiration time; populated only in the `expired` state
cancelling_at int/null Time the cancellation was initiated
cancelled_at int/null Time the cancellation completed; populated only in the `cancelled` state
request_counts object Request counter
request_counts.total int Total number of requests
request_counts.completed int Number of completed requests
request_counts.failed int Number of failed requests
### Step 3: Query Batch Inference Job Status {/* feishu-style:text-align:left */} **Request example** ```bash curl https://batch-api-${region}.xiaomimimo.com/v1/batches/$BATCH_ID \ -H "Authorization: Bearer $ARK_API_KEY" ``` {/* feishu-style:text-align:left */} Batch inference job statuses and their descriptions:
Status Status code Description
Initializing validating The job is initializing.
Running in_progress The job is running.
Completed completed The job has fully completed.
Failed failed Job execution failed, possibly due to timeout or other reasons.
Cancelling cancelling The user is actively cancelling the job
Cancelled cancelled The user's cancellation succeeded; the job has been terminated
### Step 4: Download Batch Inference Job Results #### Success file (output_file) {/* feishu-style:text-align:left */} One result object per line, mapped to the input via `custom_id`: ```bash {"id":"batch_req_xxx","custom_id":"request-1","response":{"status_code":200,"body":{"id":"chatcmpl-xxx","choices":[{"message":{"content":"Hello!"}}]}},"error":null} {"id":"batch_req_yyy","custom_id":"request-2","response":{"status_code":200,"body":{"id":"chatcmpl-yyy","choices":[{"message":{"content":"World!"}}]}},"error":null} ``` #### Error file (error_file) {/* feishu-style:text-align:left */} Generated only when there are failed requests: ```bash {"id":"batch_req_zzz","custom_id":"request-3","response":null,"error":{"code":"inference_failed","message":"400 Bad Request"}} ``` {/* feishu-style:text-align:left */} **The input** `**custom_id**` **appears verbatim in the output**, used to align inputs with outputs. {/* feishu-style:text-align:left */} After the batch inference job finishes running, you can download the result files via curl. The result files fall into 2 categories: {/* feishu-style:text-align:left */} **Download the output file** ```bash curl https://batch-api-${region}.xiaomimimo.com/v1/files/${result_file_id}/content \ -H "Authorization: Bearer YOUR_API_KEY" \ -o output.jsonl ``` {/* feishu-style:text-align:left */} Output file format (JSONL, one result per line): ```json {"id":"f7ed2","custom_id":"request-1","response":{"status_code":200,"body":{"id":"chatcmpl-xxx","choices":[{"message":{"content":"Hello!"}}]}},"error":null} {"id":"a3b21","custom_id":"request-2","response":{"status_code":200,"body":{"id":"chatcmpl-yyy","choices":[{"message":{"content":"World!"}}]}},"error":null} ``` {/* feishu-style:text-align:left */} **Download the error file** {/* feishu-style:text-align:left */} If there are failed requests, `error_file_id` is not empty: ```bash curl https://batch-api-${region}.xiaomimimo.com/v1/files/${error_file_id}/content \ -H "Authorization: Bearer YOUR_API_KEY" \ -o errors.jsonl ``` {/* feishu-style:text-align:left */} Error file format: ```json {"id":"db7da","custom_id":"request-3","response":null,"error":{"code":"inference_failed","message":"400 Bad Request"}} ``` ### Other API Examples #### Cancel a Batch Inference Job {/* feishu-style:text-align:left */} **Request example** ```bash curl -X POST https://batch-api-${region}.xiaomimimo.com/v1/batches/batch_c089f12dde664f07aeadd8c7/cancel \ -H "Authorization: Bearer YOUR_API_KEY" ``` {/* feishu-style:text-align:left */} **Response example** ```json { "id": "batch_c089f12dde664f07aeadd8c7", "object": "batch", "endpoint": "/v1/chat/completions", "errors": null, "input_file_id": "file-0f4add6a5001", "completion_window": "24h", "status": "cancelling", "output_file_id": null, "error_file_id": null, "created_at": 1711402400, "in_progress_at": 1711402410, "expires_at": 1711488800, "finalizing_at": null, "completed_at": null, "failed_at": null, "expired_at": null, "cancelling_at": 1711402600, "cancelled_at": null, "request_counts": { "total": 2, "completed": 1, "failed": 0 } } ```
Cancellation is an asynchronous operation; the returned status is `cancelling`. On subsequent queries the status becomes `cancelled`, at which point `output_file_id` and `error_file_id` are not null (the portion already completed generates an output file).
#### OpenAI SDK Compatibility {/* feishu-style:text-align:left */} Batch Inference (Batch API) is OpenAI-protocol compatible and can be used directly with the OpenAI Python SDK: ```python from openai import OpenAI client = OpenAI( api_key="your-mimo-api-key", base_url="https://batch-api-${region}.xiaomimimo.com/v1" ) # Upload the file file = client.files.create( file=open("input.jsonl", "rb"), purpose="batch" ) # Create the batch batch = client.batches.create( input_file_id=file.id, endpoint="/v1/chat/completions", completion_window="24h" ) # Query status batch = client.batches.retrieve(batch.id) print(f"Status: {batch.status}") print(f"Completed: {batch.request_counts.completed}/{batch.request_counts.total}") # Download results if batch.output_file_id: result = client.files.content(batch.output_file_id) with open("output.jsonl", "wb") as f: f.write(result.content) ``` ## FAQ ### Why doesn't the job start immediately after submission? {/* feishu-style:text-align:left */} Batch Inference (Batch API) uses an **off-peak scheduling** strategy: the system schedules jobs automatically based on online resource availability, with no manual intervention required. When resources are tight, job startup and execution may be delayed. ### How do I retry a failed job? {/* feishu-style:text-align:left */} Batch Inference (Batch API) currently does **not support in-job retry or resume**. If a job fails, you need to: 1. Download the error file; 1. Review the failure reasons in the error file; 1. Create a new job and resubmit. ### What should I do if the job is partially successful? 1. If some requests in your job succeeded and others failed, the **success file** contains all successful request results, and the **error file** contains the failed requests and their error reasons. Successfully completed requests are billed normally. 1. You can download the error file, review the failure reasons, and create a new job to resubmit. ### What happens if my balance is insufficient? 1. Balance is 0 at job creation: you can still browse and go through the job creation flow; the job creation will fail 1. Balance becomes insufficient while the job is running: the job fails. The portion already completed generates a success file and is billed normally; the portion not executed generates a failure file and is not charged ### Can I use it without real-name verification? {/* feishu-style:text-align:left */} **No.** Real-name verification must be completed before using Batch Inference (Batch API). Unverified users who visit the batch inference page will see a guide page and be redirected to the real-name verification page. ### Is streaming output supported? {/* feishu-style:text-align:left */} **No.** Batch scenarios are asynchronous processing and do not apply to the streaming protocol. ### Is Token Plan deduction supported? {/* feishu-style:text-align:left */} **No.** Batch Inference (Batch API) only deducts from your account's cash balance based on actual token usage and is not interchangeable with Token Plan package quotas. ### How long are result files retained? {/* feishu-style:text-align:left */} Input files and result files are retained for **30 days by default** and are automatically cleaned up after expiry. Please download them promptly. ### How are completed, failed, and cancelled jobs billed? - Only successful requests are billed. Requests that fail during file parsing or job execution **incur no charges** - **Requests successfully completed before cancellation are billed normally** --- DOCUMENT: Image Understanding --- URL: https://mimo.mi.com/static/docs/quick-start/usage-guide/multimodal-understanding/image-understanding.md # Image Understanding {/* feishu-style:text-align:left */} The image understanding model can answer based on the images you provide, supporting both image URL and Base64 encoding as input methods, and is suitable for scenarios such as image description and classification. ## Quick Start
For preparations such as obtaining an API Key, please refer to [First API Call](https://platform.xiaomimimo.com/#/docs/quick-start/first-api-call).
{/* feishu-style:text-align:left */} Quickly experience the effect of image understanding by passing the model through the image URL method. The sample code is as follows. ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "https://example-files.cnbj1.mi-fds.com/example-files/image/image_example.png" } }, { "type": "text", "text": "please describe the content of the image" } ] } ], max_completion_tokens=1024 ) print(completion.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "https://example-files.cnbj1.mi-fds.com/example-files/image/image_example.png" } }, { "type": "text", "text": "please describe the content of the image" } ] } ], "max_completion_tokens": 1024 }' ``` {/* feishu-style:text-align:left */} **Response** ```json { "id": "75a0045f-a28f-4b0f-8ddd-6c90c6ef6587_90e3b2e2dd1f410c8b5526a21526bdec", "choices": [ { "finish_reason": "stop", "index": 0, "message": { "content": "This is a serene, enchanted-looking ancient forest scene, rich with detail:\n1. **Foreground:** At the bottom left, lush, large green ferns with delicate fronds grow on a bed of moss-covered smooth river stones. A shallow, clear stream flows gently over the stones, its surface glinting brightly as it catches golden sunlight. Mossy rocks and small, leafy green shrubs line the right bank, their leaves glowing in the warm light.\n2. **Midground:** The stream winds deeper into the woods, surrounded by thick, vibrant moss and lush low greenery. Dominating the left is a massive, gnarled ancient tree with thick, exposed roots gripping the earth, its dark, twisted trunk curving outward. Other tall, mature trees with textured dark bark fill the space, creating a dense, old-growth forest feel.\n3. **Light & Atmosphere:** Warm, golden sunbeams (crepuscular rays) filter down through the forest canopy, cutting through a soft mist that hangs in the air. Tiny glowing particles, like sun dust or fireflies, drift in the sunlit space, adding a magical, dreamlike quality. The distant forest fades into soft haze, giving the scene deep, peaceful atmosphere.\n4. **Overall feel:** The image evokes a sense of quiet, untouched wilderness, peaceful and full of life, with the gentle flow of the stream and warm, dappled light creating an intimate, enchanted natural space.", "role": "assistant", "tool_calls": null, "reasoning_content": "Got it, let's break down this beautiful forest scene step by step.\n\nFirst, the overall mood: it's a misty, sun-dappled ancient forest, with a sense of quiet, magic (the little light specks feel like fireflies or sun particles), peaceful and lush.\n\nForeground: At the bottom left, large, vibrant green ferns with feathery fronds sit on moss-covered, smooth rounded stones. The stream here is shallow, with clear water flowing over the stones, catching the golden sunlight. Mossy rocks line the right side, with small leafy green bushes growing on them, their leaves catching the warm light.\n\nMidground: A gentle stream winds its way deeper into the forest, its surface glinting with the sunbeams. The banks are thick with moss, ferns, and low greenery, all lush and healthy. Large, ancient trees dominate here: on the left, a massive tree with thick, gnarled, exposed roots gripping the ground, its dark trunk curving outward. Other tall trees with dark, textured bark fill the space, their trunks disappearing into the mist.\n\nBackground & lighting: The scene is backlit by warm, golden sunbeams (crepuscular rays) that filter through the tree canopy, cutting through the mist that hangs in the air, creating a soft, hazy glow. The light catches tiny particles in the air, making them look like glowing embers or fireflies, adding a magical touch. The distant forest fades into soft mist, giving the scene depth.\n\nAll elements come together to feel like a serene, untouched, enchanted forest space, full of life and soft, warm light, with the sound of the gentle stream almost imaginable." } } ], "created": 1790007893, "model": "mimo-v2.6-pro", "object": "chat.completion", "usage": { "completion_tokens": 656, "prompt_tokens": 1085, "total_tokens": 1741, "completion_tokens_details": { "reasoning_tokens": 351 }, "prompt_tokens_details": { "cached_tokens": 1024, "image_tokens": 1024 } } } ``` ## Supported models {/* feishu-style:text-align:left */} Currently, the `mimo-v2.6-flash`, `mimo-v2.6-pro`, `mimo-v2.6-pro-ultraspeed` and `mimo-v2.5` models are supported. ## Image Input Method {/* feishu-style:text-align:left */} Supported ways to upload images are as follows: - Image URL Input: A publicly accessible image URL address must be provided. - Base64 Encoding Input: Convert the image to a Base64-encoded string before passing it in. ### Image URL Input {/* feishu-style:text-align:left */} Directly pass in the image via the publicly accessible image URL, which is suitable for scenarios where the image is already stored in a publicly accessible environment. The file size of a single image cannot exceed 50 MB. #### OpenAI Chat Completions API ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "https://example-files.cnbj1.mi-fds.com/example-files/image/image_example.png" } }, { "type": "text", "text": "please describe the content of the image" } ] } ], max_completion_tokens=1024 ) print(completion.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "https://example-files.cnbj1.mi-fds.com/example-files/image/image_example.png" } }, { "type": "text", "text": "please describe the content of the image" } ] } ], "max_completion_tokens": 1024 }' ``` #### Anthropic Messages API ```python import os from anthropic import Anthropic client = Anthropic( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/anthropic" ) message = client.messages.create( model="mimo-v2.6-pro", max_tokens=1024, system="You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.", messages=[ { "role": "user", "content": [ { "type": "image", "source": { "type": "url", "url": "https://example-files.cnbj1.mi-fds.com/example-files/image/image_example.png" } }, { "type": "text", "text": "please describe the content of the image" } ] } ] ) print(message.content) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/anthropic/v1/messages' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "max_tokens": 1024, "system": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.", "messages": [ { "role": "user", "content": [ { "type": "image", "source": { "type": "url", "url": "https://example-files.cnbj1.mi-fds.com/example-files/image/image_example.png" } }, { "type": "text", "text": "please describe the content of the image" } ] } ] }' ``` ### Base64 Encoded Input {/* feishu-style:text-align:left */} Convert the image file to a Base64-encoded string and then pass it in, which is suitable for scenarios where the image cannot be accessed via a public network URL. The size of the converted Base64-encoded string cannot exceed 50 MB. > In the following example, both `{MIME_TYPE}` and `$BASE64_IMAGE` are placeholders, please replace them with actual valid data before running. #### OpenAI Chat Completions API
Please include the prefix before Base64 encoding: `data:{MIME_TYPE};base64,$BASE64_IMAGE` - `{MIME_TYPE}`: The MIME type (media type) of the image, used to identify the image format, needs to be replaced with the MIME value corresponding to the actual image. - `$BASE64_IMAGE`: A pure Base64-encoded string of the image file (without any prefix).
```python import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "data:{MIME_TYPE};base64,$BASE64_IMAGE" } }, { "type": "text", "text": "please describe the content of the image" } ] } ], max_completion_tokens=1024 ) print(completion.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "data:{MIME_TYPE};base64,$BASE64_IMAGE" } }, { "type": "text", "text": "please describe the content of the image" } ] } ], "max_completion_tokens": 1024 }' ``` #### Anthropic Messages API ```python import os from anthropic import Anthropic client = Anthropic( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/anthropic" ) message = client.messages.create( model="mimo-v2.6-pro", max_tokens=1024, system="You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.", messages=[ { "role": "user", "content": [ { "type": "image", "source": { "type": "base64", "media_type": "{MIME_TYPE}", "data": "$BASE64_IMAGE" } }, { "type": "text", "text": "please describe the content of the image" } ] } ] ) print(message.content) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/anthropic/v1/messages' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "max_tokens": 1024, "system": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.", "messages": [ { "role": "user", "content": [ { "type": "image", "source": { "type": "base64", "media_type": "{MIME_TYPE}", "data": "$BASE64_IMAGE" } }, { "type": "text", "text": "please describe the content of the image" } ] } ] }' ``` ### Multi-image Input {/* feishu-style:text-align:left */} Supports simultaneously passing in public network URLs or Base64-encoded strings of multiple images, and the model can parse the image content and return responses that match the image semantics. ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "https://example-files.cnbj1.mi-fds.com/example-files/image/image_example.png" } }, { "type": "image_url", "image_url": { "url": "data:{MIME_TYPE};base64,$BASE64_IMAGE" } }, { "type": "text", "text": "please describe the connections and differences between these two pictures" } ] } ], max_completion_tokens=1024 ) print(completion.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "https://example-files.cnbj1.mi-fds.com/example-files/image/image_example.png" } }, { "type": "image_url", "image_url": { "url": "data:{MIME_TYPE};base64,$BASE64_IMAGE" } }, { "type": "text", "text": "please describe the connections and differences between these two pictures" } ] } ], "max_completion_tokens": 1024 }' ``` ## Image Restrictions - Image Formats: JPEG, PNG, GIF, WebP, BMP. - Image Size: - When passed in as a URL: single image file size does not exceed 50 MB. - When passed in as Base64 encoding: The size of the Base64 encoded string of a single image does not exceed 50 MB. - Number of images: When multiple images are passed in, the number of images is limited by the model's context length, and the total number of Tokens for all images and text must be less than the model's context length. > Note: For calculating image tokens, please refer to [Explanation of Image Token Usage](https://platform.xiaomimimo.com/#/docs/usage-guide/multimodal-understanding/image-understanding?target=explanation-of-image-token-usage-and-scaling-rules). For the model context length, please refer to [Pricing and Rate Limits](https://platform.xiaomimimo.com/#/docs/pricing). ## Explanation of Image Token Usage and Scaling Rules {/* feishu-style:text-align:left */} The calculation rules for images are relatively complex. For Token conversion and scaling rules, please refer to the following code. The estimated results are for reference only, and the actual usage shall be subject to the API response. ```python import math from PIL import Image PATCH_SIZE = 16 SPATIAL_MERGE_SIZE = 2 TEMPORAL_PATCH_SIZE = 2 IMAGE_MIN_PIXELS = 8192 IMAGE_MAX_PIXELS = 8388608 def calc_image_tokens(image_path: str) -> dict: image = Image.open(image_path) height = image.height width = image.width factor = PATCH_SIZE * SPATIAL_MERGE_SIZE # 32 h_bar = round(height / factor) * factor w_bar = round(width / factor) * factor if h_bar * w_bar > IMAGE_MAX_PIXELS: beta = math.sqrt((height * width) / IMAGE_MAX_PIXELS) h_bar = math.floor(height / beta / factor) * factor w_bar = math.floor(width / beta / factor) * factor elif h_bar * w_bar < IMAGE_MIN_PIXELS: beta = math.sqrt(IMAGE_MIN_PIXELS / (height * width)) h_bar = math.ceil(height / beta / factor) * factor w_bar = math.ceil(width / beta / factor) * factor grid_t = 1 grid_h = h_bar // PATCH_SIZE grid_w = w_bar // PATCH_SIZE num_tokens = (grid_t * grid_h * grid_w) // (SPATIAL_MERGE_SIZE ** 2) return num_tokens if __name__ == "__main__": token = calc_image_tokens(image_path="xxx/test.jpg") print(token) ``` ## Price - Billing: Total cost is calculated based on the number of input, input (cache hits), and output tokens; for pricing, please refer to [Pricing and Rate Limits](https://platform.xiaomimimo.com/#/docs/pricing). - The Token consumption of images can be calculated through [Explanation of Image Token Usage](https://platform.xiaomimimo.com/#/docs/usage-guide/multimodal-understanding/image-understanding?target=explanation-of-image-token-usage-and-scaling-rules). The estimated results are for reference only, and the actual usage shall be subject to the API response. - View Bill: You can view your bill and usage on the [Billing](https://platform.xiaomimimo.com/#/console/usage) page in the Console. ## FAQ ### Does it support local file upload? {/* feishu-style:text-align:left */} `mimo-v2.6-flash`, `mimo-v2.6-pro`, `mimo-v2.6-pro-ultraspeed` and `mimo-v2.5` models don't currently support uploading local image files. For supported upload methods, please refer to [Image Input Method](https://platform.xiaomimimo.com/#/docs/usage-guide/multimodal-understanding/image-understanding?target=image-input-method). --- DOCUMENT: Audio Understanding --- URL: https://mimo.mi.com/static/docs/quick-start/usage-guide/multimodal-understanding/audio-understanding.md # Audio Understanding {/* feishu-style:text-align:left */} The audio understanding model can answer based on the audio you provide, supporting both audio URL and Base64 encoding as input methods, and is suitable for scenarios such as audio analysis. ## Quick Start
For preparations such as obtaining an API Key, please refer to [First API Call](https://platform.xiaomimimo.com/#/docs/quick-start/first-api-call).
{/* feishu-style:text-align:left */} Quickly experience the audio understanding effect by passing the audio URL into the model. The sample code is as follows. ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "input_audio", "input_audio": { "data": "https://example-files.cnbj1.mi-fds.com/example-files/audio/audio_example.wav" } }, { "type": "text", "text": "please describe the content of the audio" } ] } ], max_completion_tokens=1024 ) print(completion.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "input_audio", "input_audio": { "data": "https://example-files.cnbj1.mi-fds.com/example-files/audio/audio_example.wav" } }, { "type": "text", "text": "please describe the content of the audio" } ] } ], "max_completion_tokens": 1024 }' ``` {/* feishu-style:text-align:left */} **Response** ```json { "id": "020d7a2c-f8e7-4e5a-aaea-790ba908b51a_4e073cbfa8994cc1832055183ac841e5", "choices": [ { "finish_reason": "stop", "index": 0, "message": { "content": "The audio features a female voice saying, \"Good morning,\" followed by the question, \"Could you tell me what the weather will be like today?\"", "role": "assistant", "tool_calls": null, "reasoning_content": "- The audio starts with a female voice saying \"Good morning.\"\n- Then she asks, \"Could you tell me what the weather will be like today?\"\n- The audio is very short, just a few seconds.\n- The content is a simple greeting followed by a question about the weather.\n- I'll summarize this clearly." } } ], "created": 1790007961, "model": "mimo-v2.6-pro", "object": "chat.completion", "usage": { "completion_tokens": 98, "prompt_tokens": 86, "total_tokens": 184, "completion_tokens_details": { "reasoning_tokens": 66 }, "prompt_tokens_details": { "audio_tokens": 25, "cached_tokens": 0 } } } ``` ## Supported models {/* feishu-style:text-align:left */} Currently, the `mimo-v2.6-flash`, `mimo-v2.6-pro`, `mimo-v2.6-pro-ultraspeed` and `mimo-v2.5` models are supported. ## Audio Input method {/* feishu-style:text-align:left */} Supported audio input methods are as follows: - Audio URL Input: A publicly accessible audio URL address must be provided. - Base64 Encoding Input: Convert the audio to a Base64-encoded string before passing it in. ### Audio URL Input {/* feishu-style:text-align:left */} Audio files can be directly passed in via a publicly accessible audio URL address, which is suitable for scenarios where the audio files are already stored in a publicly accessible environment. The size of a single audio file cannot exceed 100 MB. ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "input_audio", "input_audio": { "data": "https://example-files.cnbj1.mi-fds.com/example-files/audio/audio_example.wav" } }, { "type": "text", "text": "please describe the content of the audio" } ] } ], max_completion_tokens=1024 ) print(completion.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "input_audio", "input_audio": { "data": "https://example-files.cnbj1.mi-fds.com/example-files/audio/audio_example.wav" } }, { "type": "text", "text": "please describe the content of the audio" } ] } ], "max_completion_tokens": 1024 }' ``` ### Base64 Encoding Input {/* feishu-style:text-align:left */} Convert the audio file to a Base64-encoded string and then pass it in, which is suitable for scenarios where the audio file cannot be accessed via a public network URL. The size of the converted Base64-encoded string cannot exceed 50 MB.
Please include the prefix before Base64 encoding: `data:{MIME_TYPE};base64,$BASE64_AUDIO` - `{MIME_TYPE}`: The MIME type (media type) of the audio, used to identify the audio format, which needs to be replaced with the MIME value corresponding to the actual audio. - `$BASE64_AUDIO`: A pure Base64-encoded string of the audio file (without any prefix).
> In the following example, both `{MIME_TYPE}` and `$BASE64_AUDIO` are placeholders, please replace them with actual valid data before running. ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "input_audio", "input_audio": { "data": "data:{MIME_TYPE};base64,$BASE64_AUDIO" } }, { "type": "text", "text": "please describe the content of the audio" } ] } ], max_completion_tokens=1024 ) print(completion.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "input_audio", "input_audio": { "data": "data:{MIME_TYPE};base64,$BASE64_AUDIO" } }, { "type": "text", "text": "please describe the content of the audio" } ] } ], "max_completion_tokens": 1024 }' ``` ## Audio Restrictions - Audio Formats: MP3, WAV, FLAC, M4A, OGG. > Audio Formats variants are numerous, and it cannot be guaranteed that all files can be recognized. Please verify through testing that the files can be recognized normally. - Audio Size: - When passed in as a URL: File size does not exceed 100 MB. - When passed in as Base64 encoding: The size of the Base64 encoded string of a single audio file does not exceed 50 MB. - Number of audios: When multiple audio files are input, the number of audio files is limited by the model's context length, and the total number of tokens for all audio and text must be less than the model's context length. > Note: For calculating audio tokens, please refer to [Explanation of Audio Token Usage](https://platform.xiaomimimo.com/#/docs/usage-guide/multimodal-understanding/audio-understanding?target=explanation-of-audio-token-usage). For the model context length, please refer to [Pricing and Rate Limits](https://platform.xiaomimimo.com/#/docs/pricing). ## Explanation of Audio Token Usage {/* feishu-style:text-align:left */} For the Token conversion of audio, please refer to the following code. The estimated results are for reference only, and the actual usage is subject to the API response. ```bash Total tokens ≈ Audio duration (in seconds, e.g., 10.6 seconds) * 6.25 ``` ## Price - Billing: The total cost is calculated based on the number of input, input (cache hits), and output tokens; for pricing, please refer to [Pricing and Rate Limits](https://platform.xiaomimimo.com/#/docs/pricing). - Audio Token consumption can be calculated through [Explanation of Audio Token Usage](https://platform.xiaomimimo.com/#/docs/usage-guide/multimodal-understanding/audio-understanding?target=explanation-of-audio-token-usage). The estimated results are for reference only, and the actual usage is subject to the API response. - View Bill: You can view your bill and usage on the [Billing](https://platform.xiaomimimo.com/#/console/usage) page in the Console. ## FAQ ### Does it support local file upload? {/* feishu-style:text-align:left */} `mimo-v2.6-flash`, `mimo-v2.6-pro`, `mimo-v2.6-pro-ultraspeed` and `mimo-v2.5` models don't currently support uploading local audio files. For supported upload methods, please refer to [Audio Input Method](https://platform.xiaomimimo.com/#/docs/usage-guide/multimodal-understanding/audio-understanding?target=audio-input-method). --- DOCUMENT: Video Understanding --- URL: https://mimo.mi.com/static/docs/quick-start/usage-guide/multimodal-understanding/video-understanding.md # Video Understanding {/* feishu-style:text-align:left */} The video understanding model can answer based on the video you provide, supporting both video URL and Base64 encoding as input methods, and is suitable for scenarios such as video analysis. ## Quick Start
For preparations such as obtaining an API Key, please refer to [First API Call](https://platform.xiaomimimo.com/#/docs/quick-start/first-api-call).
{/* feishu-style:text-align:left */} Quickly experience the video understanding effect by passing the model through the video URL method. The sample code is as follows. ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "video_url", "video_url": { "url": "https://example-files.cnbj1.mi-fds.com/example-files/video/video_example.mp4" }, "fps": 2, "media_resolution": "default" }, { "type": "text", "text": "please describe the content of the video" } ] } ], max_completion_tokens=1024 ) print(completion.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "video_url", "video_url": { "url": "https://example-files.cnbj1.mi-fds.com/example-files/video/video_example.mp4" }, "fps": 2, "media_resolution": "default" }, { "type": "text", "text": "please describe the content of the video" } ] } ], "max_completion_tokens": 1024 }' ``` {/* feishu-style:text-align:left */} **Response** ```json { "id": "fbfdb524-f4c8-41e6-b0a3-31520f6198c7_0b8b3ccec31e40868b76109308ed9329", "choices": [ { "finish_reason": "stop", "index": 0, "message": { "content": "The video features a small, white kitten sitting on a polished wooden floor indoors. The scene is brightly lit by natural sunlight streaming in from a window on the left, which is partially covered by a beige curtain. The light casts distinct shadows on the floorboards. To the right, a large green potted plant (resembling a Monstera) is visible in the background. The kitten, which has striking blue eyes, sits facing forward but glances slightly to the side. It briefly opens its mouth, as if meowing or yawning. The overall atmosphere is warm, calm, and sunny.", "role": "assistant", "tool_calls": null, "reasoning_content": "The user wants a description of the video content.\n\n1. **Analyze the video:**\n * **Subject:** A small, white kitten.\n * **Setting:** Indoors, near a window. There's a wooden floor. A curtain hangs to the left. A potted plant (Monstera) is visible in the background to the right.\n * **Lighting:** Bright natural sunlight is streaming in from the left (likely a window), casting strong shadows and highlights on the floor and the kitten.\n * **Action:** The kitten is sitting on the floor. It's looking slightly to its left (viewer's right). It opens its mouth briefly, possibly meowing or yawning, but there's no audio. Then it looks forward again. It's a very short clip, essentially a still shot with minor movement.\n * **Details:** Blue eyes, white fur, pink nose. The curtain is beige/white striped. The floor is polished wood.\n\n2. **Draft the description:**\n * Start with the main subject: A small white kitten.\n * Describe the location: Sitting on a wooden floor near a window.\n * Describe the lighting: Sunlight is streaming in, creating a warm atmosphere.\n * Describe the background elements: A curtain on the left, a green plant on the right.\n * Describe the kitten's features: White fur, blue eyes.\n * Describe the action: The kitten looks around and opens its mouth slightly.\n\n3. **Refine the description:**\n * \"The video features a small white kitten with blue eyes sitting on a polished wooden floor.\"\n * \"The setting appears to be a cozy room with bright sunlight streaming in through a window on the left.\"\n * \"A beige curtain hangs beside the window.\"\n * \"To the right, a green potted plant is visible in the background.\"\n * \"The kitten is facing forward but looks slightly to the side, opening its mouth briefly as if meowing or yawning.\"\n * \"The lighting creates distinct shadows on the floor.\"\n\n4. **Final Polish:** Combine into a coherent paragraph. Ensure it covers the visual aspects clearly.\n\n * *Draft:* The video shows a white kitten with bright blue eyes sitting on a shiny wooden floor. Sunlight is pouring in from a window on the left, illuminating the kitten and the floor. There is a curtain to the left and a large leafy plant in the background on the right. The kitten opens its mouth slightly and looks around. The scene is calm and well-lit.\n\n * *More detailed:* The video captures a serene moment featuring a small, white kitten sitting on a polished wooden floor bathed in natural sunlight. The light streams in from a window on the left, partially covered by a beige curtain, casting long shadows across the floor. To the right, the large green leaves of a monstera plant add a touch of nature to the indoor setting. The kitten, with its striking blue eyes, faces the camera but glances to the side, briefly opening its mouth as if letting out a soft meow or a yawn. The overall atmosphere is warm and peaceful.\n\n5. **Output generation:** Present the description clearly to the user." } } ], "created": 1790008279, "model": "mimo-v2.6-pro", "object": "chat.completion", "usage": { "completion_tokens": 807, "prompt_tokens": 1260, "total_tokens": 2067, "completion_tokens_details": { "reasoning_tokens": 684 }, "prompt_tokens_details": { "audio_tokens": 19, "cached_tokens": 1216, "video_tokens": 1144 } } } ``` ## Supported models {/* feishu-style:text-align:left */} Currently, the `mimo-v2.6-flash`, `mimo-v2.6-pro`, `mimo-v2.6-pro-ultraspeed` and `mimo-v2.5` models are supported. ## Video Input Method {/* feishu-style:text-align:left */} Supported video input methods are as follows: - Video URL Input: A publicly accessible video URL address must be provided. - Base64 Encoding Input: Convert the video to a Base64-encoded string before inputting it. ### Video URL Input {/* feishu-style:text-align:left */} Videos can be directly passed in via a publicly accessible video URL address, which is suitable for scenarios where the video is already stored in a publicly accessible environment. The size of a single video file cannot exceed 300 MB. ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "video_url", "video_url": { "url": "https://example-files.cnbj1.mi-fds.com/example-files/video/video_example.mp4" }, "fps": 2, "media_resolution": "default" }, { "type": "text", "text": "please describe the content of the video" } ] } ], max_completion_tokens=1024 ) print(completion.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "video_url", "video_url": { "url": "https://example-files.cnbj1.mi-fds.com/example-files/video/video_example.mp4" }, "fps": 2, "media_resolution": "default" }, { "type": "text", "text": "please describe the content of the video" } ] } ], "max_completion_tokens": 1024 }' ``` ### Base64 Encoding Input {/* feishu-style:text-align:left */} Convert the video file to a Base64-encoded string and then pass it in, which is suitable for scenarios where the video cannot be accessed via a public network URL. The size of the converted Base64-encoded string cannot exceed 50 MB.
Please include the prefix before Base64 encoding: `data:{MIME_TYPE};base64,$BASE64_VIDEO` - `{MIME_TYPE}`: The MIME type (media type) of the video, used to identify the video format, and needs to be replaced with the MIME value corresponding to the actual video. - `$BASE64_VIDEO`: Pure Base64-encoded string of the video file (without any prefix).
> In the following example, both `{MIME_TYPE}` and `$BASE64_VIDEO` are placeholders, please replace them with actual valid data before running. ```python import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.6-pro", messages=[ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "video_url", "video_url": { "url": "data:{MIME_TYPE};base64,$BASE64_VIDEO" }, "fps": 2, "media_resolution": "default" }, { "type": "text", "text": "please describe the content of the video" } ] } ], max_completion_tokens=1024 ) print(completion.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": [ { "type": "video_url", "video_url": { "url": "data:{MIME_TYPE};base64,$BASE64_VIDEO" }, "fps": 2, "media_resolution": "default" }, { "type": "text", "text": "please describe the content of the video" } ] } ], "max_completion_tokens": 1024 }' ``` ## Instructions for Use ### Video Restrictions - Video Formats: MP4, MOV, AVI, WMV. > Video Formats variants are numerous, and it cannot be guaranteed that all files can be recognized. Please verify through testing that the files can be recognized normally. - Video Size: - When passed in as a URL: single video file size does not exceed 300 MB. - When passed in as Base64 encoding: The size of the Base64 encoded string of a single video does not exceed 50 MB. - Number of videos: When multiple videos are input, the number of videos is limited by the model's context length, and the total number of tokens for all video and text must be less than the model's context length. > Note: For calculating video tokens, please refer to [Explanation of Video Token Usage](https://platform.xiaomimimo.com/#/docs/usage-guide/multimodal-understanding/video-understanding?target=explanation-of-video-token-usage). For the model context length, please refer to [Pricing and Rate Limits](https://platform.xiaomimimo.com/#/docs/pricing). ### Control the fineness of video understanding {/* feishu-style:text-align:left */} You can control the granularity of video understanding through the two fields `fps` and `media_resolution` respectively. 1. `fps` is the frame number of images extracted from the video per second, used to control the fineness of understanding the time dimension of the video. The default value is 2, with a range of `[0.1, 10]`. - The higher the value, the denser the frame extraction, and the more refined the model's perception of frame changes, movements, and temporal details; - The lower the value, the sparser the frame extraction, the faster the processing speed, and the less Token consumption. 1. `media_resolution` refers to the resolution level of video frames, used to control the visual understanding fineness of a single frame. The default value is `default`. - `default`: The default level, balancing recognition effectiveness and processing efficiency; - `max`: The highest resolution level, which enhances the recognition ability for small objects and detailed textures. ## Explanation of Video Token Usage {/* feishu-style:text-align:left */} Video tokens are divided into `video_tokens` (visual) and `audio_tokens` (audio). - `video_tokens` Calculation please refer to the following code. The estimated results are for reference only, and the actual usage is subject to the API response. ```python """ Estimate the number of tokens consumed by an API call based on video duration and resolution. Two parameters control the level of detail: - fps: Frames extracted per second. Default 2, range [0.1, 10]. Higher values yield finer temporal granularity at the cost of more tokens. - media_resolution: Per-frame resolution tier. "default" balances quality and efficiency; "max" improves fine-grained detail recognition. """ import math def estimate_video_tokens( duration: float, width: int, height: int, fps: float = 2.0, media_resolution: str = "default", mute: bool = False, ) -> int: """ Estimate the token count for a video input. Args: duration: Video duration in seconds. width: Video width in pixels. height: Video height in pixels. fps: Frame extraction rate. Default 2, range [0.1, 10]. media_resolution: "default" or "max". mute: If True, audio tokens are excluded. Returns: Estimated total token count. """ # ---- Constants ---- PATCH, MERGE, T_PATCH = 16, 2, 2 SPATIAL = PATCH * MERGE # 32 PIX_PER_TOKEN = SPATIAL ** 2 # 1024 MAX_TOTAL_TOKENS = 131072 TOTAL_MAX_PIX = MAX_TOTAL_TOKENS * PIX_PER_TOKEN MIN_PIX, MAX_PIX = 8192, 8388608 MAX_FRAMES = 2048 DEFAULT_MAX_FRAME_TOKEN = 300 # ---- 1. Number of extracted frames ---- nframes = math.ceil(duration * fps) nframes = min(nframes, MAX_FRAMES) nframes = max(math.ceil(nframes / T_PATCH) * T_PATCH, T_PATCH) # ---- 2. Per-frame pixel budget ---- max_pix = TOTAL_MAX_PIX * T_PATCH // nframes if media_resolution != "max": max_pix = min(max_pix, DEFAULT_MAX_FRAME_TOKEN * PIX_PER_TOKEN) max_pix = max(MIN_PIX, min(max_pix, MAX_PIX)) # ---- 3. Resolution scaling ---- h, w = height, width if min(h, w) < SPATIAL: if h < w: w = int(w * SPATIAL / h); h = SPATIAL else: h = int(h * SPATIAL / w); w = SPATIAL h_bar = round(h / SPATIAL) * SPATIAL w_bar = round(w / SPATIAL) * SPATIAL if h_bar * w_bar > max_pix: beta = math.sqrt(h * w / max_pix) h_bar = math.floor(h / beta / SPATIAL) * SPATIAL w_bar = math.floor(w / beta / SPATIAL) * SPATIAL elif h_bar * w_bar < MIN_PIX: beta = math.sqrt(MIN_PIX / (h * w)) h_bar = math.ceil(h * beta / SPATIAL) * SPATIAL w_bar = math.ceil(w * beta / SPATIAL) * SPATIAL # ---- 4. Token calculation ---- grids = nframes // T_PATCH # temporal grid count tokens_per_grid = (h_bar // PATCH) * (w_bar // PATCH) // (MERGE ** 2) vision = grids * tokens_per_grid timestamps = grids * (5 if fps > 2 else 3) # timestamp text tokens special = grids * 2 + 2 # special markers # ---- 5. Audio tokens ---- audio = 0 if not mute: spec_len = int(duration * 24000) // 240 + 1 t = (spec_len - 1) // 2 + 1 t = t // 2 + int(t % 2 != 0) audio = math.ceil(t / 4) + 2 # +2 for audio special tokens return vision + timestamps + special + audio # ============ Example ============ if __name__ == "__main__": # A 1080p, 60-second video tokens = estimate_video_tokens(duration=60, width=1920, height=1080) print(f"Default params (fps=2, default): {tokens:,} tokens") tokens = estimate_video_tokens(duration=60, width=1920, height=1080, fps=5) print(f"High frame rate (fps=5, default): {tokens:,} tokens") tokens = estimate_video_tokens(duration=60, width=1920, height=1080, media_resolution="max") print(f"High resolution (fps=2, max): {tokens:,} tokens") tokens = estimate_video_tokens(duration=60, width=1920, height=1080, mute=True) print(f"Muted (fps=2, mute): {tokens:,} tokens") ``` - `audio_tokens` Calculation please refer to the following code. The estimated results are for reference only, and the actual usage is subject to the API response. ```bash Total tokens ≈ Audio duration (in seconds) * 6.25 ``` ## Price - Billing: The total cost is calculated based on the number of input, input (cache hits), and output tokens; for pricing, please refer to [Pricing and Rate Limits](https://platform.xiaomimimo.com/#/docs/pricing). - Video Token consumption can be calculated through [Explanation of Video Token Usage](https://platform.xiaomimimo.com/#/docs/usage-guide/multimodal-understanding/video-understanding?target=explanation-of-video-token-usage). The estimated results are for reference only, and the actual usage is subject to the API response. - View Bill: You can view your bill and usage on the [Billing](https://platform.xiaomimimo.com/#/console/usage) page in the Console. ## FAQ ### Does it support local file upload? {/* feishu-style:text-align:left */} `mimo-v2.6-flash`, `mimo-v2.6-pro`, `mimo-v2.6-pro-ultraspeed` and `mimo-v2.5` models don't currently support uploading local video files. For supported upload methods, please refer to [Video Input Method](https://platform.xiaomimimo.com/#/docs/usage-guide/multimodal-understanding/video-understanding?target=video-input-method). --- DOCUMENT: Speech Recognition(MiMo-V2.5-ASR) --- URL: https://mimo.mi.com/static/docs/quick-start/usage-guide/audio/Speech-Recognition.md # Speech Recognition (MiMo-V2.5-ASR) {/* feishu-style:text-align:left */} Speech recognition converts input audio into text output, suitable for meeting transcription, lyrics recognition, dialect transcription, noisy environment recordings, and more. You can improve recognition accuracy by specifying language parameters. {/* feishu-style:text-align:left */} **Core Capabilities** - **Broad Language and Dialect Coverage**: Supports bilingual Chinese-English recognition with automatic language detection. Natively recognizes Cantonese, Wu, Minnan, Sichuan, and other Chinese dialects. - **Robust in Complex Scenarios**: Maintains stable recognition in noisy environments, far-field pickup, and multi-speaker overlapping conversations. Also supports lyrics transcription with background music. - **Precise Handling of Specialized Content**: Accurately recognizes knowledge-intensive content such as classical poetry, technical terminology, proper nouns, and place names. Automatically generates punctuation without post-processing. ## Supported Models {/* feishu-style:text-align:left */} Currently, only the `mimo-v2.5-asr` model is supported. ## Prerequisites {/* feishu-style:text-align:left */} For API Key setup and other prerequisites, please refer to [First API Call](https://platform.xiaomimimo.com/#/docs/quick-start/first-api-call). ## Supported Audio Formats {/* feishu-style:text-align:left */} Currently, only `wav` and `mp3` audio sample files are supported. Before passing audio to the API, convert the file to a Base64 encoded string. The encoded string size must not exceed 10 MB. Currently, two audio input methods are supported: ```json "input_audio": { "data": "data:{MIME_TYPE};base64,$BASE64_AUDIO" } ``` {/* feishu-style:text-align:left */} When passing audio in pure Base64 encoding, the `format` field must be passed simultaneously to specify the audio format. ```json "input_audio": { "data": "$BASE64_AUDIO", "format": "{format}" } ``` {/* feishu-style:text-align:left */} **Supported formats and their MIME types:**
Format MIME Type
wav `audio/wav`
mp3 `audio/mpeg` or `audio/mp3`
## Code Sample
Set `asr_options.language` to specify the language. Auto detection will be used if this parameter is not configured. Manual specification is recommended when the language is confirmed to improve recognition performance. Supported values: `auto`, `zh`, `en`.
### Non-streaming Call ```python import os import base64 from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) # Replace with the actual local file path with open("audio_file.wav", "rb") as f: audio_bytes = f.read() audio_base64 = base64.b64encode(audio_bytes).decode("utf-8") completion = client.chat.completions.create( model="mimo-v2.5-asr", messages=[ { "role": "user", "content": [ { "type": "input_audio", "input_audio": { "data": f"data:audio/wav;base64,{audio_base64}" } } ] } ], extra_body={ "asr_options": { "language": "en" } } ) print(completion.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header 'Content-Type: application/json' \ --data-raw '{ "model": "mimo-v2.5-asr", "messages": [ { "role": "user", "content": [ { "type": "input_audio", "input_audio": { "data": "data:{MIME_TYPE};base64,$BASE64_AUDIO" } } ] } ], "asr_options": { "language": "en" } }' ``` ### Streaming Call ```python import os import base64 from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) # Replace with the actual local file path with open("audio_file.wav", "rb") as f: audio_bytes = f.read() audio_base64 = base64.b64encode(audio_bytes).decode("utf-8") completion = client.chat.completions.create( model="mimo-v2.5-asr", messages=[ { "role": "user", "content": [ { "type": "input_audio", "input_audio": { "data": f"data:audio/wav;base64,{audio_base64}" } } ] } ], extra_body={ "asr_options": { "language": "auto" } }, stream=True ) for chunk in completion: print(chunk.model_dump_json()) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header 'Content-Type: application/json' \ --data-raw '{ "model": "mimo-v2.5-asr", "messages": [ { "role": "user", "content": [ { "type": "input_audio", "input_audio": { "data": "data:{MIME_TYPE};base64,$BASE64_AUDIO" } } ] } ], "asr_options": { "language": "auto" }, "stream": true }' ``` ## Price - Billing: Please refer to [Pay‑As‑You‑Go API](https://platform.xiaomimimo.com/#/docs/pricing). - View Bill: You can view your usage on the [Billing](https://platform.xiaomimimo.com/#/console/usage) page in the Console. --- DOCUMENT: Speech Synthesis (MiMo-V2.5-TTS Series) --- URL: https://mimo.mi.com/static/docs/quick-start/usage-guide/audio/speech-synthesis-v2.5.md # Speech Synthesis (MiMo-V2.5-TTS Series) {/* feishu-style:text-align:left */} Speech Synthesis (Text-to-Speech) supports automatically converting input text into natural and fluent speech output. You can generate natural and vivid speech content by configuring parameters such as speech style and voice. {/* feishu-style:text-align:left */} **Core Capabilities** - **Out-of-the-box built-in voices:** A variety of high-quality built-in voices are available for quick use without additional configuration. - **Voice design and cloning:** Supports voice design via text description, or replication of arbitrary voices based on audio samples. - **Diverse speech styles:** Supports control over speed, emotion, role-play, dialects and other styles, for more vivid and natural speech expression.
Low-latency streaming output for `mimo-v2.5-tts` is now available. The streaming interface has been restored and returns responses in real time.
## List of Supported Models {/* feishu-style:text-align:left */} Currently, three models of the MiMo-V2.5-TTS series are supported, and the model list is as follows:
Model ID Function Voice Precautions
`mimo-v2.5-tts` Use built-in high-quality voices for speech synthesis Use the high-quality voices from the built-in voices list Supports singing mode, does not support voice design and voice cloning
`mimo-v2.5-tts-voicedesign` Customize voice through text description Automatically generate voices from text descriptions, without requiring presets or audio samples Does not support singing mode, built-in voices, or voice cloning
`mimo-v2.5-tts-voiceclone` Replicate any voice from audio samples Precisely replicate voices from audio samples to enable speech synthesis of any voice Does not support singing mode, built-in voices, or voice design
## Preparation {/* feishu-style:text-align:left */} For preparations such as obtaining API Key, please refer to [ First API Call ](https://platform.xiaomimimo.com/#/docs/quick-start/first-api-call). ## General Precautions
**Call Rules** - The target text for speech synthesis must be filled in the `role` of `assistant` message and cannot be placed in the `user` role message. - `user` role messages are optional parameters, and instructions can be passed in to adjust the tone and style of speech synthesis, or they can be conversation history (message content will not appear in synthesized speech). When using the `mimo-v2.5-tts-voicedesign` model, they are required parameters. - When using streaming calls, please specify the format of the output audio as `pcm16`, so that it can be spliced into a complete audio. For splicing examples, please refer to the Python calling methods in each chapter.
## Style Control {/* feishu-style:text-align:left */} The instruction-following ability of the model is sufficient to cover the following complex controls (a single natural language instruction is sufficient to take effect): - **Multi-style Switching**: A single character completes the style transition from *announcement → whisper → roar* within the same voice segment, with a natural and unobtrusive transition. - **Multi-emotion Mixing**: Supports complex emotions such as "repressed anger", "smile with a sob", "gentle but tired", "gentleness in mania", etc., rather than only allowing the selection of a single emotion. - **Multi-granularity control**: From *paragraph level* (overall tone) → *sentence level* (rhythm) → *word level* (stress) → *character granularity* (choking, dragging, or breathy sound of a specific character), all can be specified in the instruction. {/* feishu-style:text-align:left */} We currently offer two control methods: **natural language control** and **tag control** . The placement of the content for both methods in ` messages ` is different: - **Natural Language Control** → Placed in `role: user`'s `content` - **Audio Tag Control** → Placed in ` role: assistant ` 's ` content ` ### Natural Language Control {/* feishu-style:text-align:left */} Through natural language description, enable the model to understand and generate speech in the corresponding style. **The content is placed in the** `messages` **field of** `role: user` **in the** `content` **field.** You can directly describe the desired speech style in a single sentence. {/* feishu-style:text-align:left */} **Example:** > Report good news to the leader in a brisk and upbeat tone, speaking at a slightly faster pace, with the uncontrollable excitement and a touch of pride after learning the results, and a bright and energetic voice. > Looking at the results of the just-solved difficult problem, couldn't help exclaiming in a self-satisfied and overjoyed manner, with a high-pitched and bright voice, a relatively fast speaking speed, and a tone full of confidence and disbelief. > With a bright and lively teenage voice, carrying the pride and playfulness after a successful prank, speaking at a relatively fast pace with light enunciation, and the tone slightly rising when emphasizing the bet. {/* feishu-style:text-align:left */} On this basis, we also support a more complex and refined **director mode** — just like writing a script for actors, comprehensively depicting characters and voices from the three dimensions of **character, scene, and guidance**, based on which the model can generate more layered and performative voices. - **[Character]** Clearly describe the character's identity, personality traits, physical appearance and speaking habits. - **[Scene]** Describe what is happening at this moment, who you are talking to, and what emotional state you are in. The more specific the better — time, location, event, and the other person's reaction can all be included. - **[Guidance]** Similar to a director giving acting instructions to an actor: speaking speed, breath control, pauses, accents, resonance position, timbre texture, and emotional fluctuations. It can be written in detail, and the model will act according to these "stage directions". {/* feishu-style:text-align:left */} **Example:** ```python Role: The current head of the century-old noble Cen family. Since birth, she was adopted and raised by the gatekeeper of the ancestral temple, molded into a flawless, emotionless family totem. She has long lived in seclusion and has a strong sense of class alienation towards others. Scene: In the shadows of the ancestral hall, she watches the man who has broken through the security cordon at all costs to find her and attempts to elope with her. She will use the coldest and most rigid class barriers to strangle both the other person and the feelings that have just sprouted but are enough to start a prairie fire within herself. Guidance: A cold, languid yet extremely imposing deep-voiced mature woman. Her vocal tract is very relaxed, without any sign of tension, yet exuding a bone-chilling sense of oppression. - Speed and Pauses: Extremely slow, with each word rolling on the tip of her tongue before being uttered, carrying the casual arrogance of a superior. There are extremely long, unsettling pauses between sentences. - Breathiness and Full Voice: Most of the time, her voice has no obvious pitch fluctuations, with a heavy and hard full voice, like a calm yet cold undercurrent. However, a very slight breathy sound must be added at certain final sounds (such as "sincerity") to reveal a hint of weariness and longing that even she herself is unaware of. - Articulation Texture: The mixed use of literary and colloquial words bears the traces of the old era, with labiodental sounds pronounced extremely lightly but extremely clearly (such as "collision" and "cheap"), making her speech both elegant and sharp, hitting home with every word. ``` {/* feishu-style:text-align:left */} Director Mode is suitable for scenarios with high requirements for voice performance, such as character voiceovers, film-level content generation, etc. ### Audio Tag Control {/* feishu-style:text-align:left */} By embedding style tags and audio tags in the text, fine-grained control over speech can be directly achieved. The overall style tag comes at the beginning, and fine-grained control tags can be inserted in the middle. **All tag control content is placed in the** ` messages ` **of the** ` role: assistant ` ` content ` **field.** {/* feishu-style:text-align:left */} Add a **start** `(style)` tag to the target text to specify the pronunciation style of the voice. Multiple styles can be set simultaneously by placing multiple style names within the same pair of parentheses, with no restrictions on the delimiter. {/* feishu-style:text-align:left */} **Supported bracket formats:** Half-width `()`, full-width `()`, or `[]` can be used. {/* feishu-style:text-align:left */} **Format Example:** `(Style 1 Style 2)Content to be Synthesized` {/* feishu-style:text-align:left */} The following are some recommended styles, and custom styles not listed are also supported.
To experience a better singing style, you must add the `(唱歌)` tag at the very beginning of the target text, with the format: `(唱歌)lyrics`. `Lyrics` are recommended to be in Chinese for better synthesis results. The tag supports equivalent values: `唱歌`, `sing`, `singing`. Entering any of them will yield the same result.
**Style Type** **Style Example**
Basic Emotions *Happy / Sad / Angry / Fearful / Amazed / Excited / Wronged / Calm / Indifferent*
Complex Emotions *Melancholy / Relieved / Helpless / Guilty / Relieved / Jealous / Tired / Apprehensive / Emotional*
Overall tone *Gentle / Cold / Lively / Serious / Lazy / Playful / Deep / Capable / Sharp*
Timbre Positioning *Magnetic / Mellow / Clear / Ethereal / Innocent / Old / Sweet / Hoarse / Elegant*
Character Tone *Clamp voice / Big Sister voice / Shota voice / Uncle voice / Taiwanese accent*
Dialect *Northeast dialect / Sichuan dialect / Henan dialect / Cantonese*
Role-playing *Sun Wukong / Lin Daiyu*
Singing *singing*
{/* feishu-style:text-align:left */} **Example:** - `(Sighing)After all these years, when I walked down that street again, a part of my heart suddenly felt empty.` - `(Lazy)Let me sleep for five more minutes... just five minutes, really, for the last time.` - `(Magnetic)The night is already deep, but the city is still breathing. I'm the one accompanying you tonight. Welcome to listen to .` - `(Northeastern dialect)Oh my goodness, it's so cold today! You know that wind, it's whistling like a knife, cutting into your face!` - `(Cantonese)This is really amazing! Once you've tasted it, you won't forget!` - `(singing)Forgive me for my unruly and unrestrained love for freedom throughout my life, and I'm also afraid that one day I'll fall, Oh no. Abandoning ideals, anyone can do it, so how could I be afraid that one day it'll only be you and me.` {/* feishu-style:text-align:left */} On this basis, we also support inserting `[audio tag]` at any position in the text. Through the [audio tag], you can perform fine-grained control over the sound, precisely adjusting tone, mood, and expression style—whether it's a whisper, a hearty laugh, or a little complaint with a touch of emotion. You can also flexibly insert breathing sounds, pauses, coughs, etc., all of which can be easily achieved. The speaking speed can also be flexibly adjusted, allowing each sentence to have its proper rhythm.
**Style Type** **Style Example**
Speech Rate and Rhythm *Inhale / Take a deep breath / Sigh / Let out a long sigh / Pant / Hold one's breath*
Emotional State *nervous / scared / excited / tired / wronged / coquettish / guilty / shocked / impatient*
Speech Features *Trembling / Voice trembling / Pitch change / Cracked voice / Nasal voice / Breathiness / Hoarseness*
Laughing and crying tone *Smile / Chuckle / Laugh out loud / Sneer / Sob / Whimper / Choke / Wail*
{/* feishu-style:text-align:left */} **Example:** - (nervously, takes a deep breath) Hoo... Calm down, calm down. It's just an interview... (speaking faster, muttering) I've rehearsed my self-introduction fifty times, it should be okay. Come on, you can do it... (softly) Oh, is my tie crooked? - (extremely exhausted, listless) Master... wake me up when we get there... (sighs deeply) I'll take a little nap first. This overtime has made me feel like my soul is about to scatter. - If I had... (pauses for a moment) even if I had persisted for just one more second, would the outcome have been different? (forced smile) Oh, there are no "what ifs" anymore. - (Rapid breathing due to the cold) Hoo—hoo—This, this snow in the Greater Khingan Mountains... (cough) It can literally freeze one's bones... Don't, don't stop, keep moving, move quickly. - (raising voice and shouting) Sister! This fish is fresh! Just caught this morning! Hey! You there, stop rummaging around! If you crush it, you'll have to pay for it! ## Speech Synthesis Using Built-in Voices - It comes with multiple high-quality voices and can be used directly without additional configuration. Currently, only the `mimo-v2.5-tts` model is supported - Supports controlling the style of synthetic speech by passing natural language instructions in the user message - Supports controlling the style of synthesized speech through audio tags ### Built-in Voice List {/* feishu-style:text-align:left */} When in use, you can set the preset timbre in `{"audio": {"voice": "mimo_default"}}`.
**Voice Name** **Voice ID** Language Gender
MiMo-默认 mimo_default It varies depending on the deployed cluster. The default for the China cluster is `冰糖`, and the default for other clusters is `Mia`
冰糖 冰糖 Chinese Female
茉莉 茉莉 Chinese Female
苏打 苏打 Chinese Male
白桦 白桦 Chinese Male
Mia Mia English Female
Chloe Chloe English Female
Milo Milo English Male
Dean Dean English Male
### Code Sample #### Non-streaming Call ```python import os from openai import OpenAI import base64 client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.5-tts", messages=[ { "role": "user", "content": "Bright, bouncy, slightly sing-song tone — like you're bursting with good news you can barely hold in. Fast pace, rising pitch at the end." }, { "role": "assistant", "content": "Hey boss — guess what, guess what? I just got the results back and I actually passed! Not just passed, I got a distinction! I know, I know — you told me I was cutting it close, but hey, here we are. Drinks are on me tonight, okay?" } ], audio={ "format": "wav", "voice": "Chloe" } ) message = completion.choices[0].message audio_bytes = base64.b64decode(message.audio.data) with open("audio_file.wav", "wb") as f: f.write(audio_bytes) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header 'Content-Type: application/json' \ --data-raw '{ "model": "mimo-v2.5-tts", "messages": [ { "role": "user", "content": "Bright, bouncy, slightly sing-song tone — like you are bursting with good news you can barely hold in. Fast pace, rising pitch at the end." }, { "role": "assistant", "content": "Hey boss — guess what, guess what? I just got the results back and I actually passed! Not just passed, I got a distinction! I know, I know — you told me I was cutting it close, but hey, here we are. Drinks are on me tonight, okay?" } ], "audio": { "format": "wav", "voice": "Chloe" } }' ``` #### Streaming Call
Low-latency streaming output for `mimo-v2.5-tts` is now available. The streaming interface has been restored and returns responses in real time.
```python import base64 import os import numpy as np import soundfile as sf from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.5-tts", messages=[ { "role": "user", "content": "Bright, bouncy, slightly sing-song tone — like you're bursting with good news you can barely hold in. Fast pace, rising pitch at the end." }, { "role": "assistant", "content": "Hey boss — guess what, guess what? I just got the results back and I actually passed! Not just passed, I got a distinction! I know, I know — you told me I was cutting it close, but hey, here we are. Drinks are on me tonight, okay?" } ], audio={ "format": "pcm16", "voice": "Chloe" }, stream=True ) # 24kHz PCM16LE mono audio collected_chunks: np.ndarray = np.array([], dtype=np.float32) for chunk in completion: if not chunk.choices: continue delta = chunk.choices[0].delta audio = getattr(delta, "audio", None) if audio is not None: assert isinstance(audio, dict), f"Expected audio to be a dict, got {type(audio)}" pcm_bytes = base64.b64decode(audio["data"]) np_pcm = np.frombuffer(pcm_bytes, dtype=np.int16).astype(np.float32) / 32768.0 collected_chunks = np.concatenate((collected_chunks, np_pcm)) print(f"Received audio chunk of size {len(pcm_bytes)} bytes") # Save the collected audio to a file os.makedirs("tmp", exist_ok=True) sf.write("tmp/output.wav", collected_chunks, samplerate=24000) print("Audio saved to tmp/output.wav") ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header 'Content-Type: application/json' \ --data-raw '{ "model": "mimo-v2.5-tts", "messages": [ { "role": "user", "content": "Bright, bouncy, slightly sing-song tone — like you are bursting with good news you can barely hold in. Fast pace, rising pitch at the end." }, { "role": "assistant", "content": "Hey boss — guess what, guess what? I just got the results back and I actually passed! Not just passed, I got a distinction! I know, I know — you told me I was cutting it close, but hey, here we are. Drinks are on me tonight, okay?" } ], "audio": { "format": "pcm16", "voice": "Chloe" }, "stream": true }' ``` ## Speech Synthesis Using Voice Design {/* feishu-style:text-align:left */} There is no need to provide an audio file. Simply add voice description text to the message with the role of `user`, and a customized voice can be generated. Currently, only the `mimo-v2.5-tts-voicedesign` model is supported. ### How to Write a Good Voice Design Prompt {/* feishu-style:text-align:left */} When using the `mimo-v2.5-tts-voicedesign` model, the text in the `user` message is the voice design description. The more specific and vivid the description, the closer the generated voice will be to the expected one. #### Key Dimension {/* feishu-style:text-align:left */} A good voice description usually covers the following multiple dimensions (not necessarily comprehensive):
Dimension Example
Gender and Age "young woman in her mid-20s", "middle-aged man in his 50s"
Voice / Texture "deep and gravelly", "silky, mellow, and magnetic"
Mood / Tone "warm and confident", "gentle but with a hint of weariness"
Speech speed / Rhythm "slow and deliberate", "speaking at an extremely fast pace, like a machine gun."
{/* feishu-style:text-align:left */} The following dimensions can be optionally added to increase richness: - **Role / Character**: narrator, podcast host, storyteller, late-night radio DJ - **Speaking style**: casual and colloquial, seriously, lowering one's voice as if plotting - **Scene description**: narrating a nature documentary, during a roadshow for investors - **Era reference**: 1940s film noir, dubbed voices of translated films from the 1980s #### Writing Suggestions {/* feishu-style:text-align:left */} **Concise descriptive** -- quickly outline the sound profile using keywords or a single sentence ```bash Heavy Russian accent, gruff middle-aged male, blunt and matter-of-fact. ``` {/* feishu-style:text-align:left */} **Professional Descriptive** -- Three-dimensional portrayal of sound through scenarios, character design, or multi-dimensional details ```bash Young female, extreme close-up with a binaural, ear-to-ear ASMR feel. Audible breathing, subtle swallowing, and soft natural lip sounds. She speaks very slowly, creating a deeply relaxing and immersive experience. ``` ```json An elderly gentleman, speaking Mandarin with a northern accent, his speech slow and steady, his voice slightly hoarse and weathered, as if an old and seasoned grandfather were telling a story, full of the wisdom of years. ``` #### Precautions - **Length**: 1-4 sentences are sufficient; there's no need to write a long text. Clearly describing the core features is more important than piling up dimensions - **Avoid conflicts**: Do not simultaneously request contradictory characteristics (e.g., "innocent childish voice + CEO aura") - **Avoid using audio quality effect terms**: Do not write descriptions related to post-processing such as reverb, echo, EQ, compression, etc - **Avoid vague words**: Do not use descriptions lacking specific references such as "ordinary," "normal," or "foreign" - **Both Chinese and English are supported**: the model supports both Chinese and English voice timbre descriptions, so choose the language in which you can express most precisely - **Synthetic text should match the voice tone**: The synthetic text in the `assistant` message should match the voice tone description to achieve the best results. For example, pair a goodnight monologue with a "gentle and soothing female voice" instead of a passionate sports commentary. It is recommended to use LLM to automatically generate matching synthetic text based on your voice tone description; on the Studio page, you can directly click the "Generate Text" button after entering the voice tone description. ### Code Sample
`mimo-v2.5-tts-voicedesign` supports the optional parameter `optimize_text_preview` to control whether the target broadcast text is intelligently polished. When set to `true`, the `assistant` role message can be omitted.
#### Non-streaming Call ```python import os from openai import OpenAI import base64 client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.5-tts-voicedesign", messages=[ { "role": "user", "content": "Give me a young male tone." }, { "role": "assistant", "content": "Yes, I had a sandwich." } ], audio={ "format": "wav", "optimize_text_preview": True } ) message = completion.choices[0].message audio_bytes = base64.b64decode(message.audio.data) with open("audio_file.wav", "wb") as f: f.write(audio_bytes) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header 'Content-Type: application/json' \ --data-raw '{ "model": "mimo-v2.5-tts-voicedesign", "messages": [ { "role": "user", "content": "Give me a young male tone." }, { "role": "assistant", "content": "Yes, I had a sandwich." } ], "audio": { "format": "wav", "optimize_text_preview": true } }' ``` #### Streaming Call
**Note** - The low-latency streaming output feature of `mimo-v2.5-tts-voicedesign` is not yet available. If you have relevant requirements, please follow the upcoming feature updates. - The streaming call interface is currently downgraded to compatibility mode, and only **returns the results once in streaming format after all inferences are completed.**
```python import base64 import os import numpy as np import soundfile as sf from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1" ) completion = client.chat.completions.create( model="mimo-v2.5-tts-voicedesign", messages=[ { "role": "user", "content": "Give me a young male tone." }, { "role": "assistant", "content": "You are UN-BE-LIEVABLE! I am sooooo done with your constant lies. GET. OUT!" } ], audio={ "format": "pcm16", "optimize_text_preview": True }, stream=True ) # 24kHz PCM16LE mono audio collected_chunks: np.ndarray = np.array([], dtype=np.float32) for chunk in completion: if not chunk.choices: continue delta = chunk.choices[0].delta audio = getattr(delta, "audio", None) if audio is not None: assert isinstance(audio, dict), f"Expected audio to be a dict, got {type(audio)}" pcm_bytes = base64.b64decode(audio["data"]) np_pcm = np.frombuffer(pcm_bytes, dtype=np.int16).astype(np.float32) / 32768.0 collected_chunks = np.concatenate((collected_chunks, np_pcm)) print(f"Received audio chunk of size {len(pcm_bytes)} bytes") # Save the collected audio to a file os.makedirs("tmp", exist_ok=True) sf.write("tmp/output.wav", collected_chunks, samplerate=24000) print("Audio saved to tmp/output.wav") ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header 'Content-Type: application/json' \ --data-raw '{ "model": "mimo-v2.5-tts-voicedesign", "messages": [ { "role": "user", "content": "Give me a young male tone." }, { "role": "assistant", "content": "You are UN-BE-LIEVABLE! I am sooooo done with your constant lies. GET. OUT!" } ], "audio": { "format": "pcm16", "optimize_text_preview": true }, "stream": true }' ``` ## Speech Synthesis Using Voice Cloning - By passing in audio samples, you can accurately replicate the target timbre and generate speech. Currently, only the `mimo-v2.5-tts-voiceclone` model is supported - Supports controlling the style of synthetic speech by passing natural language instructions in the user message - Supports controlling the style of synthesized speech through audio tags ### Code Sample {/* feishu-style:text-align:left */} Convert the audio file sample to a Base64-encoded string and then pass it in. The size of the converted Base64-encoded string cannot exceed 10 MB, and currently only `mp3` and `wav` format audio sample files are supported.
Please include the prefix before Base64 encoding: `data:{MIME_TYPE};base64,$BASE64_AUDIO` - `{MIME_TYPE}`: The MIME type (media type) of the audio, used to identify the audio format, needs to be replaced with the MIME value corresponding to the actual audio. The values here can be: ` audio/mpeg ` (or ` audio/mp3 `), ` audio/wav `. - `$BASE64_AUDIO`: A pure Base64-encoded string of the audio file (without any prefix).
> In the following example, both `{MIME_TYPE}` and `$BASE64_AUDIO` are placeholders, please replace them with actual valid data before running. #### Non-streaming Call ```python import base64 import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1", ) with open("voice.mp3", "rb") as f: voice_bytes = f.read() voice_base64 = base64.b64encode(voice_bytes).decode("utf-8") completion = client.chat.completions.create( model="mimo-v2.5-tts-voiceclone", messages=[ { "role": "user", "content": "" }, { "role": "assistant", "content": "Yes, I had a sandwich." } ], audio={ "format": "wav", "voice": f"data:audio/mpeg;base64,{voice_base64}" } ) message = completion.choices[0].message audio_bytes = base64.b64decode(message.audio.data) with open("audio_file.wav", "wb") as f: f.write(audio_bytes) ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header 'Content-Type: application/json' \ --data-raw '{ "model": "mimo-v2.5-tts-voiceclone", "messages": [ { "role": "user", "content": "" }, { "role": "assistant", "content": "Yes, I had a sandwich." } ], "audio": { "format": "wav", "voice": "data:{MIME_TYPE};base64,$BASE64_AUDIO" } }' ``` #### Streaming Call
**Note** - The low-latency streaming output feature of `mimo-v2.5-tts-voiceclone` is not yet available. If you have relevant requirements, please follow the upcoming feature updates. - The streaming call interface is currently downgraded to compatibility mode, and only **returns the results oncein streaming formatafter all inferences are completed.**
```python import base64 import os import numpy as np import soundfile as sf from openai import OpenAI client = OpenAI( api_key=os.environ.get("MIMO_API_KEY"), base_url="https://api.xiaomimimo.com/v1", ) with open("voice.mp3", "rb") as f: voice_bytes = f.read() voice_base64 = base64.b64encode(voice_bytes).decode("utf-8") completion = client.chat.completions.create( model="mimo-v2.5-tts-voiceclone", messages=[ { "role": "user", "content": "" }, { "role": "assistant", "content": "Yes, I had a sandwich." } ], audio={ "format": "wav", "voice": f"data:audio/mpeg;base64,{voice_base64}", }, stream=True ) # 24kHz PCM16LE mono audio collected_chunks: np.ndarray = np.array([], dtype=np.float32) for chunk in completion: if not chunk.choices: continue delta = chunk.choices[0].delta audio = getattr(delta, "audio", None) if audio is not None: assert isinstance(audio, dict), ( f"Expected audio to be a dict, got {type(audio)}" ) pcm_bytes = base64.b64decode(audio["data"]) np_pcm = np.frombuffer(pcm_bytes, dtype=np.int16).astype(np.float32) / 32768.0 collected_chunks = np.concatenate((collected_chunks, np_pcm)) print(f"Received audio chunk of size {len(pcm_bytes)} bytes") # Save the collected audio to a file os.makedirs("tmp", exist_ok=True) sf.write("tmp/output.wav", collected_chunks, samplerate=24000) print("Audio saved to tmp/output.wav") ``` ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header 'Content-Type: application/json' \ --data-raw '{ "model": "mimo-v2.5-tts-voiceclone", "messages": [ { "role": "user", "content": "" }, { "role": "assistant", "content": "You are UN-BE-LIEVABLE! I am sooooo done with your constant lies. GET. OUT!" } ], "audio": { "format": "pcm16", "voice": "data:{MIME_TYPE};base64,$BASE64_AUDIO" }, "stream": true }' ``` ## Price - Billing: Free for a limited time. - View Bill: You can view your usage on the [Billing](https://platform.xiaomimimo.com/#/console/usage) page in the Console. --- DOCUMENT: Refer & earn --- URL: https://mimo.mi.com/static/docs/quick-start/promotions/refer.md # Refer & earn ## 1. Overview {/* feishu-style:text-align:left */} The Xiaomi MiMo Open Platform Refer & Earn program is a long-term referral activity. Any registered user can share a 6-character invite code or invite link with friends. Once a friend redeems the code, a referral relationship is created. When the friend completes their first paid Token Plan order, the inviter earns an additional 10% of the paid amount in credits, and the friend gets 10% off on the first order. ## 2. Eligibility #### 2.1 Inviter {/* feishu-style:text-align:left */} Any registered Xiaomi MiMo Open Platform user is eligible to invite. #### 2.2 Invitee {/* feishu-style:text-align:left */} Users who signed up to the Xiaomi MiMo Open Platform within the last 3 days. The following accounts are not eligible as invitees: - Accounts older than 3 days - Accounts that have already acted as an inviter or invitee - Mutual invites (if A invited B, B cannot redeem A's code in return) - Your own other accounts ## 3. Reward rules #### 3.1 Credits back on paid orders - **Trigger:** Friend completes their first paid order (definition in §3.2) - **Reward:** - Inviter earns an additional 10% of the friend's paid amount in credits - Friend gets 10% off on the first order (stacks with any existing first-order discount) - Special bonus: Standard/Pro annual and Max monthly/annual plans receive a 20% referral bonus with no 30-transaction cap. - **Crediting:** To make sure rewards are paid out correctly, credits are released 3 days after your friend's first paid order (if the friend refunds within 3 days, the credits won't be paid out) #### 3.2 First-paid-order definition - **First order:** The invitee's first paid Token Plan subscription within the lifetime of their account (does not include: API top-ups) - **Paid amount:** Actual cash payment after discounts (i.e. the amount charged to bank card / third-party payment) - **30-day first-order window:** The friend must complete the first paid order within 30 days of binding the invite code; otherwise the 10% paid-order reward will not be issued. The 10% first-order discount follows the same 30-day window: first orders completed after 30 days no longer get the discount. - **Free Token Plan exclusion:** If the friend has already received 1 month or more of free Token Plan, their first order no longer gets 10% off, and the purchase does not count toward referral rewards (no credits back will be issued) ## 4. Limits and quotas {/* feishu-style:text-align:left */} To prevent abuse, the following two limits apply per account: #### 4.1 Bindings: unlimited - No friend limit: inviter can invite any number of friends to bind - After 30 bindings, new invitees still get 10% off on Token Plan first order - 10% off has no cap #### 4.2 Credits-back quota: 30 paid orders - First 30 paid orders earn 10% credits back; beyond 30, no further credits back, but friends still get 10% off ## 5. Using your rewards #### 5.1 Credits redemption scope {/* feishu-style:text-align:left */} Credits can offset your Xiaomi MiMo model API call charges. #### 5.2 Not redeemable for - Token Plan subscription purchase #### 5.3 Credit validity - **Valid for 40 days:** Counted from the credited date; unused balance after 40 days expires automatically - **Deduction order:** Earliest-expiring first (FIFO) · credits before cash - Non-cashable, non-transferable, no change given #### 5.4 Refund handling - **Credits back (refund within 3-day pending window):** The credits will not be issued - **Credits back (refund after credited):** The unused portion of the inviter's credits is immediately voided - **Refunds do not affect counts:** The 30-paid-order cap is counted by “paid order completed,” and counts are not reversed by refunds ## 6. Violations {/* feishu-style:text-align:left */} To keep things fair, the following will result in rewards being withheld or revoked: - Inviting your own other accounts - Mass registration via bots, virtual numbers, or other technical means - Misleading promotion or fraudulent referrals - Any other violation of these rules > **About account identity:** Xiaomi MiMo identifies accounts belonging to the same user using account bindings, login environment, network signals, and other signals. ## 7. Adjustments and sunset {/* feishu-style:text-align:left */} **This is a long-term feature with no fixed end date.** Xiaomi MiMo may adjust, pause, or sunset this feature based on operational needs. ## 8. Contact us {/* feishu-style:text-align:left */} For support or business inquiries: - Email support-mimo@xiaomi.com ## 9. Final interpretation {/* feishu-style:text-align:left */} Within the limits of applicable law, Xiaomi MiMo holds final interpretation rights for these rules. {/* feishu-style:text-align:left */} Update Time August 13, 2026 --- DOCUMENT: Account and Authentication --- URL: https://mimo.mi.com/static/docs/quick-start/faq/account.md # Account and Authentication ### Domestic/Overseas - **What are the differences between overseas users and domestic users?** {/* feishu-style:text-align:left */} Main differences: {/* feishu-style:text-align:left */} 1) Different payment methods (domestic supports WeChat Pay, Alipay, and Xiaomi Pay; overseas payments are made through the Waffo Console (settled in US dollars $)). {/* feishu-style:text-align:left */} 2) Return different Base URLs + Keys based on the region where the account is located, which are not interoperable. {/* feishu-style:text-align:left */} 3) The overseas version currently has no invoicing function. - **Can domestic and overseas usage be aggregated for calculation?** {/* feishu-style:text-align:left */} No, it is not allowed. Different Base URLs and Keys are returned based on the region where the account is located, and usage is calculated separately. {/* feishu-style:text-align:left */}

### What login methods are supported? {/* feishu-style:text-align:left */} The platform uses a Xiaomi account to log in. If you have already registered a Xiaomi account, you can log in directly; if you do not have a Xiaomi account, you can visit the Console to register, or register in advance at [id.mi.com](https://id.mi.com/): - Registration Methods for Mainland Chinese Users - Mobile Number Registration - Third-party account authorized registration (supports registration via WeChat, Weibo, QQ, Alipay, and Apple ID; to ensure account security, a third-party account may be used to bind a commonly used mobile phone number) - Registration Methods for Overseas Users - Email Registration - Third-party account authorized registration (supports Google/Facebook registration; to ensure account security, the email bound to the third-party account may be used as the security email for the Xiaomi account) {/* feishu-style:text-align:left */}

### Unable to log in to the account? {/* feishu-style:text-align:left */} If you have trouble logging in, please visit [ Xiaomi Account Help Center ](https://account.xiaomi.com/helpcenter?_locale=zh_CN) to view more account issues and solutions. {/* feishu-style:text-align:left */}

### Does recharge/purchase of a package require real-name authentication? {/* feishu-style:text-align:left */} Domestic users need to complete personal or corporate real-name authentication before they can recharge or purchase a package. Both types of authentication can be used for purchase. {/* feishu-style:text-align:left */}

### What are the differences between personal certification and enterprise certification? {/* feishu-style:text-align:left */} Real-name authentication is divided into personal authentication and enterprise authentication, and the same account only supports authentication as one type. After completing real-name authentication, a personal account can be changed to an enterprise account through enterprise authentication, while an enterprise account cannot be changed to a personal account. {/* feishu-style:text-align:left */} **The payment process for enterprise accounts is exactly the same as that for personal accounts,** and enterprise accounts also support WeChat/Alipay/Xiaomi Pay. {/* feishu-style:text-align:left */}

### Why does it prompt that "the account has abnormal usage behavior" and cannot continue to be used? {/* feishu-style:text-align:left */} It may be because the system captured certain keywords during detection, triggering the automatic protection mechanism. We highly value the experience of each user and will conduct regular manual reviews to minimize the occurrence of misjudgments. If you encounter this issue, you can provide your user ID, and we will verify it for you and assist with the unblocking as soon as possible. {/* feishu-style:text-align:left */}

### Can I still continue to use the model service after canceling my Xiaomi account? {/* feishu-style:text-align:left */} After the Xiaomi account is cancelled, the platform data will be cleared, and the API Key will expire synchronously. Please carefully evaluate and operate with caution. If you have any questions, please contact us support-mimo@xiaomi.com . --- DOCUMENT: Payment --- URL: https://mimo.mi.com/static/docs/quick-start/faq/payment.md # Payment ### How to recharge on the Open Platform? - **Pay-as-you-go API:** Go to the [ Account Balance ](https://platform.xiaomimimo.com/#/console/balance) page to top up. The platform provides three top-up methods for domestic users: Xiaomi Pay, Alipay, and WeChat Pay; for overseas users, it provides **commonly used top-up methods such as Apple Pay, Google Pay, credit/debit cards, etc.** Top-ups generally arrive in real-time, and you can check the balance in the [ Account Balance ](https://platform.xiaomimimo.com/#/console/balance), and view the cumulative top-up amount and top-up records on the [ Top-up Details ](https://platform.xiaomimimo.com/#/console/recharge) page. - **Token Plan:** The plan does not currently support deduction using account balance or bonus, and you need to go to [Subscribe to Token Plan](https://platform.xiaomimimo.com/token-plan) to purchase separately. {/* feishu-style:text-align:left */}

### Does purchasing a Token Plan package count towards cumulative recharge? {/* feishu-style:text-align:left */} Subscriptions are not counted, and orders for subscription packages are not included in cumulative recharge. {/* feishu-style:text-align:left */}

### What payment methods are supported? {/* feishu-style:text-align:left */} Domestically in China, it supports WeChat Pay, Alipay, and Xiaomi Pay, while overseas, it uses the Waffo payment gateway for payment (settled in US dollars $). {/* feishu-style:text-align:left */}

### Is there a time limit for payment? {/* feishu-style:text-align:left */} The effective payment duration is subject to the display on the page. After the timeout, the order will automatically close, and you need to place a new order. {/* feishu-style:text-align:left */}

### How to set up balance alerts? {/* feishu-style:text-align:left */} On the [Account Balance](https://platform.xiaomimimo.com/#/console/balance) page, you can enable the balance alert setting. Once enabled, when the account balance falls below the alert threshold, we will send a notification to your registered mobile number/email. Please check it carefully. {/* feishu-style:text-align:left */}

### Does it support refund applications? - **Pay-as-you-go API:** Account balance supports full refund. If you have a refund request, you can open the contact pop-up window via the "Apply for Refund" button in the upper right corner of [Recharge Details](https://platform.xiaomimimo.com/#/console/recharge) , select the "Refund" option, and state the reason to initiate a refund request. The account balance will be refunded back to the original payment method after the review is approved (the consumed amount, the invoiced amount, and the platform-gifted amount cannot be refunded). After the refund request is accepted, you will no longer be able to continue calling the model service and will not be able to continue recharging. Billing may be delayed, and the refund amount shall be subject to the actual amount received. Refunds are generally processed and returned to the original payment method within 3-5 working days. - **Token Plan:** Once a package is paid, it cannot be refunded. {/* feishu-style:text-align:left */}

### How to issue an invoice? - **Chinese User:** {/* feishu-style:text-align:left */} Visit the [Invoice](https://platform.xiaomimimo.com/#/console/invoice) page, select the successfully recharged order, and issue an electronic invoice. Both individual and corporate invoices can be issued. Fill in your email or mobile phone number, and the invoice will be sent via email/sms upon completion of issuance. {/* feishu-style:text-align:left */} Note: - The amount eligible for invoicing is the actual payment amount. Platform coupons, discounts, or amounts gifted by the platform cannot be invoiced. Refunded amounts cannot be invoiced. - According to regulations, invoices issued in the name of individuals can only be issued as digital electronic ordinary invoices, while invoices issued in the name of enterprises can be issued as digital electronic ordinary invoices or digital electronic special invoices. - We generally issue invoices within 48 hours after receiving the application, and delays may occur in case of special circumstances. - Invoices support red stamping. If the invoice has been deducted or recorded, you need to log in to the Electronic Tax Bureau and confirm the information (red stamping confirmation form) within 72 hours to successfully complete the red stamping. After red stamping, the original order can be reissued with an invoice. - The invoicing entity is: Beijing Xiaomi Mobile Software Co., Ltd. - **Overseas Users:** {/* feishu-style:text-align:left */} Each recharge order will automatically generate an invoice. When you complete a recharge, you can view the invoice on the order page. You can also enter the [Recharge Details](https://platform.xiaomimimo.com/#/console/recharge) page to download historical invoices. {/* feishu-style:text-align:left */}

### Can I still call the API if my balance is insufficient? {/* feishu-style:text-align:left */} Before the billing system goes live, models can be called for inference services normally when the balance is 0. {/* feishu-style:text-align:left */} After the billing system goes live, due to a certain time delay, the balance may be ≤ 0. Once the balance becomes negative, the model inference service can no longer be used, and the next recharge order will first deduct the overdue amount. {/* feishu-style:text-align:left */}

### Will there still be charges after the API Key is deleted? - **Real-time inference API**: If the API Key is deleted, it will no longer be able to call the interface, and no charges will be incurred. The historical consumption records of this API Key can still be queried in [Billing](https://platform.xiaomimimo.com/#/console/usage). - **Batch inference API**: If the API Key is deleted, batch tasks already in execution will not be interrupted, and the billing will still be attributed to this API Key. The historical billing records of this API Key can still be viewed in [Billing](https://platform.xiaomimimo.com/#/console/usage). --- DOCUMENT: API Integration --- URL: https://mimo.mi.com/static/docs/quick-start/faq/api-integration.md # API Integration ### How to obtain API Key? - **API Key for Pay-as-you-go API Calls:** After logging in to the Xiaomi MiMo API Open Platform, apply for an API Key on the [Console - API Keys](https://platform.xiaomimimo.com/#/console/api-keys) page. When using the model via API, please include your API Key in the request header:` api-key: $MIMO_API_KEY ` or ` Authorization: Bearer $MIMO_API_KEY `. - **API Key of Token Plan:** After successful purchase, you can see the exclusive API Key on the [Token Plan ](https://platform.xiaomimimo.com/#/console/plan-manage)page. **Note: API Key is only visible and can be copied when created, please save it properly.** - **The API Key format for Token Plan is** `tp-xxxxx`**, which is only used for Token Plan subscription services; the API Key format for pay-as-you-go API calls is** `sk-xxxxx`**, used for pay-as-you-go billing. The two are independent of each other and cannot be mixed. The API Key for Token Plan is only available within the validity period of the Token Plan package you have subscribed to.** {/* feishu-style:text-align:left */}

### What if the API Key is lost or leaked? {/* feishu-style:text-align:left */} can be reset on the [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) page. {/* feishu-style:text-align:left */}

### How can I obtain the Base URL of the Token Plan? {/* feishu-style:text-align:left */} The Base URL provided on the Token Plan page shall prevail: Two types of Base URLs are provided, one compatible with the OpenAI interface protocol and the other compatible with the Anthropic interface protocol, which can be copied and used as needed. {/* feishu-style:text-align:left */}

### Which programming tools does Token Plan support? {/* feishu-style:text-align:left */} Supports mainstream programming tools and model frameworks, such as Claude Code, OpenClaw, OpenCode, Kilo Code, Cline, Hermes Agent, CodeBuddy Code, etc. For specific access methods, please refer to [Overview of AI Tools](https://platform.xiaomimimo.com/#/docs/integration/tools-overview). {/* feishu-style:text-align:left */}

### Can Token Plan be used in multiple programming tools at the same time? {/* feishu-style:text-align:left */} The same package can be used across all supported tools, but the quota is shared, and usage of all tools will consume the same package quota. {/* feishu-style:text-align:left */}

### What's the difference between OpenAI and Anthropic interfaces? - OpenAI interface `/v1/chat/completions` follows OpenAI format, including developer/system/user/assistant roles - Anthropic interface `/anthropic/v1/messages` follows Claude format, with a separate system parameter {/* feishu-style:text-align:left */}

### How to make multi-turn tool calls in thinking mode? {/* feishu-style:text-align:left */} During the multi-turn tool calls process in thinking mode, the model returns a `reasoning_content` field alongside `tool_calls`. To continue the conversation, it is recommended to keep all previous `reasoning_content` in the `messages` array for each subsequent request to achieve the best performance. {/* feishu-style:text-align:left */} The requested example is as follows: ```bash curl --location --request POST 'https://api.xiaomimimo.com/v1/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "messages": [ { "role": "assistant", "content": "Hello! I am MiMo.", "reasoning_content": "Okay, the user just asked me to introduce myself. That is a pretty straightforward request, but I should think about why they are asking this." }, { "role": "user", "content": "What is the weather like in Hebei?" } ], "model": "mimo-v2.6-pro", "max_completion_tokens": 1024, "temperature": 1.0, "stream": false, "tools": [ { "type": "function", "function": { "name": "get_current_weather", "description": "Get the current weather in a given location", "parameters": { "type": "object", "properties": { "location": { "type": "string", "description": "The city and state, e.g. San Francisco, CA" }, "unit": { "type": "string", "enum": [ "celsius", "fahrenheit" ] } }, "required": [ "location" ] } } } ], "tool_choice": "auto" }' ``` {/* feishu-style:text-align:left */}

### Why aretool_calls sometimes included in the reasoning_content field and sometimes in a separate tool_calls field? {/* feishu-style:text-align:left */} The appearance of `tool_calls` in the reasoning content indicates instability and incomplete output caused by the model having `thinking` enabled when calling `tool`. It is recommended to disable `thinking` when calling `tool` calls and to adjust the settings according to [Model Hyperparameters](https://platform.xiaomimimo.com/#/docs/quick-start/model-hyperparameters) to achieve a more stable and better user experience. {/* feishu-style:text-align:left */}

### What's the response speed? {/* feishu-style:text-align:left */} Response speed depends on: - Request length and complexity - Server load and geographic location - Whether streaming response is used {/* feishu-style:text-align:left */}

### How to handle timeouts? {/* feishu-style:text-align:left */} Please implement reasonable timeout handling on the client side: - Set reasonable connection and read timeout times - Use exponential backoff for retries - For long responses, it's recommended to use streaming mode {/* feishu-style:text-align:left */}

### What if the API returns inappropriate content? {/* feishu-style:text-align:left */} The platform has added content review for both user input and model output. If violations occur, the returned content will be automatically intercepted to ensure the content you receive is safe. {/* feishu-style:text-align:left */}

### Why doesn’t the model perform a web search after enabling online search? {/* feishu-style:text-align:left */} There may be three reasons: - **Cache**: There is a 5-minute cache period after enabling / disabling online search. The online search switch will not take effect immediately within 5 minutes. - **Model determines no need for search**: The model judges that the current query does not involve real-time information and can be answered directly with its own knowledge. To force a search, set `forced_search: true`. - **Only some models are supported**: Currently only `mimo-v2.6-flash`, `mimo-v2.6-pro`, `mimo-v2.6-pro-ultraspeed`, `mimo-v2.5-pro` and `mimo-v2.5` supports web search. {/* feishu-style:text-align:left */}

### Does it support local file upload? - Real-time Inference, local file upload is not supported at this time. - Batch Inference, supporting the upload of local files (in JSONL format). --- DOCUMENT: Plans and Pricing --- URL: https://mimo.mi.com/static/docs/quick-start/faq/token-plan/Plans&Pricing.md # Plans and Pricing - **What packages are available for the Token Plan?** {/* feishu-style:text-align:left */} We currently offer four subscription plans: Lite (available for individual users only), Standard, Pro, and Max. Both monthly and annual auto-renewal subscriptions are supported, while team plans additionally support one-time purchases for custom durations (1/3/6/12 months). Each plan comes with distinct usage quotas and benefits. - **Which models does Token Plan support?** {/* feishu-style:text-align:left */} Both the individual and team versions support a total of **6** **models, including the MiMo-V2.6 series**, which are accessible across all subscription tiers. > mimo-v2.6-pro: Flagship Full-Modal Inference Model > > mimo-v2.6-flash: An Ultra-Full-Modal Inference Model with Extreme Speed > > mimo-v2.5-asr: Speech Recognition Model > > mimo-v2.5-tts/ v2.5-tts-voiceclone / v2.5-tts-voicedesign: Speech synthesis model (limited-time free)
**The Individual edition additionally supports mimo-v2.5-pro and mimo-v2.5.Both of these two models will be officially taken offline at 10:00 on October 21, 2026 Beijing Time, and it is recommended to switch to the new version of the models as soon as possible.**
- **How much is the package deal? Are there any discounts available?** {/* feishu-style:text-align:left */} The specific prices of the four packages are subject to those displayed on the landing page. The platform currently offers the following time-limited promotional activities: - First-purchase Discount: Users purchasing the Individual Edition for the first time can enjoy a 12% discount, with only one discount eligible per account; the Team Edition is not eligible for this discount. - Annual auto-renewal plan: Enjoy a 12% discount compared to the monthly auto-renewal plan; first-purchase offers are not applicable to the annual plan. - Night discount rate: During off-peak hours (Beijing time 00:00–08:00, i. e. UTC 16:00–24:00), the Credits consumption coefficient is 0.8x. - **Is continuous subscription supported?** {/* feishu-style:text-align:left */} Supported. Two subscription cycles are available: monthly auto-renewal and annual auto-renewal, with automatic renewal upon expiration. You may cancel the auto-renewal at any time on the Token Plan page. The annual auto-renewal plan offers a 12% discount, which is a more favorable deal compared to the monthly auto-renewal plan. - **Can I purchase multiple packages or upgrade my package?** {/* feishu-style:text-align:left */} **You can purchase both the Individual Plan and the Team Plan at the same time, but only one subscription per plan type can be active simultaneously.** - **Individual plan: If you wish to get more credits before your current plan expires**, you can convert the Credit quota you have already used into the equivalent amount, then pay the price difference on that basis to upgrade to a higher plan and get more Credits. Cross-tier plan upgrade by paying the price difference is supported, while plan downgrade is not allowed. If you have already upgraded to the top-tier Max plan, further upgrade is unavailable. **After your current plan expires, you can repurchase a plan of any tier.** > Price difference = New package price -(Remaining balance of original package / Total amount of original package) * Original package price - **Team Plan:** When the monthly quota for a seat is exhausted, that seat will be suspended from service, without affecting other seats that still have remaining quota. If you need to continue using the service, you can contact the team owner to purchase additional seats as needed, or switch to the pay-as-you-go standard API service. Plan upgrades and downgrades are temporarily not supported. A team-shared usage package is coming soon, stay tuned. --- DOCUMENT: Validity and Expiry --- URL: https://mimo.mi.com/static/docs/quick-start/faq/token-plan/Validity&Expiry.md # Validity and Expiry - **How long is the package valid after purchase?** {/* feishu-style:text-align:left */} It takes effect immediately upon purchase, and is valid for one calendar month/year from the date of purchase based on UTC time. {/* feishu-style:text-align:left */} For example, if you subscribe to a certain monthly plan on March 28, the plan will expire at 23:59:59 UTC on April 28. - **Will the subscription be automatically renewed after it expires?** {/* feishu-style:text-align:left */} It depends on whether you have enabled auto-renewal. If auto-renewal is enabled, upon the expiration of your plan, the system will initiate automatic deduction via Alipay/WeChat/Mi Pay (for domestic users) or via waffo/Stripe (for overseas users). After successful deduction, your subscription will automatically enter the next billing cycle without any manual operation required; if auto-renewal is not enabled, the service will be suspended once your plan expires, and you will need to manually resubscribe. - **My credit limit has been used up but it hasn't expired yet. Can I still use it?** {/* feishu-style:text-align:left */} No. {/* feishu-style:text-align:left */} **Individual Edition**: When either "expiration" or "all Credits limit is used up" is met, the service will be stopped. The system will not continue to consume your bonus or account balance. We support package upgrade function. Regardless of your package consumption, we support automatically converting your remaining Credit into an equivalent amount. You can upgrade the package by making up the difference to obtain more Credits. If you need to continue using it, please upgrade the package or switch to the pay-as-you-go API. {/* feishu-style:text-align:left */} **Team Edition:** After a single seat quota is exhausted, the seat will not be available for the month. But it does not affect the use of other seats that have not been exhausted. You can wait for quota reset, add purchase and reassign seats, or switch to normal API. - **My subscription plan has expired before I've used up its benefits. Can the remaining balance be rolled over?** {/* feishu-style:text-align:left */} No. The service will be terminated immediately upon the expiration of your current plan. You will need to subscribe to a new plan separately, and any unused quota from your old plan will not be carried over to the new one. - **Will I receive a reminder when my data plan expires or is about to run out?** {/* feishu-style:text-align:left */} Yes. If auto-renewal is enabled for your plan, you will receive renewal reminders via SMS, email, and in-app messages on payment apps (WeChat/Alipay) five days before the expiration date; if auto-renewal is not enabled for your plan, you will receive reminders via SMS and email two days before the expiration date and on the expiration date itself. - **Will I get a reminder when my package quota is almost used up?** {/* feishu-style:text-align:left */} Yes. You will receive SMS and email alerts when your current plan usage reaches 50%, 90% and 100% respectively. --- DOCUMENT: Usage and Quota --- URL: https://mimo.mi.com/static/docs/quick-start/faq/token-plan/Usage&Quota.md # Usage and Quota - **Are the quotas of different models consumed independently?** {/* feishu-style:text-align:left */} The available models in the package are consumed in parallel at different ratios, rather than independently. All TTS models are free for a limited time and do not consume package credits. ASR models deduct credits based on the duration of the input audio (the duration is counted to the nearest second and finally converted to hours for statistics). The table below lists the credit deduction per token for **Cache, input and output** of each model. {/* feishu-style:text-align:left */} Language Model Language Model
model Input (Cache Hit) Token Input (cache miss) Token Output Token
mimo-v2.6-pro 2.5 Credits 300 Credits 600 Credits
mimo-v2.6-flash 2 Credits 100 Credits 200 Credits
mimo-v2.5-pro 2.5 Credits 300 Credits 600 Credits
mimo-v2.5 2 Credits 100 Credits 200 Credits
{/* feishu-style:text-align:left */} ASR Model
model Input audio duration (h)
mimo-v2.5-asr 30M Credits
{/* feishu-style:text-align:left */} The TTS series models are available for free for a limited time and will not consume your package credits. - **What is the off-peak 0.8x coefficient** {/* feishu-style:text-align:left */} To balance resource pressure and pass on benefits to users, the Credit consumption coefficient will be 0.8 times when using the model during off-peak hours (Beijing time 00:00-08:00, i. e. UTC 16:00-24:00). {/* feishu-style:text-align:left */} For example, in a scenario where you use the mimo-v2.5-pro model and consume 10M Credits during peak hours, only 8M Credits will be consumed during off-peak hours. - **What should I do if I'm worried that Credits will be consumed quickly since the usage is calculated based on Tokens?** - We recommend that you check your historical token usage in each AI Agent framework before making a purchase, and select an appropriate plan based on your actual usage. - We have designed a Progress Bar system, and you can view the progress and make advance planning in the Token Plan. - **For the Individual plan**, we offer a package upgrade feature. Regardless of your current package usage, we will automatically convert your remaining Credits into an equivalent monetary value, allowing you to upgrade your plan by paying the price difference to get more Credits. **For the Team plan**, you can add seats mid-subscription. Meanwhile, we are also developing the shared usage package feature, so stay tuned. - **What is the remaining value of a service package?** {/* feishu-style:text-align:left */} When you renew or upgrade your current unused Individual Plan, the system will calculate the equivalent value based on the Credits consumption of the current plan, and this remaining value will be used to offset part of the payment for the new plan. Please note that for users who participated in the platform's Credits giveaway activities before August 25,2026, and each subsequent package payment amount is less than ¥7 (domestic China) / 1$(overseas), the remaining value will be calculated based on your last actual payment, which can be viewed in the renewal/upgrade pop-up window for details. - **Why are there still compensation Credits on my Token Plan page?** {/* feishu-style:text-align:left */} During the process of renewing your current plan, since the remaining value of your previous plan is higher than the value of the current plan, the platform will compensate you with Credits equivalent to the value of the price difference. --- DOCUMENT: Team Edition Benefits Related --- URL: https://mimo.mi.com/static/docs/quick-start/faq/token-plan/Team-Benefits.md # Team Edition Benefits Related - **What is the difference between the Token Plan Individual Edition and Team Edition?** {/* feishu-style:text-align:left */} The Team version supports seat-based purchase, unified payment, and centralized management of members and usage, making it ideal for teams or enterprises requiring multi-person collaboration. It maintains the same pricing as the Individual version and inherits all the benefits of the Individual version(with the exception of the two models mimo-v2.5-pro and mimo-v2.5). - **How to invite members? Is bulk invitation supported?** {/* feishu-style:text-align:left */} Members shall first register a Xiaomi account, and provide the account owner with their full name, Xiaomi account ID, as well as a mobile phone number or email address for receiving notifications. At least one of the mobile phone number and email address must be filled in; if both are provided, notifications will be sent via both SMS and email simultaneously. The owner may invite members manually, or download the template, fill in the required information and then send bulk invitations. - **Why didn't the member receive the notification after I sent the invitation?** {/* feishu-style:text-align:left */} Please confirm that the mobile phone number or email address you filled in when receiving the invitation is correct and can normally receive notifications, and check the SMS interception records or the junk email folder. If you still have not received it, the owner can copy the unexpired invitation link of this member and directly provide it to the member. - **Why can't I join the team after opening the invitation link?** {/* feishu-style:text-align:left */} Invitation is valid for 7 days from the date of sending. The invitation link will become invalid if the invitation expires, the member is removed by the owner, or the team subscription expires. Please check whether the Xiaomi ID of the logged-in account matches the invited ID, and whether the account's region matches the team. - **What is the difference between an owner and an ordinary member? Does the owner need a seat for Individual use?** {/* feishu-style:text-align:left */} The owner is responsible for purchasing plans, inviting and managing members, assigning or revoking seats, and can view the team's overall usage, member usage and transaction information. Regular members can only view their own seats, call information and usage details. {/* feishu-style:text-align:left */} When the owner invokes the model, seats also need to be allocated, and the quota of the allocated seats will be consumed. - **I have already joined the team, but why is the API Key still unavailable?** {/* feishu-style:text-align:left */} Once a seat is revoked, a member is removed, or the plan expires, the Team Edition API Key will no longer be valid. Please ensure that you have accessed the correct team workspace, been assigned a seat with remaining quota, and that the plan has not expired, then use your exclusive API Key for this workspace and the Base URL corresponding to the applicable protocol. - **What should I do if the seat usage is insufficient? Will the quota exhaustion affect other members?** {/* feishu-style:text-align:left */} Team packages are billed independently per seat. Once the monthly quota for a given seat is exhausted, that seat will be suspended from service, while other seats with remaining quota will remain unaffected. {/* feishu-style:text-align:left */} If you wish to continue using the service, you may switch to the pay-as-you-go standard API service. **The team-shared quota package feature is currently under development, stay tuned**. - **If I purchase an annual plan or buy multiple months at one time, will I get all the quota at once?** {/* feishu-style:text-align:left */} No. Whether it is monthly auto-renewal, annual auto-renewal, or one-time purchase of multiple months for the Team Plan, the seat quota will be reset on a monthly basis. {/* feishu-style:text-align:left */} Standard, Pro, and Max tiers offer per seat per month respectively **11 billion, 38 billion, and 82 billion Credits**. Please check the specific reset time on the seat page. - **Can the unused quota this month be carried over to next month?** {/* feishu-style:text-align:left */} No. The quota of Team Edition seats is reset on a monthly basis, and any unused quota in the current month will not be carried over to the next month. This rule also applies to annual continuous subscription plans and multi-month one-time purchase plans. The specific reset time is subject to the display on the seat page. - **After a member leaves, can the seat be allocated to another person? Will the quota be recalculated?** {/* feishu-style:text-align:left */} Yes. The owner may reclaim a member seat, or assign a seat that automatically reverts upon member removal to another member. {/* feishu-style:text-align:left */} When reallocating a consumed seat, the new member inherits the remaining quota of that seat for the current billing cycle, the consumed portion will not be reset, and historical consumption will be retained in the seat usage record. - **How are additional seats purchased mid-subscription charged? When will the new seats expire?** {/* feishu-style:text-align:left */} The owner may purchase additional seats of the same tier during the validity period of the current plan, with a minimum purchase of 1 seat per transaction. The amount for additional purchase shall be calculated according to the following formula: > **Additional purchase amount = (Remaining billing duration ÷ Total billing duration of the current period) × Price per seat billing cycle × Quantity of additional purchases** {/* feishu-style:text-align:left */} The quota for the first usage cycle of the newly added seats shall also be converted based on the remaining validity period, and the expiration time shall be consistent with that of the current package. After the additional purchase is successful, the owner can assign the seats to members for use, and the specific amount and quota shall be subject to the display on the additional purchase page. - **Is it possible to upgrade or downgrade the package, or unsubscribe from some of the seats?** {/* feishu-style:text-align:left */} During the validity period of the current plan, changing the plan tier or unsubscribing from individual or partial seats is temporarily unavailable. If you need additional usage resources, you may purchase extra seats for the current tier as needed. Once purchased, the subscription service takes effect immediately; no refunds are supported, and unused quotas are non-refundable. - **Can I turn off auto-renewal? What will happen after the subscription expires?** {/* feishu-style:text-align:left */} Monthly and annual auto-renewal subscriptions support automatic renewal, which can be turned off by the subscriber at any time. {/* feishu-style:text-align:left */} After the plan expires, team seats will become unassigned, the Team Edition API Key will no longer be usable, and pending invitation links will also become invalid; historical usage details will still be retained. Owners can repurchase via the "Subscribe Now" entry in the plan overview. - **How to make payment for a large amount?** {/* feishu-style:text-align:left */} There are certain payment limits for monthly and annual auto-renewal subscriptions, subject to the display on the relevant page. If you receive a limit-exceeded prompt when attempting to sign up for auto-renewal, it is recommended to switch to one-time purchase. The actual payment shall still be subject to the limits set by your selected payment channel and issuing bank. - **I already have an individual plan. Can I still use the team plan? Is it possible to join multiple teams?** {/* feishu-style:text-align:left */} Yes. Individual plan and team plan can take effect at the same time, using their own API Key and quota respectively. You can also join multiple teams and use the seat rights allocated to the corresponding team; When calling, please configure the exclusive API Key in the corresponding team space. --- DOCUMENT: Token Plan and Desktop Membership --- URL: https://mimo.mi.com/static/docs/quick-start/faq/token-plan/desktop-guide.md # Token Plan and Desktop Membership ## Overview {/* feishu-style:text-align:left */} Token Plan and Xiaomi MiMo Desktop Membership are two independent subscription services that differ in terms of application scope, supported models, benefits and usage quotas: - **Token Plan:** A subscription plan that allows users to access the MiMo model via a dedicated API Key on Desktop and other supported AI tools, available in both individual and team editions. - **Desktop Membership:** A membership service exclusive to the Desktop version, covering the usage quotas and benefits of the corresponding plan. It is suitable for users who directly use the MiMo large language model on the desktop to complete various tasks.
**The benefits, pricing, and quotas of the two subscriptions are independent of each other.** Purchasing a Token Plan does not automatically activate a Desktop membership, nor does purchasing a Desktop membership entitle you to a Token Plan package. **Users who have subscribed to the Token Plan can use the Desktop function normally after completing the access configuration, and there is no need to subscribe to an additional Desktop membership.** Token Plan package quota will be consumed during use, and the mimo-v2.6-pro-ultraspeed model is not supported for the time being.
## Rights and Scope of Application
Comparison Item Token Plan Xiaomi MiMo Desktop Membership
Subscription Content Call the specified MiMo model services, package quotas and benefits in supported tools Member services, package quotas and benefits within Desktop
Purchase Discounts
  • First-purchase discount: Get 12% off your first Personal plan purchase, once per account. Team plans are not eligible.
  • Annual subscriptions: Both Personal and Team plans offer 12% off compared with monthly subscriptions. The first-purchase discount does not apply to annual subscriptions.
  • Off-peak usage rate: During off-peak hours (00:00–08:00 Beijing time, or 16:00–24:00 UTC), usage is charged at a 0.8× Credits consumption rate.
  • Annual subscriptions: Offer 12% off compared with monthly subscriptions.

Subject to the actual display on the page.
Scope of Application Compatible with supported AI tools including Desktop, OpenClaw, Claude Code, Codex, etc. Member benefits and quotas are only available for use within the Desktop app
Usage in Desktop Configure a dedicated API Key for the Token Plan and the corresponding service address via custom model settings Log in to the subscribed membership Xiaomi account and use the corresponding membership services
Desktop Features All functions can be used normally after access All function properly
Model Equity Subject to the supported model list announced in the Token Plan Subject to the model benefits announced for each membership tier on Desktop
UltraSpeed Model Not supported yet Currently included in the "Premium" and "Elite" membership tiers, but not included in the "Basic" and "Intermediate" membership tiers
Credit Utilization Consumes Token Plan Credits, and multiple integrated tools share the package quota Consume the usage quota corresponding to the Desktop membership
Whether it includes another subscription Excludes Desktop membership Excluding the Token Plan package
{/* feishu-style:text-align:left */} For details on supported plans and models, please refer to the [Token Plan Subscription Guide ](https://mimo.mi.com/docs/zh-CN/tokenplan/Token%20Plan/subscription)and [Desktop Membership Plan](https://mimo.xiaomimimo.com/pricing/). ## How to choose - Token Plan - In addition to Desktop, you want to use the MiMo model in other supported tools (such as Codex, ClaudeCode, OpenCode, OpenClaw) - Desktop Membership - Hope to use it directly in Desktop without configuring an API Key - Hope to use UltraSpeed in Desktop ## Frequently Asked Questions - **I have purchased the Token Plan, but why does the Desktop app still prompt that my membership has not been activated?** {/* feishu-style:text-align:left */} The membership status in Desktop refers to **the Desktop membership subscription status**. The Token Plan is a standalone model subscription package, and purchasing it will not activate the Desktop membership. {/* feishu-style:text-align:left */} If you want to use the Token Plan in Desktop, please configure the exclusive credentials for the package via custom models. After integration, you can use Desktop functions normally without subscribing to an additional Desktop membership. {/* feishu-style:text-align:left */} For configuration instructions, please refer to [MiMo Desktop Access Guide](https://mimo.mi.com/docs/zh-CN/tokenplan/integration/mimo-desktop). - **Are there any feature restrictions when accessing Desktop via the Token Plan?** {/* feishu-style:text-align:left */} All features of Desktop are fully functional. The scope of supported models is subject to the model list specified in the Token Plan; mimo-v2.6-pro-ultraspeed is not temporarily supported. Usage will consume the quota of your Token Plan package, which is shared with other integrated tools. - **Why can't I use UltraSpeed after purchasing the Token Plan? Can I use it by upgrading my Token Plan tier?** {/* feishu-style:text-align:left */} None of the current Token Plan tiers include **mimo-v2.6-pro-ultraspeed**, and upgrading your Token Plan tier will not grant you access to this model. If you wish to use UltraSpeed via a Desktop membership, you currently need to subscribe to the "Premium" or "Elite" Desktop membership tier. - **Can the Desktop membership quota be used for other AI tools?** {/* feishu-style:text-align:left */} No. The benefits and quota of Desktop membership are only valid within Xiaomi MiMo Desktop, and do not include the quota of the Token Plan package for other tools. {/* feishu-style:text-align:left */} If you need to call the MiMo model in other supported AI tools, you can select the Token Plan and complete the configuration according to the access guide of the corresponding tool. --- DOCUMENT: Promotions --- URL: https://mimo.mi.com/static/docs/quick-start/faq/promotions.md # Promotions ### What promotions and benefits are currently available? {/* feishu-style:text-align:left */} The main benefit currently offered is bonus credit, which can be used to offset pay-as-you-go API usage and quickly try MiMo models. You can receive bonus credit through **Refer & Earn** or new user registration. Eligibility, credit amounts, and payout rules vary by campaign. Bonus credit is currently valid for 40 days and expires after that. {/* feishu-style:text-align:left */}

### How do I join "Refer & Earn"? {/* feishu-style:text-align:left */} Any registered user can invite; new users (signed up within 3 days) can be invited. After your friend enters via your invite link and the code is bound, only the referral relationship is created. When your friend completes their first paid order, you earn an additional 10% of their paid amount in credits and they get 10% off. {/* feishu-style:text-align:left */}

### Where is my invite code? How do I share it? {/* feishu-style:text-align:left */} After signing in, use the "Invite friends" feature to view your 6-character invite code. You can copy the invite link or save the invite poster to share with friends. {/* feishu-style:text-align:left */}

### How does my friend redeem the code? {/* feishu-style:text-align:left */} After your friend clicks your invite link and signs in to the console, the system will automatically bind the invite code (no manual entry required). Once the binding succeeds, only the referral relationship is created. If the invite code is not filled in automatically after opening the link, your friend can manually enter the 6-character code in the console to bind. Each invitee can redeem only 1 invite code. The binding cannot be changed afterward. {/* feishu-style:text-align:left */}

### When will I earn credits back? {/* feishu-style:text-align:left */} Binding an invite code only creates the referral relationship. If your friend later completes a first paid order, Tier 2 (Credits back) kicks in — you earn 10% of their paid amount. The invitee receives an additional 10% off on top of the existing first-order offer. Eligible credits are issued 3 days after the order is completed. {/* feishu-style:text-align:left */}

### What counts as a "first paid order"? {/* feishu-style:text-align:left */} A "first paid order" is your friend's first paid Token Plan subscription within the lifetime of their account. Includes: Token Plan subscription first payment Does not include: API top-ups; unpaid orders 30-day first-order window: Your friend must complete their first paid order within 30 days of binding the invite code; after 30 days, credits back will no longer be issued. The 10% first-order discount follows the same 30-day window: first orders completed after 30 days no longer get the discount. Free Token Plan exclusion: If your friend has already received 1 month or more of free Token Plan, their first order no longer gets 10% off, and the purchase does not count toward referral rewards (no credits back will be issued). A refund does not reset first-order eligibility. {/* feishu-style:text-align:left */}

### Is there a cap on credits back? {/* feishu-style:text-align:left */} Credits-back quota: 30 paid orders — first 30 paid orders earn 10% credits back. Beyond 30 orders, the inviter no longer earns credits back, while invitees can still receive 10% off. Special bonus: Standard annual, Pro annual, and Max (any duration) members earn 20% credits back on paid orders, with no 30-order cap. {/* feishu-style:text-align:left */}

### What can credits be used for? {/* feishu-style:text-align:left */} Credits can only offset Xiaomi MiMo model API call charges. Redeemable for: API call charges Not redeemable for: Token Plan subscription Non-cashable, non-transferable, no change given Valid for 40 days: counted from the credited date, earliest-expiring first (FIFO); unused balance after 40 days expires automatically. {/* feishu-style:text-align:left */}

### Does my friend's refund affect my rewards? {/* feishu-style:text-align:left */} Credits back (refund within 3-day pending window): the credits will not be issued Credits back (refund after credited): the unused portion of the inviter's credits is immediately voided Refunds do not affect counts: the 30-paid-order cap is counted by "paid order completed", and counts are not reversed by refunds {/* feishu-style:text-align:left */}

### My friend redeemed my code, but I don't see the reward. {/* feishu-style:text-align:left */} Check in order: 1. Confirm the binding succeeded: a successful binding only creates the referral relationship. 1. Credits back arrive 3 days after the friend's first paid order: please wait 3 days 1. Check the 30-day first-order window: your friend must complete the first paid order within 30 days of binding; after 30 days, no credits back are issued 1. Check the cap: After 30 paid referrals, the 10% bonus stops — unless you are on Standard annual, Pro annual, or Max monthly/annual plan. 1. If the invitee has already received at least 1 month of free Token Plan, the first paid order does not receive 10% off or qualify for credits back. {/* feishu-style:text-align:left */}

### What should I do if the invite code cannot be bound? {/* feishu-style:text-align:left */} Check whether the invite code is correct, the invitee is using their own code, the two accounts already have a reverse referral relationship, or the invitee registered within the last 3 days. If the page says the account is restricted, email support-mimo@xiaomi.com. If there is a network error or the page is unresponsive, try again later. --- DOCUMENT: Others --- URL: https://mimo.mi.com/static/docs/quick-start/faq/others.md # Others ## Model Deprecation {/* feishu-style:text-align:left */} With the continuous iteration of the MiMo model, the new version has comprehensively outperformed the old version in terms of effectiveness and performance. We will gradually deprecate the legacy models, and the specific plan will be announced in advance via SMS, email, website announcements, etc. Please pay attention to the relevant messages and complete the switch in a timely manner.For information on models that have been or will be deprecated, please refer to [Model Deprecation](https://platform.xiaomimimo.com/docs/zh-CN/updates/deprecate). ## Model Capabilities and More Usage Methods ### How is the experience of the Xiaomi MiMo model? 1. Ordinary users can directly experience online through the official website **Xiaomi MiMO Studio (** [**https://aistudio.xiaomimimo.com**](https://aistudio.xiaomimimo.com/) **)** 1. Developers can obtain resources through Xiaomi's official channels: 1. GitHub repository ( [https://github.com/xiaomimimo/MiMo-V2-Flash](https://github.com/xiaomimimo/MiMo-V2-Flash) ) to obtain open-source model weights and code 2. Xiaomi MiMo Open Platform ( [https://platform.xiaomimimo.com/](https://platform.xiaomimimo.com/) ) Apply for API Service 1. It should be noted that Xiaomi has not yet launched an official MiMo standalone app. Downloading related software from unofficial channels poses security risks, so please do not download it. {/* feishu-style:text-align:left */}

### Does the model support local deployment? How can I perform local deployment? {/* feishu-style:text-align:left */} The large model released this time has been open-sourced, please see: https://github.com/XiaomiMiMo/MiMo-V2-Flash. Local deployment primarily supports sglang, for related issues please refer to: https://github.com/sgl-project/sglang/pull/15207. {/* feishu-style:text-align:left */}

### Why does the message "Server is busy, please try again later" always appear during the conversation? {/* feishu-style:text-align:left */} Currently, MiMo's output is affected by various factors such as server load, question content, and model output. You can start a new conversation or try regenerating to get a better communication experience. {/* feishu-style:text-align:left */}

### When will the model support the function of uploading files? {/* feishu-style:text-align:left */} Support will be provided as soon as possible, please wait patiently. ## Contact Us {/* feishu-style:text-align:left */} If you need any assistance or business consultation, please feel free to contact us at any time. - Email sent to [support-mimo@xiaomi.com](mailto:support-mimo@xiaomi.com) - Scan the QR code in the lower left corner to join the developer communication group - Fill out [ Contact Us ](https://platform.xiaomimimo.com/#/contact) Questionnaire Feedback --- DOCUMENT: Rate Limit --- URL: https://mimo.mi.com/static/docs/api/guidance/rate-limit.md # Rate Limit {/* feishu-style:text-align:left */} This page lists all the models currently supported by the Xiaomi MiMo API Open Platform and their rate-limiting quotas, helping you plan your request frequency before integration. ### Rate Limiting Instructions {/* feishu-style:text-align:left */} The platform sets a model concurrency limit for each account. When the server load is high, response delays or ` 429 ` error may occur. We recommend that you plan your request frequency reasonably and implement request retry and backoff strategies in high-concurrency scenarios to avoid triggering rate limits.
- **RPM (Requests Per Minute)** : The maximum number of requests initiated per minute. The calculation scope is the sum of the total number of requests from all API Keys under a single account when calling the same model. - **TPM (Tokens Per Minute)** : The maximum number of Tokens that can be interacted with per minute. The calculation scope is the sum of the total number of requested Tokens for all API Keys under a single account when calling the same model.
### Text Generation Model
Model ID RPM TPM
`mimo-v2.6-pro` 100 10M
`mimo-v2.6-flash` 100 10M
`mimo-v2.6-pro-ultraspeed` Customized services available, please [contact us](https://platform.xiaomimimo.com/contact?userId=2451887661)
`mimo-v2.5-pro`(to be deprecated) 100 10M
`mimo-v2.5`(to be deprecated) 100 10M
### Automatic Speech Recognition Model (ASR)
Model ID RPM TPM
`mimo-v2.5-asr` 100 10K
### Text-to-Speech (TTS) Model
Model ID RPM TPM
`mimo-v2.5-tts` 100 10M
`mimo-v2.5-tts-voiceclone` 100 10M
`mimo-v2.5-tts-voicedesign` 100 10M
--- DOCUMENT: Model Hyperparameters --- URL: https://mimo.mi.com/static/docs/api/guidance/model-hyperparameters.md # Model Hyperparameters {/* feishu-style:text-align:left */} `temperature` represents the sampling temperature. Higher values (such as 0.8) will make the output more random, while lower values (such as 0.2) will make the output more deterministic. {/* feishu-style:text-align:left */} `top_p` represents the probability threshold for nucleus sampling, used to control the diversity of text generated by the model. The higher the value, the greater the diversity of the generated text.
In thinking mode, the `mimo-v2.6-flash`, `mimo-v2.6-pro`, `mimo-v2.6-pro-ultraspeed`, `mimo-v2.5-pro` and `mimo-v2.5` models do not support customizing the `temperature` and `top_p` parameters. Even if these parameters are passed in, the actual effective values will be forcibly set by the model to its recommended default values of `1.0` and `0.95`.
{/* feishu-style:text-align:left */} The default values and parameter ranges of `temperature` and `top_p` for different models are as follows:
**Model Name** **temperature** **top_p**
`mimo-v2.6-flash`
  • Default value: 1.0
  • Range: [0, 1.5]
  • Default value: 0.95
  • Range: [0.01, 1.0]
`mimo-v2.6-pro`
  • Default value: 1.0
  • Range: [0, 1.5]
  • Default value: 0.95
  • Range: [0.01, 1.0]
`mimo-v2.6-pro-ultraspeed`
  • Default value: 1.0
  • Range: [0, 1.5]
  • Default value: 0.95
  • Range: [0.01, 1.0]
`mimo-v2.5-pro`(to be deprecated)
  • Default value: 1.0
  • Range: [0, 1.5]
  • Default value: 0.95
  • Range: [0.01, 1.0]
`mimo-v2.5`(to be deprecated)
  • Default value: 1.0
  • Range: [0, 1.5]
  • Default value: 0.95
  • Range: [0.01, 1.0]
--- DOCUMENT: Error Codes --- URL: https://mimo.mi.com/static/docs/api/guidance/error-codes.md # Error Codes {/* feishu-style:text-align:left */} When using API calls to the MiMo model, common error codes and solutions are as follows:
**Error Code** **Causes** **Solutions**
400 - Invalid Format Invalid request format
  • Check if the JSON format is correct
  • Check if all required parameters are included
  • Check if parameter values are within the valid range
  • Check if the message format meets the interface requirements
  • Check if the model exists
  • Check if the fields are entered correctly
  • Check multimodal file input for compliance with format, size and other restrictions.
  • Check if multimodal file input is publicly accessible
  • In multi-turn conversations under thinking mode, the `reasoning_content` field must be fully passed back to the API.
401 - Authentication Fails
  • Missing or invalid API Key, or incorrect Authorization request header format
  • API Key that mixes Token Plan and Pay-as-you-go API
  • Check if the API key and request header format are correct
  • Check if a dedicated Base URL and API Key are used when using the Token Plan
402 - Insufficient Balance Insufficient account balance Check your account balance and recharge in a timely manner
403 - Forbidden Access The service is currently not available in the current region, or the API Key has been restricted by risk control Create a new API Key and pay attention to the security of input content
404 - Not Found The requested endpoint or model does not support image input capability Verify that the model / endpoint being used supports image input capability
421 - Content Filter Content moderation and blocking Avoid entering unsafe or sensitive content
429 - Too Many Requests Requests are too frequent, or the quota of Token Plan has been exhausted
  • Implement exponential backoff and retry logic, or reduce the request frequency
  • Upgrade the Token Plan package or switch to pay-as-you-go API
500 - Server Error Our server encounters an issue Please try again later, or contact us for resolution
503 - Server Overloaded The server is overloaded due to high traffic Please try again later
--- DOCUMENT: OpenAI Chat Completion API --- URL: https://mimo.mi.com/static/docs/api/chat/openai-api.md # OpenAI Chat Completions API Compatibility
Please refer to [Speech Synthesis (MiMo-TTS Series) - OpenAI API Compatibility](https://mimo.mi.com/docs/en-US/api/audio/tts) for the protocol documentation of the MiMo-TTS series models.
## Request Address ```bash https://api.xiaomimimo.com/v1/chat/completions ``` ## Request Headers {/* feishu-style:text-align:left */} The API supports the following two authentication methods. Please choose one and add it to the request headers: ```json api-key: $MIMO_API_KEY Content-Type: application/json ``` ```json Authorization: Bearer $MIMO_API_KEY Content-Type: application/json ``` ## Request body text is supported.", "children": [ { "name": "Text content part", "type": "object", "isBold": false, "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The text content." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content part." } ] } ] } ] }, { "name": "role", "type": "string", "isBold": true, "required": true, "description": "The role of the message author.
Available options: developer" }, { "name": "name", "type": "string", "isBold": true, "required": false, "description": "An optional name for the participant. Provides the model information to differentiate between participants of the same role." } ] }, { "name": "System message", "type": "object", "isBold": false, "description": "Developer-provided instructions that the model should follow, regardless of messages sent by the user.", "children": [ { "name": "content", "type": [ "string", "array" ], "isBold": true, "required": true, "description": "The contents of the system message.", "children": [ { "name": "Text content", "type": "string", "isBold": false, "description": "The contents of the system message." }, { "name": "Array of content parts", "type": "array", "isBold": false, "description": "An array of content parts with a defined type. For system messages, only type text is supported.", "children": [ { "name": "Text content part", "type": "object", "isBold": false, "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The text content." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content part." } ] } ] } ] }, { "name": "role", "type": "string", "isBold": true, "required": true, "description": "Role of the message author.
Available options: system" }, { "name": "name", "type": "string", "isBold": true, "required": false, "description": "An optional name for the participant. Provides the model information to differentiate between participants of the same role." } ] }, { "name": "User message", "type": "object", "isBold": false, "description": "Messages sent by an end user, containing prompts or additional context information.", "children": [ { "name": "content", "type": [ "string", "array" ], "isBold": true, "required": true, "description": "The contents of the user message.", "children": [ { "name": "Text content", "type": "string", "isBold": false, "description": "The text contents of the message." }, { "name": "Array of content parts", "type": "array", "isBold": false, "description": "An array of content parts with a defined type. Supported options differ based on the model being used to generate the response. Can contain text, image, audio or video inputs.
Currently, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed and mimo-v2.5 models support image, audio or video input.
", "children": [ { "name": "Text content part", "type": "object", "isBold": false, "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The text content." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content part." } ] }, { "name": "Image content part", "type": "object", "isBold": false, "children": [ { "name": "image_url", "type": "object", "isBold": true, "required": true, "children": [ { "name": "url", "type": "string", "isBold": true, "required": true, "description": "Either a URL of the image or the base64 encoded image data." } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content part.
Available options: image_url" } ] }, { "name": "Audio content part", "type": "object", "isBold": false, "children": [ { "name": "input_audio", "type": "object", "isBold": true, "required": true, "children": [ { "name": "data", "type": "string", "isBold": true, "required": true, "description": "Either a URL of the audio or the base64 encoded audio data." } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content part.
Available options: input_audio" } ] }, { "name": "Video content part", "type": "object", "isBold": false, "children": [ { "name": "video_url", "type": "object", "isBold": true, "required": true, "children": [ { "name": "url", "type": "string", "isBold": true, "required": true, "description": "Either a URL of the video or the base64 encoded video data." } ] }, { "name": "fps", "type": "number", "isBold": true, "required": false, "defaultValue": "2", "description": "Number of frames sampled per second.
Required range: [0.1, 10.0]" }, { "name": "media_resolution", "type": "string", "isBold": true, "required": false, "defaultValue": "default", "description": "Resolution level.
Available options: default, max" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content part.
Available options: video_url" } ] } ] } ] }, { "name": "role", "type": "string", "isBold": true, "required": true, "description": "Role of the message author.
Available options: user" }, { "name": "name", "type": "string", "isBold": true, "required": false, "description": "An optional name for the participant. Provides model information to differentiate between participants of the same role." } ] }, { "name": "Assistant message", "type": "object", "isBold": false, "description": "Messages sent by the model in response to user messages.", "children": [ { "name": "role", "type": "string", "isBold": true, "required": true, "description": "Role of the message author.
Available options: assistant" }, { "name": "content", "type": [ "string", "array" ], "isBold": true, "required": false, "description": "The contents of the assistant message. Required unless tool_calls is specified.", "children": [ { "name": "Text content", "type": "string", "isBold": false, "description": "The contents of the assistant message." }, { "name": "Array of content parts", "type": "array", "isBold": false, "description": "An array of content parts with a defined type. Can be one or more of type text.", "children": [ { "name": "Text content part", "type": "object", "isBold": false, "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The text content." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content part." } ] } ] } ] }, { "name": "name", "type": "string", "isBold": true, "required": false, "description": "An optional name for the participant. Provides model information to differentiate between participants of the same role." }, { "name": "tool_calls", "type": "array", "isBold": true, "required": false, "description": "The tool calls generated by the model, such as function calls.", "children": [ { "name": "Function tool call", "type": "object", "isBold": false, "description": "A call to a function tool created by the model.", "children": [ { "name": "function", "type": "object", "isBold": true, "required": true, "description": "The function that the model called.", "children": [ { "name": "arguments", "type": "string", "isBold": true, "required": true, "description": "The arguments to call the function with, as generated by the model in JSON format. Note that the model does not always generate valid JSON, and may hallucinate parameters not defined by your function schema. Validate the arguments in your code before calling your function." }, { "name": "name", "type": "string", "isBold": true, "required": true, "description": "The name of the function to call." } ] }, { "name": "id", "type": "string", "isBold": true, "required": true, "description": "The ID of the tool call." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Tool type. Currently, only function is supported." } ] } ] } ] }, { "name": "Tool message", "type": "object", "isBold": false, "children": [ { "name": "content", "type": [ "string", "array" ], "isBold": true, "required": true, "description": "The contents of the tool message.", "children": [ { "name": "Text content", "type": "string", "isBold": false, "description": "The contents of the tool message." }, { "name": "Array of content parts", "type": "array", "isBold": false, "description": "An array of content parts with a defined type. For tool messages, text, image, audio and video are supported.
Currently, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed and mimo-v2.5 models support image, audio or video input.
", "children": [ { "name": "Text content part", "type": "object", "isBold": false, "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The text content." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content part." } ] }, { "name": "Image content part", "type": "object", "isBold": false, "children": [ { "name": "image_url", "type": "object", "isBold": true, "required": true, "children": [ { "name": "url", "type": "string", "isBold": true, "required": true, "description": "Either a URL of the image or the base64 encoded image data." } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content part.
Available options: image_url" } ] }, { "name": "Audio content part", "type": "object", "isBold": false, "children": [ { "name": "input_audio", "type": "object", "isBold": true, "required": true, "children": [ { "name": "data", "type": "string", "isBold": true, "required": true, "description": "Either a URL of the audio or the base64 encoded audio data." } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content part.
Available options: input_audio" } ] }, { "name": "Video content part", "type": "object", "isBold": false, "children": [ { "name": "video_url", "type": "object", "isBold": true, "required": true, "children": [ { "name": "url", "type": "string", "isBold": true, "required": true, "description": "Either a URL of the video or the base64 encoded video data." } ] }, { "name": "fps", "type": "number", "isBold": true, "required": false, "defaultValue": "2", "description": "Number of frames sampled per second.
Required range: [0.1, 10.0]" }, { "name": "media_resolution", "type": "string", "isBold": true, "required": false, "defaultValue": "default", "description": "Resolution level.
Available options: default, max" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content part.
Available options: video_url" } ] } ] } ] }, { "name": "role", "type": "string", "isBold": true, "required": true, "description": "Role of the message author.
Available options: tool" }, { "name": "tool_call_id", "type": "string", "isBold": true, "required": true, "description": "Tool call that this message is responding to." } ] } ] }, { "name": "model", "type": "string", "isBold": true, "required": true, "description": "Model ID is used to generate the response.
Available options: mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro, mimo-v2.5" }, { "name": "frequency_penalty", "type": [ "number", "null" ], "isBold": true, "required": false, "defaultValue": "0", "description": "Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.
Required range: [-2.0, 2.0]" }, { "name": "max_completion_tokens", "type": [ "integer", "null" ], "isBold": true, "required": false, "description": "An upper bound for the number of tokens that can be generated for a completion, including visible output tokens and reasoning tokens.
  • mimo-v2.6-flash: default 131072
  • mimo-v2.6-pro: default 131072
  • mimo-v2.6-pro-ultraspeed: default 131072
  • mimo-v2.5-pro: default 131072
  • mimo-v2.5: default 32768
Required range: [1, 131072]" }, { "name": "presence_penalty", "type": [ "number", "null" ], "isBold": true, "required": false, "defaultValue": "0", "description": "Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.
Required range: [-2.0, 2.0]" }, { "name": "response_format", "type": "object", "isBold": true, "required": false, "description": "An object specifying the format that the model must output.", "children": [ { "name": "Text", "type": "object", "isBold": false, "description": "Default response format. Used to generate text responses.", "children": [ { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of response format being defined. Always text." } ] }, { "name": "JSON object", "type": "object", "isBold": false, "description": "JSON object response format. Note that the model will not generate JSON without a system or user message instructing it to do so.", "children": [ { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of response format being defined. Always json_object." } ] } ] }, { "name": "stop", "type": [ "string", "array", "null" ], "isBold": true, "required": false, "defaultValue": "null", "description": "Up to 4 sequences where the API will stop generating further tokens. The returned text will not contain the stop sequence." }, { "name": "stream", "type": [ "boolean", "null" ], "isBold": true, "required": false, "defaultValue": "false", "description": "If set to true, the model response data will be streamed to the client as it is generated using server-sent events." }, { "name": "thinking", "type": "object", "isBold": true, "required": false, "description": "This parameter is used to control whether the model enables the chain of thought.
Note: During the multi-turn tool calls process in thinking mode, the model returns a reasoning_content field alongside tool_calls. To continue the conversation, it is recommended to keep all previous reasoning_content in the messages array for each subsequent request to achieve the best performance.
In thinking mode, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro and mimo-v2.5 models do not support customizing the temperature and top_p parameters. Even if these parameters are passed in, the actual effective values will be forcibly set by the model to its recommended default values of 1.0 and 0.95.
", "children": [ { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Whether to enable the chain of thought.
  • mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro, mimo-v2.5: default enabled
Available options: enabled, disabled" } ] }, { "name": "temperature", "type": "number", "isBold": true, "required": false, "description": "What sampling temperature to use, between 0 and 1.5. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic. We generally recommend altering this or top_p but not both.
In thinking mode, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro and mimo-v2.5 models do not support customizing the temperature parameter. Even if this parameter is passed in, it will be forcibly overridden and take effect with the model's recommended default value of 1.0.
  • mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro, mimo-v2.5: default 1.0
Required range: [0, 1.5]" }, { "name": "tool_choice", "type": "string", "isBold": true, "required": false, "description": "Controls how the model selects a tool.
Note: When a value other than auto is passed to tool_choice, the backend will remove this field by default, and the model response behavior will still be equivalent to the auto mode (this logic is subject to future adjustments).
Available options: auto" }, { "name": "tools", "type": "array", "isBold": true, "required": false, "description": "A list of tools the model may call. You can provide function tools.
Note: During the multi-turn tool calls process in thinking mode, the model returns a reasoning_content field alongside tool_calls. To continue the conversation, it is recommended to keep all previous reasoning_content in the messages array for each subsequent request to achieve the best performance.
", "children": [ { "name": "Function tool", "type": "object", "isBold": false, "description": "A function tool that can be used to generate a response.", "children": [ { "name": "function", "type": "object", "isBold": true, "required": true, "children": [ { "name": "name", "type": "string", "isBold": true, "required": true, "description": "The name of the tool function. Must be a-z, A-Z, 0-9, or contain underscores (_) and dashes (-), with a maximum length of 64.
Required string length: 1 - 64" }, { "name": "description", "type": "string", "isBold": true, "required": false, "description": "A description of what the function does, used by the model to choose when and how to call the function." }, { "name": "parameters", "type": "object", "isBold": true, "required": false, "description": "The parameters the functions accept, described as a JSON Schema object.
Omitting parameters defines a function with an empty parameter list." }, { "name": "strict", "type": "boolean", "isBold": true, "required": false, "defaultValue": "false", "description": "Whether to enable strict schema adherence when generating the function call. If set to true, the model will follow the exact schema defined in the parameters field. Only a subset of JSON Schema is supported when strict is true." } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Tool type. Currently, only function is supported." } ] }, { "name": "Web search tool", "type": "object", "isBold": false, "description": "A web search tool that can be used to generate a response.For details, please refer to Web Search.
Note:Web Search plugin must be activated before use.
", "children": [ { "name": "user_location", "type": "object", "isBold": true, "required": false, "children": [ { "name": "type", "type": "string", "isBold": true, "required": true, "description": "approximate" }, { "name": "country", "type": "string", "isBold": true, "required": false, "description": "country" }, { "name": "region", "type": "string", "isBold": true, "required": false, "description": "region" }, { "name": "city", "type": "string", "isBold": true, "required": false, "description": "city" }, { "name": "district", "type": "string", "isBold": true, "required": false, "description": "district" }, { "name": "longitude", "type": "long", "isBold": true, "required": false, "description": "longitude " }, { "name": "latitude", "type": "long", "isBold": true, "required": false, "description": "latitude" } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Tool type. Currently, only web_search is supported." }, { "name": "force_search", "type": "boolean", "isBold": true, "required": false, "defaultValue": "false", "description": "Whether to enable forced search. true for forced search, false for the model to decide whether search is needed." }, { "name": "max_keyword", "type": "integer", "isBold": true, "required": false, "defaultValue": "5", "description": "Limit the maximum number of keywords that can be used in a single search.
Required range: [1, 50]" }, { "name": "limit", "type": "integer", "isBold": true, "required": false, "defaultValue": "5", "description": "Limit the maximum number of results returned by a single search operation.
Required range: [1, 50]" } ] } ] }, { "name": "top_p", "type": "number", "isBold": true, "required": false, "defaultValue": "0.95", "description": "The probability threshold for nucleus sampling, which controls the diversity of the text that the model generates. A higher top_p value results in more diverse text. A lower top_p value results in more deterministic text.
Because both temperature and top_p control the diversity of the generated text, we recommend that you set only one of them.
In thinking mode, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro, mimo-v2.5 models do not support customizing the top_p parameter. Even if this parameter is passed in, it will be forcibly overridden and take effect with the model's recommended default value of 0.95.
Required range: [0.01, 1.0]" } ]`} /> ## Chat response object (non-streaming output) stop if the model hit a natural stop point or a provided stop sequence, length if the maximum number of tokens specified in the request was reached, tool_calls if the model called a tool, content_filter if content was omitted due to a flag from our content filters, repetition_truncation if the model detects repetition." }, { "name": "index", "type": "integer", "isBold": true, "description": "The index of the choice in the list of choices." }, { "name": "message", "type": "object", "isBold": true, "description": "A chat completion message generated by the model.", "children": [ { "name": "content", "type": "string", "isBold": true, "description": "The contents of the message." }, { "name": "reasoning_content", "type": "string", "isBold": true, "description": "The reasoning contents of the assistant message, before the final answer." }, { "name": "role", "type": "string", "isBold": true, "description": "The role of the author of this message." }, { "name": "tool_calls", "type": "array", "isBold": true, "description": "After a function call is initiated, the model returns the tool to be called and the parameters that are Required for the call. This parameter can contain one or more tool response objects.", "children": [ { "name": "Function tool call", "type": "object", "isBold": false, "description": "A call to a function tool created by the model.", "children": [ { "name": "function", "type": "object", "isBold": true, "description": "The function that the model called.", "children": [ { "name": "arguments", "type": "string", "isBold": true, "description": "The arguments to call the function with, as generated by the model in JSON format. Note that the model does not always generate valid JSON, and may hallucinate parameters not defined by your function schema. Validate the arguments in your code before calling your function." }, { "name": "name", "type": "string", "isBold": true, "description": "The name of the function to call." } ] }, { "name": "id", "type": "string", "isBold": true, "description": "The ID of the tool call." }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the tool. Currently, only function is supported." } ] } ] }, { "name": "annotations", "type": "array", "isBold": true, "description": "After web search, the model returns annotations for all referenced URLs.", "children": [ { "name": "web_search tool call", "type": "object", "isBold": false, "description": "A call to a web search tool created by the model.", "children": [ { "name": "logo_url", "type": "string", "isBold": true, "description": "Logo url." }, { "name": "publish_time", "type": "string", "isBold": true, "description": "Publish time." }, { "name": "site_name", "type": "string", "isBold": true, "description": "Site name." }, { "name": "summary", "type": "string", "isBold": true, "description": "Summary." }, { "name": "title", "type": "string", "isBold": true, "description": "Title." }, { "name": "type", "type": "string", "isBold": true, "description": "Type." }, { "name": "url", "type": "string", "isBold": true, "description": "Url." } ] } ] }, { "name": "error_message", "type": "string", "isBold": true, "description": "Error message of web search." } ] } ] }, { "name": "created", "type": "integer", "isBold": true, "description": "The Unix timestamp (in seconds) of when the chat completion was created." }, { "name": "id", "type": "string", "isBold": true, "description": "A unique identifier for the chat completion." }, { "name": "model", "type": "string", "isBold": true, "description": "The model to generate the completion." }, { "name": "object", "type": "string", "isBold": true, "description": "The object type, which is always chat.completion." }, { "name": "usage", "type": [ "object", "null" ], "isBold": true, "description": "Usage statistics for the completion request.", "children": [ { "name": "completion_tokens", "type": "integer", "isBold": true, "description": "Number of tokens in the generated completion." }, { "name": "prompt_tokens", "type": "integer", "isBold": true, "description": "Number of tokens in the prompt." }, { "name": "total_tokens", "type": "integer", "isBold": true, "description": "Total number of tokens used in the request (prompt + completion)." }, { "name": "completion_tokens_details", "type": "object", "isBold": true, "description": "Breakdown of tokens used in a completion.", "children": [ { "name": "reasoning_tokens", "type": "integer", "isBold": true, "description": "Tokens generated by the model for reasoning." } ] }, { "name": "prompt_tokens_details", "type": "object", "isBold": true, "description": "Breakdown of tokens used in the prompt.", "children": [ { "name": "cached_tokens", "type": "integer", "isBold": true, "description": "Number of tokens served from cache." }, { "name": "audio_tokens", "type": "integer", "isBold": true, "description": "Audio input tokens present in the prompt." }, { "name": "image_tokens", "type": "integer", "isBold": true, "description": "Image input tokens present in the prompt." }, { "name": "video_tokens", "type": "integer", "isBold": true, "description": "Video input tokens present in the prompt." } ] }, { "name": "web_search_usage", "type": "object", "isBold": true, "description": "Detailed usage of the web search API.", "children": [ { "name": "tool_usage", "type": "integer", "isBold": true, "description": "Number of API calls in web search." }, { "name": "page_usage", "type": "integer", "isBold": true, "description": "Number of web pages returned by the web search API." } ] } ] } ]`} /> ## Chat response chunk object (streaming output) tool_calls list, starting from 0." }, { "name": "function", "type": "object", "isBold": true, "description": "The function to be called.", "children": [ { "name": "arguments", "type": "string", "isBold": true, "description": "The arguments to call the function with, as generated by the model in JSON format. Note that the model does not always generate valid JSON, and may hallucinate parameters not defined by your function schema. Validate the arguments in your code before calling your function." }, { "name": "name", "type": "string", "isBold": true, "description": "The name of the function to call." } ] }, { "name": "id", "type": "string", "isBold": true, "description": "The ID of the tool call." }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the tool. Currently, only function is supported." } ] }, { "name": "annotations", "type": "array", "isBold": true, "description": "After web search, the model returns annotations for all referenced URLs.", "children": [ { "name": "web_search tool call", "type": "object", "isBold": false, "description": "A call to a web search tool created by the model.", "children": [ { "name": "logo_url", "type": "string", "isBold": true, "description": "Logo url." }, { "name": "publish_time", "type": "string", "isBold": true, "description": "Publish time." }, { "name": "site_name", "type": "string", "isBold": true, "description": "Site name." }, { "name": "summary", "type": "string", "isBold": true, "description": "Summary." }, { "name": "title", "type": "string", "isBold": true, "description": "Title." }, { "name": "type", "type": "string", "isBold": true, "description": "Type." }, { "name": "url", "type": "string", "isBold": true, "description": "Url." } ] } ] }, { "name": "error_message", "type": "string", "isBold": true, "description": "Error message of web search." } ] }, { "name": "finish_reason", "type": [ "string", "null" ], "isBold": true, "description": "The reason the model stopped generating tokens. This will be stop if the model hit a natural stop point or a provided stop sequence, length if the maximum number of tokens specified in the request was reached, tool_calls if the model called a tool, content_filter if content was omitted due to a flag from our content filters, repetition_truncation if the model detects repetition." }, { "name": "index", "type": "integer", "isBold": true, "description": "The index of the choice in the list of choices." } ] }, { "name": "created", "type": "integer", "isBold": true, "description": "The Unix timestamp (in seconds) of when the chat completion was created. Each chunk has the same timestamp." }, { "name": "id", "type": "string", "isBold": true, "description": "A unique identifier for the chat completion. Each chunk has the same ID." }, { "name": "model", "type": "string", "isBold": true, "description": "The model to generate the completion." }, { "name": "object", "type": "string", "isBold": true, "description": "The object type, which is always chat.completion.chunk." }, { "name": "usage", "type": [ "object", "null" ], "isBold": true, "description": "Usage statistics for the completion request.", "children": [ { "name": "completion_tokens", "type": "integer", "isBold": true, "description": "Number of tokens in the generated completion." }, { "name": "prompt_tokens", "type": "integer", "isBold": true, "description": "Number of tokens in the prompt." }, { "name": "total_tokens", "type": "integer", "isBold": true, "description": "Total number of tokens used in the request (prompt + completion)." }, { "name": "completion_tokens_details", "type": "object", "isBold": true, "description": "Breakdown of tokens used in a completion.", "children": [ { "name": "reasoning_tokens", "type": "integer", "isBold": true, "description": "Tokens generated by the model for reasoning." } ] }, { "name": "prompt_tokens_details", "type": "object", "isBold": true, "description": "Breakdown of tokens used in the prompt.", "children": [ { "name": "cached_tokens", "type": "integer", "isBold": true, "description": "Number of tokens served from cache." }, { "name": "audio_tokens", "type": "integer", "isBold": true, "description": "Audio input tokens present in the prompt." }, { "name": "image_tokens", "type": "integer", "isBold": true, "description": "Image input tokens present in the prompt." }, { "name": "video_tokens", "type": "integer", "isBold": true, "description": "Video input tokens present in the prompt." } ] }, { "name": "web_search_usage", "type": "object", "isBold": true, "description": "Detailed usage of the web search API.", "children": [ { "name": "tool_usage", "type": "integer", "isBold": true, "description": "Number of API calls in web search." }, { "name": "page_usage", "type": "integer", "isBold": true, "description": "Number of web pages returned by the web search API." } ] } ] } ]`} /> --- DOCUMENT: OpenAI Responses API --- URL: https://mimo.mi.com/static/docs/api/chat/responses.md # OpenAI Responses API Compatibility {/* feishu-style:text-align:left */} MiMo provides a calling interface compatible with the OpenAI Responses API format. This document covers request parameters, response schemas and code examples. {/* feishu-style:text-align:left */} **Compatibility Notes & Limitations:** {/* feishu-style:text-align:left */} This interface aligns with the OpenAI Responses API specification to facilitate quick integration for developers. Only the parameters documented here will be processed normally; undefined parameters will be filtered out and may cause request errors. See below for specific behavioral differences. - **Incompatible parameters**: Fields such as `background`, `previous_response_id`, and `context_management` are not currently supported. Carrying these parameters in a request will be ignored or trigger an error. - **Reasoning level control**: `reasoning.effort` controls model reasoning. `none` disables thinking; Every other level enables thinking with identical behavior — The reasoning intensity is not differentiated at this stage. ## Request Address ```bash https://api.xiaomimimo.com/v1/responses ``` ## Request Headers {/* feishu-style:text-align:left */} The API supports the following two authentication methods. Please choose one and add it to the request headers: ```json api-key: $MIMO_API_KEY Content-Type: application/json ``` ```json Authorization: Bearer $MIMO_API_KEY Content-Type: application/json ``` ## Request Body user role." }, { "name": "InputItemList", "type": "array", "isBold": false, "description": "A list of one or many input items to the model, containing different content types.", "children": [ { "name": "EasyInputMessage", "type": "object", "isBold": false, "description": "A message input to the model with a role indicating instruction following hierarchy. Instructions given with the developer or system role take precedence over instructions given with the user role.", "children": [ { "name": "content", "type": [ "string", "array" ], "isBold": true, "required": true, "description": "Text, image, audio and video inputs to the model, used to generate a response.", "children": [ { "name": "TextInput", "type": "string", "isBold": false, "description": "A text input to the model." }, { "name": "ResponseInputMessageContentList", "type": "array", "isBold": false, "description": "A list of one or many input items to the model, containing different content types.
Currently, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed and mimo-v2.5 models support image, audio or video input.
", "children": [ { "name": "ResponseInputText", "type": "object", "isBold": false, "description": "A text input to the model.", "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The text input to the model." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_text" } ] }, { "name": "ResponseInputImage", "type": "object", "isBold": false, "description": "An image input to the model.", "children": [ { "name": "image_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the image to be sent to the model. A fully qualified URL or base64 encoded image in a data URL." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_image" } ] }, { "name": "ResponseInputAudio", "type": "object", "isBold": false, "description": "An audio input to the model.", "children": [ { "name": "audio_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the audio to be sent to the model. A fully qualified URL or base64 encoded audio in a data URL." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_audio" } ] }, { "name": "ResponseInputVideo", "type": "object", "isBold": false, "description": "A video input to the model.", "children": [ { "name": "video_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the video to be sent to the model. A fully qualified URL or base64 encoded video in a data URL." }, { "name": "fps", "type": "number", "isBold": true, "required": false, "defaultValue": "2", "description": "Number of frames sampled per second.
Required range: [0.1, 10.0]" }, { "name": "media_resolution", "type": "string", "isBold": true, "required": false, "defaultValue": "default", "description": "Resolution level.
Available options: default, max" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_video" } ] } ] } ] }, { "name": "role", "type": "string", "isBold": true, "required": true, "description": "The role of the message input.
Available options: user, assistant, system, developer" }, { "name": "type", "type": "string", "isBold": true, "required": false, "description": "The type of the message input.
Available options: message" } ] }, { "name": "Message", "type": "object", "isBold": false, "description": "A message input to the model with a role indicating instruction following hierarchy.", "children": [ { "name": "content", "type": "array", "isBold": true, "required": true, "description": "A list of one or many input items to the model, containing different content types.
Currently, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed and mimo-v2.5 models support image, audio or video input.
", "children": [ { "name": "ResponseInputText", "type": "object", "isBold": false, "description": "A text input to the model.", "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The text input to the model." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_text" } ] }, { "name": "ResponseInputImage", "type": "object", "isBold": false, "description": "An image input to the model.", "children": [ { "name": "image_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the image to be sent to the model. A fully qualified URL or base64 encoded image in a data URL." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_image" } ] }, { "name": "ResponseInputAudio", "type": "object", "isBold": false, "description": "An audio input to the model.", "children": [ { "name": "audio_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the audio to be sent to the model. A fully qualified URL or base64 encoded audio in a data URL." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_audio" } ] }, { "name": "ResponseInputVideo", "type": "object", "isBold": false, "description": "A video input to the model.", "children": [ { "name": "video_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the video to be sent to the model. A fully qualified URL or base64 encoded video in a data URL." }, { "name": "fps", "type": "number", "isBold": true, "required": false, "defaultValue": "2", "description": "Number of frames sampled per second.
Required range: [0.1, 10.0]" }, { "name": "media_resolution", "type": "string", "isBold": true, "required": false, "defaultValue": "default", "description": "Resolution level.
Available options: default, max" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_video" } ] } ] }, { "name": "role", "type": "string", "isBold": true, "required": true, "description": "The role of the message input.
Available options: user, system, developer" }, { "name": "status", "type": "string", "isBold": true, "required": false, "description": "The status of item. Populated when items are returned via API.
Available options: in_progress, completed" }, { "name": "type", "type": "string", "isBold": true, "required": false, "description": "The type of the message input.
Available options: message" } ] }, { "name": "ResponseOutputMessage", "type": "object", "isBold": false, "description": "An output message from the model.", "children": [ { "name": "id", "type": "string", "isBold": true, "required": true, "description": "The unique ID of the output message." }, { "name": "content", "type": "array", "isBold": true, "required": true, "description": "The content of the output message.", "children": [ { "name": "ResponseOutputText", "type": "object", "isBold": false, "description": "A text output from the model.", "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The text output from the model." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the output text.
Available options: output_text" } ] } ] }, { "name": "role", "type": "string", "isBold": true, "required": true, "description": "The role of the output message.
Available options: assistant" }, { "name": "status", "type": "string", "isBold": true, "required": true, "description": "The status of the message input. Populated when input items are returned via API.
Available options: in_progress, completed" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the output message.
Available options: message" } ] }, { "name": "FunctionCall", "type": "object", "isBold": false, "description": "A tool call to run a function.", "children": [ { "name": "arguments", "type": "string", "isBold": true, "required": true, "description": "A JSON string of the arguments to pass to the function." }, { "name": "call_id", "type": "string", "isBold": true, "required": true, "description": "The unique ID of the function tool call generated by the model." }, { "name": "name", "type": "string", "isBold": true, "required": true, "description": "The name of the function to run." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the function tool call.
Available options: function_call" }, { "name": "id", "type": "string", "isBold": true, "required": false, "description": "The unique ID of the function tool call." }, { "name": "namespace", "type": "string", "isBold": true, "required": false, "description": "The namespace of the function to run." }, { "name": "status", "type": "string", "isBold": true, "required": false, "description": "The status of the item. Populated when items are returned via API.
Available options: in_progress, completed" } ] }, { "name": "FunctionCallOutput", "type": "object", "isBold": false, "description": "The output of a function tool call.", "children": [ { "name": "call_id", "type": "string", "isBold": true, "required": true, "description": "The unique ID of the function tool call generated by the model." }, { "name": "output", "type": [ "string", "array" ], "isBold": true, "required": true, "description": "Text, image, audio or video output of the function tool call.", "children": [ { "name": "StringOutput", "type": "string", "isBold": false, "description": "A JSON string of the output of the function tool call." }, { "name": "OutputContentList", "type": "array", "isBold": false, "description": "An array of content outputs for the function tool call.
Currently, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed and mimo-v2.5 models support image, audio or video input.
", "children": [ { "name": "ResponseInputTextContent", "type": "object", "isBold": false, "description": "A text input to the model.", "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The text input to the model." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_text" } ] }, { "name": "ResponseInputImageContent", "type": "object", "isBold": false, "description": "An image input to the model.", "children": [ { "name": "image_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the image to be sent to the model. A fully qualified URL or base64 encoded image in a data URL." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_image" } ] }, { "name": "ResponseInputAudioContent", "type": "object", "isBold": false, "description": "An audio input to the model.", "children": [ { "name": "audio_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the audio to be sent to the model. A fully qualified URL or base64 encoded audio in a data URL." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_audio" } ] }, { "name": "ResponseInputVideoContent", "type": "object", "isBold": false, "description": "A video input to the model.", "children": [ { "name": "video_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the video to be sent to the model. A fully qualified URL or base64 encoded video in a data URL." }, { "name": "fps", "type": "number", "isBold": true, "required": false, "defaultValue": "2", "description": "Number of frames sampled per second.
Required range: [0.1, 10.0]" }, { "name": "media_resolution", "type": "string", "isBold": true, "required": false, "defaultValue": "default", "description": "Resolution level.
Available options: default, max" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_video" } ] } ] } ] }, { "name": "id", "type": "string", "isBold": true, "required": false, "description": "The unique ID of the function tool call output. Populated when this item is returned via API." }, { "name": "status", "type": "string", "isBold": true, "required": false, "description": "The status of the item. Populated when items are returned via API.
Available options: in_progress, completed" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the function tool call output.
Available options: function_call_output" } ] }, { "name": "AgentMessage[Multi-Agent]", "type": "object", "isBold": false, "description": "A message routed between agents.", "children": [ { "name": "author", "type": "string", "isBold": true, "required": true, "description": "The sending agent identity." }, { "name": "content", "type": "array", "isBold": true, "required": true, "description": "Plaintext, image, audio, video or encrypted content sent between agents.
Currently, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed and mimo-v2.5 models support image, audio or video input.
", "children": [ { "name": "ResponseInputTextContent", "type": "object", "isBold": false, "description": "A text input to the model.", "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The text input to the model." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_text" } ] }, { "name": "ResponseInputImageContent", "type": "object", "isBold": false, "description": "An image input to the model.", "children": [ { "name": "image_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the image to be sent to the model. A fully qualified URL or base64 encoded image in a data URL." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_image" } ] }, { "name": "ResponseInputAudioContent", "type": "object", "isBold": false, "description": "An audio input to the model.", "children": [ { "name": "audio_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the audio to be sent to the model. A fully qualified URL or base64 encoded audio in a data URL." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_audio" } ] }, { "name": "ResponseInputVideoContent", "type": "object", "isBold": false, "description": "A video input to the model.", "children": [ { "name": "video_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the video to be sent to the model. A fully qualified URL or base64 encoded video in a data URL." }, { "name": "fps", "type": "number", "isBold": true, "required": false, "defaultValue": "2", "description": "Number of frames sampled per second.
Required range: [0.1, 10.0]" }, { "name": "media_resolution", "type": "string", "isBold": true, "required": false, "defaultValue": "default", "description": "Resolution level.
Available options: default, max" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_video" } ] }, { "name": "EncryptedContent", "type": "object", "isBold": false, "description": "Opaque encrypted content that Responses API decrypts inside trusted model execution. For MiMo, the content is plaintext and no decryption is performed.", "children": [ { "name": "encrypted_content", "type": "string", "isBold": true, "required": true, "description": "Opaque encrypted content." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: encrypted_content" } ] } ] }, { "name": "recipient", "type": "string", "isBold": true, "required": true, "description": "The destination agent identity." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The item type.
Available options: agent_message" }, { "name": "id", "type": [ "string", "null" ], "isBold": true, "required": false, "description": "The unique ID of this agent message item." }, { "name": "agent", "type": [ "object", "null" ], "isBold": true, "required": false, "description": "The agent that produced this item.", "children": [ { "name": "agent_name", "type": "string", "isBold": true, "required": true, "description": "The canonical name of the agent that produced this item." } ] } ] }, { "name": "AdditionalTools", "type": "object", "isBold": false, "children": [ { "name": "role", "type": "string", "isBold": true, "required": true, "description": "The role that provided the additional tools.
Available options: developer" }, { "name": "tools", "type": "array", "isBold": true, "required": true, "description": "A list of additional tools made available at this item.", "children": [ { "name": "Function", "type": "object", "isBold": false, "description": "Defines a function in your own code the model can choose to call.", "children": [ { "name": "name", "type": "string", "isBold": true, "required": true, "description": "The name of the tool function. Must be a-z, A-Z, 0-9, or contain underscores (_) and dashes (-), with a maximum length of 64.
Required string length: 1 - 64" }, { "name": "parameters", "type": "object", "isBold": true, "required": true, "description": "A JSON schema object describing the parameters of the function." }, { "name": "strict", "type": "boolean", "isBold": true, "required": true, "defaultValue": "false", "description": "Whether to enable strict schema adherence when generating the function call." }, { "name": "description", "type": [ "string", "null" ], "isBold": true, "required": false, "description": "A description of the function. Used by the model to determine whether or not to call the function." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the tool.
Available options: function" } ] }, { "name": "Custom", "type": "object", "isBold": false, "description": "A custom tool that processes input using a specified format.", "children": [ { "name": "name", "type": "string", "isBold": true, "required": true, "description": "The name of the custom tool, used to identify it in tool calls." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the custom tool.
Available options: custom" }, { "name": "description", "type": "string", "isBold": true, "required": false, "description": "Optional description of the custom tool, used to provide more context." }, { "name": "format", "type": "object", "isBold": true, "required": false, "description": "The input format for the custom tool. Default is unconstrained text.", "children": [ { "name": "CustomToolInputFormat", "type": "object", "isBold": false, "children": [ { "name": "Text", "type": "object", "isBold": false, "description": "Unconstrained free-form text.", "children": [ { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Unconstrained text format.
Available options: text" } ] }, { "name": "Grammar", "type": "object", "isBold": false, "description": "A grammar defined by the user.", "children": [ { "name": "definition", "type": "string", "isBold": true, "required": true, "description": "The grammar definition." }, { "name": "syntax", "type": "string", "isBold": true, "required": true, "description": "The syntax of the grammar definition.
Available options: lark, regex" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Grammar format.
Available options: grammar" } ] } ] } ] } ] }, { "name": "Namespace", "type": "object", "isBold": false, "description": "Groups function tools under a shared namespace.", "children": [ { "name": "description", "type": "string", "isBold": true, "required": true, "description": "A description of the namespace shown to the model." }, { "name": "name", "type": "string", "isBold": true, "required": true, "description": "The namespace name used in tool calls.
MinLength: 1" }, { "name": "tools", "type": "array", "isBold": true, "required": true, "description": "The function tools available inside this namespace.", "children": [ { "name": "Function", "type": "object", "isBold": false, "children": [ { "name": "name", "type": "string", "isBold": true, "required": true, "description": "Required string length: 1 - 64" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Available options: function" }, { "name": "description", "type": [ "string", "null" ], "isBold": true, "required": false }, { "name": "parameters", "type": "object", "isBold": true, "required": false }, { "name": "strict", "type": "boolean", "isBold": true, "required": false } ] }, { "name": "Custom", "type": "object", "isBold": false, "description": "A custom tool that processes input using a specified format.", "children": [ { "name": "name", "type": "string", "isBold": true, "required": true, "description": "The name of the custom tool, used to identify it in tool calls." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the custom tool.
Available options: custom" }, { "name": "description", "type": "string", "isBold": true, "required": false, "description": "Optional description of the custom tool, used to provide more context." }, { "name": "format", "type": "object", "isBold": true, "required": false, "description": "The input format for the custom tool. Default is unconstrained text.", "children": [ { "name": "CustomToolInputFormat", "type": "object", "isBold": false, "children": [ { "name": "Text", "type": "object", "isBold": false, "description": "Unconstrained free-form text.", "children": [ { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Unconstrained text format.
Available options: text" } ] }, { "name": "Grammar", "type": "object", "isBold": false, "description": "A grammar defined by the user.", "children": [ { "name": "definition", "type": "string", "isBold": true, "required": true, "description": "The grammar definition." }, { "name": "syntax", "type": "string", "isBold": true, "required": true, "description": "The syntax of the grammar definition.
Available options: lark, regex" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Grammar format.
Available options: grammar" } ] } ] } ] } ] } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the tool.
Available options: namespace" } ] } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The item type.
Available options: additional_tools" }, { "name": "id", "type": [ "string", "null" ], "isBold": true, "required": false, "description": "The unique ID of this additional tools item." } ] }, { "name": "Reasoning", "type": "object", "isBold": false, "description": "A description of the chain of thought used by a reasoning model while generating a response.", "children": [ { "name": "id", "type": "string", "isBold": true, "required": true, "description": "The unique identifier of the reasoning content." }, { "name": "content", "type": "array", "isBold": true, "required": false, "description": "Reasoning text content.", "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The reasoning text from the model." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the object.
Available options: reasoning_text" } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the object.
Available options: reasoning" }, { "name": "status", "type": "string", "isBold": true, "required": false, "description": "The status of the item. Populated when items are returned via API.
Available options: in_progress, completed" } ] }, { "name": "CustomToolCallOutput", "type": "object", "isBold": false, "description": "The output of a custom tool call from your code, being sent back to the model.", "children": [ { "name": "call_id", "type": "string", "isBold": true, "required": true, "description": "The call ID, used to map this custom tool call output to a custom tool call." }, { "name": "output", "type": [ "string", "array" ], "isBold": true, "required": true, "description": "The output from the custom tool call generated by your code. Can be a string or a list of output content.", "children": [ { "name": "StringOutput", "type": "string", "isBold": false, "description": "A string of the output of the custom tool call." }, { "name": "OutputContentList", "type": "array", "isBold": false, "description": "An array of content outputs for the function tool call.
Currently, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed and mimo-v2.5 models support image, audio or video input.
", "children": [ { "name": "ResponseInputText", "type": "object", "isBold": false, "description": "A text input to the model.", "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The text input to the model." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_text" } ] }, { "name": "ResponseInputImage", "type": "object", "isBold": false, "description": "An image input to the model.", "children": [ { "name": "image_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the image to be sent to the model. A fully qualified URL or base64 encoded image in a data URL." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_image" } ] }, { "name": "ResponseInputAudio", "type": "object", "isBold": false, "description": "An audio input to the model.", "children": [ { "name": "audio_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the audio to be sent to the model. A fully qualified URL or base64 encoded audio in a data URL." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_audio" } ] }, { "name": "ResponseInputVideo", "type": "object", "isBold": false, "description": "A video input to the model.", "children": [ { "name": "video_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the video to be sent to the model. A fully qualified URL or base64 encoded video in a data URL." }, { "name": "fps", "type": "number", "isBold": true, "required": false, "defaultValue": "2", "description": "Number of frames sampled per second.
Required range: [0.1, 10.0]" }, { "name": "media_resolution", "type": "string", "isBold": true, "required": false, "defaultValue": "default", "description": "Resolution level.
Available options: default, max" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_video" } ] } ] } ] }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the custom tool call output.
Available options: custom_tool_call_output" }, { "name": "id", "type": "string", "isBold": true, "required": false, "description": "The unique ID of the custom tool call output." } ] }, { "name": "CustomToolCall", "type": "object", "isBold": false, "description": "A call to a custom tool created by the model.", "children": [ { "name": "call_id", "type": "string", "isBold": true, "required": true, "description": "An identifier used to map this custom tool call to a tool call output." }, { "name": "input", "type": "string", "isBold": true, "required": true, "description": "The input for the custom tool call generated by the model." }, { "name": "name", "type": "string", "isBold": true, "required": true, "description": "The name of the custom tool being called." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the custom tool call.
Available options: custom_tool_call" }, { "name": "id", "type": "string", "isBold": true, "required": false, "description": "The unique ID of the custom tool call." }, { "name": "namespace", "type": "string", "isBold": true, "required": false, "description": "The namespace of the custom tool being called." } ] } ] } ] }, { "name": "instructions", "type": "string", "isBold": true, "required": false, "description": "A system (or developer) message inserted into the model's context." }, { "name": "max_output_tokens", "type": "integer", "isBold": true, "required": false, "description": "An upper bound for the number of tokens that can be generated for a response, including visible output tokens and reasoning tokens.
  • mimo-v2.6-flash: default 131072
  • mimo-v2.6-pro: default 131072
  • mimo-v2.6-pro-ultraspeed: default 131072
  • mimo-v2.5-pro: default 131072
  • mimo-v2.5: default 32768
Required range: [1, 131072]" }, { "name": "model", "type": "string", "isBold": true, "required": true, "description": "Model ID used to generate the response.
Available options: mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro, mimo-v2.5" }, { "name": "stream", "type": "boolean", "isBold": true, "required": false, "defaultValue": "false", "description": "If set to true, the model response data will be streamed to the client as it is generated using server-sent events." }, { "name": "reasoning", "type": "object", "isBold": true, "required": false, "description": "Configuration options for reasoning models.
Note: During the multi-turn tool calls process in thinking mode, the model returns the reasoning content alongside the tool calls field. To continue the conversation, it is recommended to keep all previous reasoning content in the input array for each subsequent request to achieve the best performance.
In thinking mode, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro and mimo-v2.5 models do not support customizing the temperature and top_p parameters. Even if these parameters are passed in, the actual effective values will be forcibly set by the model to its recommended default values of 1.0 and 0.95.
", "children": [ { "name": "effort", "type": "string", "isBold": true, "required": true, "description": "Constrains effort on reasoning for reasoning models. Reducing reasoning effort can result in faster responses and fewer tokens used on reasoning in a response.
Custom tuning of reasoning effort is currently unsupported. When set to none, reasoning is disabled; all other valid values map to enabled reasoning. The value minimal is mapped to low. Values xhigh, max and ultra are mapped to high.
  • mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro, mimo-v2.5: default enabled
Available options: none, minimal, low, medium, high, xhigh, max, ultra" } ] }, { "name": "temperature", "type": "number", "isBold": true, "required": false, "description": "What sampling temperature to use, between 0 and 1.5. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic. We generally recommend altering this or top_p but not both.
In thinking mode, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro and mimo-v2.5 models do not support customizing the temperature parameter. Even if this parameter is passed in, it will be forcibly overridden and take effect with the model's recommended default value of 1.0.
  • mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro, mimo-v2.5: default 1.0
Required range: [0, 1.5]" }, { "name": "text", "type": "object", "isBold": true, "required": false, "description": "Configuration options for a text response from the model. Can be plain text or structured JSON data.", "children": [ { "name": "format", "type": "object", "isBold": true, "required": false, "description": "An object specifying the format that the model must output. The default format is { "type": "text" } with no additional options.", "children": [ { "name": "ResponseFormatText", "type": "object", "isBold": false, "description": "Default response format. Used to generate text responses.", "children": [ { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of response format being defined.
Available options: text" } ] }, { "name": "ResponseFormatJSONObject", "type": "object", "isBold": false, "description": "JSON object response format.
Note: If the output of JSON is not required by system instructions or user instructions, the model will not actively generate JSON.
", "children": [ { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of response format being defined.
Available options: json_object" } ] } ] } ] }, { "name": "tool_choice", "type": "string", "isBold": true, "required": false, "description": "Controls how the model calls tools.
Note: When a value other than auto is passed to tool_choice, the backend will remove this field by default, and the model response behavior will still be equivalent to the auto mode (this logic is subject to future adjustments).
Available options: auto" }, { "name": "tools", "type": "array", "isBold": true, "required": false, "description": "An array of tools the model may call while generating a response. You can specify which tool to use by setting the tool_choice parameter.
Note: During the multi-turn tool calls process in thinking mode, the model returns the reasoning content alongside the tool calls field. To continue the conversation, it is recommended to keep all previous reasoning content in the input array for each subsequent request to achieve the best performance.
", "children": [ { "name": "Function", "type": "object", "isBold": false, "description": "Defines a function in your own code the model can choose to call.", "children": [ { "name": "name", "type": "string", "isBold": true, "required": true, "description": "The name of the tool function. Must be a-z, A-Z, 0-9, or contain underscores (_) and dashes (-), with a maximum length of 64.
Required string length: 1 - 64" }, { "name": "parameters", "type": "object", "isBold": true, "required": true, "description": "A JSON schema object describing the parameters of the function." }, { "name": "strict", "type": "boolean", "isBold": true, "required": true, "defaultValue": "false", "description": "Whether to enable strict schema adherence when generating the function call." }, { "name": "description", "type": "string", "isBold": true, "required": false, "description": "A description of the function. Used by the model to determine whether or not to call the function." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the tool.
Available options: function" } ] }, { "name": "Custom", "type": "object", "isBold": false, "description": "A custom tool that processes input using a specified format.", "children": [ { "name": "name", "type": "string", "isBold": true, "required": true, "description": "The name of the custom tool, used to identify it in tool calls." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the custom tool.
Available options: custom" }, { "name": "description", "type": "string", "isBold": true, "required": false, "description": "Optional description of the custom tool, used to provide more context." }, { "name": "format", "type": "object", "isBold": true, "required": false, "description": "The input format for the custom tool. Default is unconstrained text.", "children": [ { "name": "CustomToolInputFormat", "type": "object", "isBold": false, "children": [ { "name": "Text", "type": "object", "isBold": false, "description": "Unconstrained free-form text.", "children": [ { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Unconstrained text format.
Available options: text" } ] }, { "name": "Grammar", "type": "object", "isBold": false, "description": "A grammar defined by the user.", "children": [ { "name": "definition", "type": "string", "isBold": true, "required": true, "description": "The grammar definition." }, { "name": "syntax", "type": "string", "isBold": true, "required": true, "description": "The syntax of the grammar definition.
Available options: lark, regex" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Grammar format.
Available options: grammar" } ] } ] } ] } ] }, { "name": "Namespace", "type": "object", "isBold": false, "description": "Groups function tools under a shared namespace.", "children": [ { "name": "description", "type": "string", "isBold": true, "required": true, "description": "A description of the namespace shown to the model." }, { "name": "name", "type": "string", "isBold": true, "required": true, "description": "The namespace name used in tool calls.
MinLength: 1" }, { "name": "tools", "type": "array", "isBold": true, "required": true, "description": "The function tools available inside this namespace.", "children": [ { "name": "Function", "type": "object", "isBold": false, "children": [ { "name": "name", "type": "string", "isBold": true, "required": true, "description": "Required string length: 1 - 64" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Available options: function" }, { "name": "description", "type": [ "string", "null" ], "isBold": true, "required": false }, { "name": "parameters", "type": "object", "isBold": true, "required": false }, { "name": "strict", "type": "boolean", "isBold": true, "required": false } ] }, { "name": "Custom", "type": "object", "isBold": false, "description": "A custom tool that processes input using a specified format.", "children": [ { "name": "name", "type": "string", "isBold": true, "required": true, "description": "The name of the custom tool, used to identify it in tool calls." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the custom tool.
Available options: custom" }, { "name": "description", "type": "string", "isBold": true, "required": false, "description": "Optional description of the custom tool, used to provide more context." }, { "name": "format", "type": "object", "isBold": true, "required": false, "description": "The input format for the custom tool. Default is unconstrained text.", "children": [ { "name": "CustomToolInputFormat", "type": "object", "isBold": false, "children": [ { "name": "Text", "type": "object", "isBold": false, "description": "Unconstrained free-form text.", "children": [ { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Unconstrained text format.
Available options: text" } ] }, { "name": "Grammar", "type": "object", "isBold": false, "description": "A grammar defined by the user.", "children": [ { "name": "definition", "type": "string", "isBold": true, "required": true, "description": "The grammar definition." }, { "name": "syntax", "type": "string", "isBold": true, "required": true, "description": "The syntax of the grammar definition.
Available options: lark, regex" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Grammar format.
Available options: grammar" } ] } ] } ] } ] } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the tool.
Available options: namespace" } ] } ] }, { "name": "top_p", "type": "number", "isBold": true, "required": false, "defaultValue": "0.95", "description": "An alternative to sampling with temperature, called nucleus sampling. We generally recommend altering this or temperature but not both.
In thinking mode, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro and mimo-v2.5 models do not support customizing the top_p parameter. Even if this parameter is passed in, it will be forcibly overridden and take effect with the model's recommended default value of 0.95.
Required range: [0.01, 1.0]" } ]`} /> ## Response Object (non-streaming output) Available options: max_output_tokens, content_filter" } ] }, { "name": "model", "type": "string", "isBold": true, "description": "Model ID used to generate the response." }, { "name": "object", "type": "string", "isBold": true, "description": "Available options: response" }, { "name": "output", "type": "array", "isBold": true, "description": "An array of content items generated by the model.
  • The length and order of items in the output array is dependent on the model’s response.
  • Rather than accessing the first item in the output array and assuming it’s an assistant message with the content generated by the model, you might consider using the output_text property where supported in SDKs.
", "children": [ { "name": "ResponseOutputMessage", "type": "object", "isBold": false, "description": "A message output from the model.", "children": [ { "name": "id", "type": "string", "isBold": true, "description": "The unique ID of the output message." }, { "name": "content", "type": "array", "isBold": true, "description": "The content of the output message.", "children": [ { "name": "ResponseOutputText", "type": "object", "isBold": false, "description": "A text output from the model.", "children": [ { "name": "text", "type": "string", "isBold": true, "description": "The text output from the model." }, { "name": "type", "type": "string", "isBold": true, "description": "Available options: output_text" } ] } ] }, { "name": "role", "type": "string", "isBold": true, "description": "The role of the output message.
Available options: assistant" }, { "name": "status", "type": "string", "isBold": true, "description": "The status of the message.
Available options: in_progress, completed" }, { "name": "type", "type": "string", "isBold": true, "description": "Available options: message" } ] }, { "name": "FunctionCall", "type": "object", "isBold": false, "description": "A tool call to a function tool.", "children": [ { "name": "arguments", "type": "string", "isBold": true, "description": "The arguments that the model generated for the tool call, as a JSON string." }, { "name": "call_id", "type": "string", "isBold": true, "description": "An identifier used when responding to the tool call with output." }, { "name": "name", "type": "string", "isBold": true, "description": "The name of the tool to call." }, { "name": "id", "type": "string", "isBold": true, "description": "The unique ID of the tool call." }, { "name": "namespace", "type": "string", "isBold": true, "description": "The namespace of the function to run." }, { "name": "status", "type": "string", "isBold": true, "description": "The status of the item.
Available options: in_progress, completed" }, { "name": "type", "type": "string", "isBold": true, "description": "Available options: function_call" } ] }, { "name": "FunctionCallOutput", "type": "object", "isBold": false, "description": "The output of a function tool call.", "children": [ { "name": "call_id", "type": "string", "isBold": true, "description": "The unique ID of the function tool call generated by the model." }, { "name": "output", "type": [ "string", "array" ], "isBold": true, "description": "The output from the function call generated by your code. Can be a string or a list of output content.", "children": [ { "name": "StringOutput", "type": "string", "isBold": false, "description": "A string of the output of the function tool call." }, { "name": "OutputContentList", "type": "array", "isBold": false, "description": "Text, image, audio or video output of the function call.", "children": [ { "name": "ResponseInputText", "type": "object", "isBold": false, "description": "A text input to the model.", "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The text input to the model." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_text" } ] }, { "name": "ResponseInputImage", "type": "object", "isBold": false, "description": "An image input to the model.", "children": [ { "name": "image_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the image to be sent to the model. A fully qualified URL or base64 encoded image in a data URL." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_image." } ] }, { "name": "ResponseInputAudioContent", "type": "object", "isBold": false, "description": "An audio input to the model.", "children": [ { "name": "audio_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the audio to be sent to the model. A fully qualified URL or base64 encoded audio in a data URL." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_audio" } ] }, { "name": "ResponseInputVideoContent", "type": "object", "isBold": false, "description": "A video input to the model.", "children": [ { "name": "video_url", "type": "string", "isBold": true, "required": true, "description": "The URL of the video to be sent to the model. A fully qualified URL or base64 encoded video in a data URL." }, { "name": "fps", "type": "number", "isBold": true, "required": false, "defaultValue": "2", "description": "Number of frames sampled per second.
Required range: [0.1, 10.0]" }, { "name": "media_resolution", "type": "string", "isBold": true, "required": false, "defaultValue": "default", "description": "Resolution level.
Available options: default, max" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the input item.
Available options: input_video" } ] } ] } ] }, { "name": "id", "type": "string", "isBold": true, "description": "The unique ID of the function tool call output. Populated when this item is returned via API." }, { "name": "status", "type": "string", "isBold": true, "description": "The status of the item.
Available options: in_progress, completed" }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the function tool call output.
Always function_call_output." } ] }, { "name": "Reasoning", "type": "object", "isBold": false, "description": "A description of the chain of thought used by a reasoning model while generating a response. Be sure to include these items in your input to the Responses API for subsequent turns of a conversation if you are manually managing context.", "children": [ { "name": "id", "type": "string", "isBold": true, "description": "The unique identifier of the reasoning content." }, { "name": "content", "type": "array", "isBold": true, "description": "Reasoning text content.", "children": [ { "name": "text", "type": "string", "isBold": true, "description": "The reasoning text from the model." }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the reasoning text.
Available options: reasoning_text" } ] }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the object.
Available options: reasoning" }, { "name": "status", "type": "string", "isBold": true, "description": "The status of the item.
Available options: in_progress, completed" } ] }, { "name": "CustomToolCall", "type": "object", "isBold": false, "description": "A call to a custom tool created by the model.", "children": [ { "name": "call_id", "type": "string", "isBold": true, "description": "An identifier used to map this custom tool call to a tool call output." }, { "name": "input", "type": "string", "isBold": true, "description": "The input for the custom tool call generated by the model." }, { "name": "name", "type": "string", "isBold": true, "description": "The name of the custom tool being called." }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the custom tool call.
Available options: custom_tool_call" }, { "name": "id", "type": "string", "isBold": true, "description": "The unique ID of the custom tool call." }, { "name": "namespace", "type": "string", "isBold": true, "description": "The namespace of the custom tool being called." } ] }, { "name": "CustomToolCallOutput", "type": "object", "isBold": false, "description": "The output of a custom tool call from your code, being sent back to the model.", "children": [ { "name": "id", "type": "string", "isBold": true, "description": "The unique ID of the custom tool call output." }, { "name": "call_id", "type": "string", "isBold": true, "description": "The call ID, used to map this custom tool call output to a custom tool call." }, { "name": "output", "type": [ "string", "array" ], "isBold": true, "description": "The output from the custom tool call generated by your code. Can be a string or a list of output content.", "children": [ { "name": "StringOutput", "type": "string", "isBold": false, "description": "A string of the output of the custom tool call." }, { "name": "OutputContentList", "type": "array", "isBold": false, "description": "An array of content outputs for the function tool call.", "children": [ { "name": "ResponseInputText", "type": "object", "isBold": false, "description": "A text input to the model.", "children": [ { "name": "text", "type": "string", "isBold": true, "description": "The text input to the model." }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the input item.
Available options: input_text" } ] }, { "name": "ResponseInputImage", "type": "object", "isBold": false, "description": "An image input to the model.", "children": [ { "name": "image_url", "type": "string", "isBold": true, "description": "The URL of the image to be sent to the model. A fully qualified URL or base64 encoded image in a data URL." }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the input item.
Available options: input_image" } ] }, { "name": "ResponseInputAudio", "type": "object", "isBold": false, "description": "An audio input to the model.", "children": [ { "name": "audio_url", "type": "string", "isBold": true, "description": "The URL of the audio to be sent to the model. A fully qualified URL or base64 encoded audio in a data URL." }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the input item.
Available options: input_audio" } ] }, { "name": "ResponseInputVideo", "type": "object", "isBold": false, "description": "A video input to the model.", "children": [ { "name": "video_url", "type": "string", "isBold": true, "description": "The URL of the video to be sent to the model. A fully qualified URL or base64 encoded video in a data URL." }, { "name": "fps", "type": "number", "isBold": true, "defaultValue": "2", "description": "Number of frames sampled per second.
Required range: [0.1, 10.0]" }, { "name": "media_resolution", "type": "string", "isBold": true, "defaultValue": "default", "description": "Resolution level.
Available options: default, max" }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the input item.
Available options: input_video" } ] } ] } ] }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the custom tool call output.
Available options: custom_tool_call_output" } ] } ] }, { "name": "output_text", "type": "string", "isBold": true, "description": "SDK-only convenience property that contains the aggregated text output from all output_text items in the output array, if any are present." }, { "name": "status", "type": "string", "isBold": true, "description": "The status of the response.
Available options: completed, in_progress, incomplete" }, { "name": "usage", "type": "object", "isBold": true, "description": "Usage statistics for the response.", "children": [ { "name": "ResponseUsage", "type": "object", "isBold": false, "children": [ { "name": "input_tokens", "type": "integer", "isBold": true, "description": "The number of input tokens." }, { "name": "input_tokens_details", "type": "object", "isBold": true, "description": "Details about input tokens.", "children": [ { "name": "cached_tokens", "type": "integer", "isBold": true, "description": "The number of cached input tokens." } ] }, { "name": "output_tokens", "type": "integer", "isBold": true, "description": "The number of output tokens." }, { "name": "output_tokens_details", "type": "object", "isBold": true, "description": "Details about output tokens.", "children": [ { "name": "reasoning_tokens", "type": "integer", "isBold": true, "description": "The number of reasoning tokens." } ] }, { "name": "total_tokens", "type": "integer", "isBold": true, "description": "The total number of tokens." } ] } ] } ]`} /> ## Response chunk object (streaming output) {/* feishu-style:text-align:left */} When you create a Response with `stream` set to `true`, the server will emit server-sent events to the client as the Response is generated. ### response.created > An event that is emitted when a response is created. response.created." } ]`} /> ### response.in_progress > Emitted when the response is in progress. response.in_progress." } ]`} /> ### response.completed > Emitted when the model response is complete. response.completed." } ]`} /> ### response.incomplete > An event that is emitted when a response finishes as incomplete. response.incomplete." } ]`} /> ### response.output_item.added > Emitted when a new output item is added. output field returned by the model creation request in non-streaming mode." }, { "name": "output_index", "type": "number", "isBold": true, "description": "The index of the output item that was added." }, { "name": "sequence_number", "type": "number", "isBold": true, "description": "The sequence number of this event." }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the event. Always response.output_item.added." } ]`} /> ### response.output_item.done > Emitted when an output item is marked done. output field returned by the model creation request in non-streaming mode." }, { "name": "output_index", "type": "number", "isBold": true, "description": "The index of the output item that was marked done." }, { "name": "sequence_number", "type": "number", "isBold": true, "description": "The sequence number of this event." }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the event. Always response.output_item.done." } ]`} /> ### response.content_part.added > Emitted when a new content part is added. output_text." } ] }, { "name": "ReasoningText", "type": "object", "isBold": false, "description": "Reasoning text from the model.", "children": [ { "name": "text", "type": "string", "isBold": true, "description": "The reasoning text from the model." }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the reasoning text. Always reasoning_text." } ] } ] }, { "name": "sequence_number", "type": "number", "isBold": true, "description": "The sequence number of this event." }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the event. Always response.content_part.added." } ]`} /> ### response.content_part.done > Emitted when a content part is done. output_text." } ] } ] }, { "name": "sequence_number", "type": "number", "isBold": true, "description": "The sequence number of this event." }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the event. Always response.content_part.done." } ]`} /> ### response.output_text.delta > Emitted when there is an additional text delta. response.output_text.delta." } ]`} /> ### response.output_text.done > Emitted when text content is finalized. response.output_text.done." } ]`} /> ### response.function_call_arguments.delta > Emitted when there is a partial function-call arguments delta. response.function_call_arguments.delta." } ]`} /> ### response.function_call_arguments.done > Emitted when function-call arguments are finalized. response.function_call_arguments.done." } ]`} /> ### response.reasoning_text.delta > Emitted when a delta is added to a reasoning text. response.reasoning_text.delta." } ]`} /> ### response.reasoning_text.done > Emitted when a reasoning text is completed. response.reasoning_text.done." } ]`} /> ### response.custom_tool_call_input.delta > Event representing a delta (partial update) to the input of a custom tool call. response.custom_tool_call_input.delta." } ]`} /> ### response.custom_tool_call_input.done > Event indicating that input for a custom tool call is complete. response.custom_tool_call_input.done." } ]`} /> --- DOCUMENT: Anthropic API --- URL: https://mimo.mi.com/static/docs/api/chat/anthropic-api.md # Anthropic Messages API Compatibility ## Request Address ```bash https://api.xiaomimimo.com/anthropic/v1/messages ``` ## Request Headers {/* feishu-style:text-align:left */} The API supports the following two authentication methods. Please choose one and add it to the request headers: ```json api-key: $MIMO_API_KEY Content-Type: application/json ``` ```json Authorization: Bearer $MIMO_API_KEY Content-Type: application/json ``` ## Request Body role and content.
Each input message content may be either a single string or an array of content blocks, where each block has a specific type. Using a string for content is shorthand for an array of one content block of type text.", "children": [ { "name": "role", "type": "string", "isBold": true, "required": true, "description": "Role of the message.
Available options: user, assistant, system" }, { "name": "content", "type": [ "string", "array" ], "isBold": true, "required": true, "children": [ { "name": "Text content", "type": "string", "isBold": false, "description": "The text contents of the message." }, { "name": "Array of content parts", "type": "array", "isBold": false, "description": "An array of content parts with a defined type. Such text, image, audio, video, tool use, tool result, and thinking.
Currently, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed and mimo-v2.5 models support image, audio or video input.
", "children": [ { "name": "Text", "type": "object", "isBold": false, "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The content of the text block.
Minimum length: 1" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content.
Available options: text" } ] }, { "name": "Image", "type": "object", "isBold": false, "children": [ { "name": "source", "type": "object", "isBold": true, "required": true, "description": "Image data is provided via URL or Base64.", "children": [ { "name": "Base64ImageSource", "type": "object", "isBold": false, "children": [ { "name": "data", "type": "string", "isBold": true, "required": true, "description": "Base64 encoded image data." }, { "name": "media_type", "type": "string", "isBold": true, "required": true, "description": "Media type.
Available options: image/jpeg, image/png, image/gif, image/webp, image/bmp" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Image source type.
Available options: base64" } ] }, { "name": "URLImageSource", "type": "object", "isBold": false, "children": [ { "name": "url", "type": "string", "isBold": true, "required": true, "description": "A URL of the image." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Image source type.
Available options: url" } ] } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content.
Available options: image" } ] }, { "name": "Audio", "type": "object", "isBold": false, "children": [ { "name": "source", "type": "object", "isBold": true, "required": true, "description": "Audio data is provided via URL or Base64.", "children": [ { "name": "Base64AudioSource", "type": "object", "isBold": false, "children": [ { "name": "data", "type": "string", "isBold": true, "required": true, "description": "Base64 encoded audio data." }, { "name": "media_type", "type": "string", "isBold": true, "required": true, "description": "Media type.
Available options: audio/mpeg, audio/wav, audio/flac, audio/mp4, audio/ogg" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Audio source type.
Available options: base64" } ] }, { "name": "URLAudioSource", "type": "object", "isBold": false, "children": [ { "name": "url", "type": "string", "isBold": true, "required": true, "description": "A URL of the audio." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Audio source type.
Available options: url" } ] } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content.
Available options: audio" } ] }, { "name": "Video", "type": "object", "isBold": false, "children": [ { "name": "source", "type": "object", "isBold": true, "required": true, "description": "Video data is provided via URL or Base64.", "children": [ { "name": "Base64VideoSource", "type": "object", "isBold": false, "children": [ { "name": "data", "type": "string", "isBold": true, "required": true, "description": "Base64 encoded video data." }, { "name": "media_type", "type": "string", "isBold": true, "required": true, "description": "Media type.
Available options: video/mp4, video/quicktime, video/x-msvideo, video/x-ms-wmv" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Video source type.
Available options: base64" } ] }, { "name": "URLVideoSource", "type": "object", "isBold": false, "children": [ { "name": "url", "type": "string", "isBold": true, "required": true, "description": "A URL of the video." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Video source type.
Available options: url" } ] }, { "name": "fps", "type": "number", "isBold": true, "required": false, "defaultValue": "2", "description": "Number of frames sampled per second.
Required range: [0.1, 10.0]" }, { "name": "media_resolution", "type": "string", "isBold": true, "required": false, "defaultValue": "default", "description": "Resolution level.
Available options: default, max" } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content.
Available options: video" } ] }, { "name": "Tool use", "type": "object", "isBold": false, "children": [ { "name": "id", "type": "string", "isBold": true, "required": true, "description": "The unique identifier for tool use." }, { "name": "input", "type": "object", "isBold": true, "required": true, "description": "The parameter object passed when using the tool." }, { "name": "name", "type": "string", "isBold": true, "required": true, "description": "Tool name." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content.
Available options: tool_use" } ] }, { "name": "Tool result", "type": "object", "isBold": false, "children": [ { "name": "tool_use_id", "type": "string", "isBold": true, "required": true, "description": "The tool_use ID corresponding to this result." }, { "name": "content", "type": [ "string", "array" ], "isBold": true, "description": "The result returned after the tool is executed.", "children": [ { "name": "Text content", "type": "string", "isBold": false, "description": "The text contents of the message." }, { "name": "Array of content parts", "type": "array", "isBold": false, "description": "An array of content parts with a defined type. Such text, image, audio and video.", "children": [ { "name": "Text", "type": "object", "isBold": false, "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The content of the text block." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content.
Available options: text" } ] }, { "name": "Image", "type": "object", "isBold": false, "children": [ { "name": "source", "type": "object", "isBold": true, "required": true, "description": "Image data is provided via URL or Base64.", "children": [ { "name": "Base64ImageSource", "type": "object", "isBold": false, "children": [ { "name": "data", "type": "string", "isBold": true, "required": true, "description": "Base64 encoded image data." }, { "name": "media_type", "type": "string", "isBold": true, "required": true, "description": "Media type.
Available options: image/jpeg, image/png, image/gif, image/webp, image/bmp" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Image source type.
Available options: base64" } ] }, { "name": "URLImageSource", "type": "object", "isBold": false, "children": [ { "name": "url", "type": "string", "isBold": true, "required": true, "description": "A URL of the image." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Image source type.
Available options: url" } ] } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content.
Available options: image" } ] }, { "name": "Audio", "type": "object", "isBold": false, "children": [ { "name": "source", "type": "object", "isBold": true, "required": true, "description": "Audio data is provided via URL or Base64.", "children": [ { "name": "Base64AudioSource", "type": "object", "isBold": false, "children": [ { "name": "data", "type": "string", "isBold": true, "required": true, "description": "Base64 encoded audio data." }, { "name": "media_type", "type": "string", "isBold": true, "required": true, "description": "Media type.
Available options: audio/mpeg, audio/wav, audio/flac, audio/mp4, audio/ogg" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Audio source type.
Available options: base64" } ] }, { "name": "URLAudioSource", "type": "object", "isBold": false, "children": [ { "name": "url", "type": "string", "isBold": true, "required": true, "description": "A URL of the audio." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Audio source type.
Available options: url" } ] } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content.
Available options: audio" } ] }, { "name": "Video", "type": "object", "isBold": false, "children": [ { "name": "source", "type": "object", "isBold": true, "required": true, "description": "Video data is provided via URL or Base64.", "children": [ { "name": "Base64VideoSource", "type": "object", "isBold": false, "children": [ { "name": "data", "type": "string", "isBold": true, "required": true, "description": "Base64 encoded video data." }, { "name": "media_type", "type": "string", "isBold": true, "required": true, "description": "Media type.
Available options: video/mp4, video/quicktime, video/x-msvideo, video/x-ms-wmv" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Video source type.
Available options: base64" } ] }, { "name": "URLVideoSource", "type": "object", "isBold": false, "children": [ { "name": "url", "type": "string", "isBold": true, "required": true, "description": "A URL of the video." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "Video source type.
Available options: url" } ] }, { "name": "fps", "type": "number", "isBold": true, "required": false, "defaultValue": "2", "description": "Number of frames sampled per second.
Required range: [0.1, 10.0]" }, { "name": "media_resolution", "type": "string", "isBold": true, "required": false, "defaultValue": "default", "description": "Resolution level.
Available options: default, max" } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content.
Available options: video" } ] } ] } ] }, { "name": "is_error", "type": "boolean", "isBold": true }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content.
Available options: tool_result" } ] }, { "name": "Thinking", "type": "object", "isBold": false, "children": [ { "name": "signature", "type": "string", "isBold": true, "description": "The signature of the thinking block." }, { "name": "thinking", "type": "string", "isBold": true, "required": true, "description": "Thinking content." }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content.
Available options: thinking" } ] } ] } ] } ] }, { "name": "model", "type": "string", "isBold": true, "required": true, "description": "The model that will complete your prompt.
Available options: mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro, mimo-v2.5" }, { "name": "max_tokens", "type": "integer", "isBold": true, "required": false, "description": "The maximum number of tokens to generate before stopping.
Note that our models may stop before reaching this maximum. This parameter only specifies the absolute maximum number of tokens to generate.
  • mimo-v2.6-flash: default 131072
  • mimo-v2.6-pro: default 131072
  • mimo-v2.6-pro-ultraspeed: default 131072
  • mimo-v2.5-pro: default 131072
  • mimo-v2.5: default 32768
Required range: [1, 131072]" }, { "name": "stop_sequences", "type": "array", "isBold": true, "required": false, "description": "Custom text sequences that will cause the model to stop generating.
Our models will normally stop when they have naturally completed their turn, which will result in a response stop_reason of end_turn.
If you want the model to stop generating when it encounters custom strings of text, you can use the stop_sequences parameter." }, { "name": "stream", "type": "boolean", "isBold": true, "required": false, "defaultValue": "false", "description": "Whether to incrementally stream the response using server-sent events." }, { "name": "system", "type": [ "string", "array" ], "isBold": true, "required": false, "description": "A system prompt is a way of providing context and instructions to model, such as specifying a particular goal or role.", "children": [ { "name": "Text content", "type": "string", "isBold": false, "description": "The content of the system prompt." }, { "name": "Array of content parts", "type": "array", "isBold": false, "children": [ { "name": "text", "type": "string", "isBold": true, "required": true, "description": "The text content.
Minimum length: 1" }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content.
Available options: text" } ] } ] }, { "name": "temperature", "type": "number", "isBold": true, "required": false, "description": "Sampling temperature controls the diversity of the text generated by the model.
The higher the temperature, the more diverse the generated text will be; conversely, the lower the temperature, the more deterministic the generated text will be.
In thinking mode, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro and mimo-v2.5 models do not support customizing the temperature parameter. Even if this parameter is passed in, it will be forcibly overridden and take effect with the model's recommended default value of 1.0.
  • mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro, mimo-v2.5: default 1.0
Required range: [0, 1.5]" }, { "name": "thinking", "type": "object", "isBold": true, "required": false, "description": "Configuration for enabling model's extended thinking.
Note: During the multi-turn tool calls process in thinking mode, the model returns a thinking content block alongside tool_use content block. To continue the conversation, it is recommended to keep all previous thinking content block in the messages array for each subsequent request to achieve the best performance.
In thinking mode, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro and mimo-v2.5 models do not support customizing the temperature and top_p parameters. Even if these parameters are passed in, the actual effective values will be forcibly set by the model to its recommended default values of 1.0 and 0.95.
", "children": [ { "name": "type", "type": "string", "isBold": true, "required": true, "description": "
  • mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro, mimo-v2.5: default enabled
Available options: enabled, disabled" } ] }, { "name": "tool_choice", "type": "object", "isBold": true, "required": false, "description": "How the model should use the provided tools.", "children": [ { "name": "type", "type": "string", "isBold": true, "required": true, "description": "
  • auto means the model will automatically decide whether to use tools.
Note: When a value other than auto is passed to type, the backend will remove this field by default, and the model response behavior will still be equivalent to the auto mode (this logic is subject to future adjustments).
Available options: auto" }, { "name": "disable_parallel_tool_use", "type": "boolean", "isBold": true, "defaultValue": "false", "description": "Whether to disable parallel tool use.
If set to true:
  • When type is auto, the model will output at most one tool use.
" } ] }, { "name": "tools", "type": "array", "isBold": true, "required": false, "description": "Definitions of tools that the model may use.
If you include tools in your API request, the model may return tool_use content blocks that represent the model's use of those tools. You can then run those tools using the tool input generated by the model and then optionally return results back to the model using tool_result content blocks.
Note: During the multi-turn tool calls process in thinking mode, the model returns a thinking content block alongside tool_use content block. To continue the conversation, it is recommended to keep all previous thinking content block in the messages array for each subsequent request to achieve the best performance.
Each tool definition includes:
  • name: Name of the tool.
  • description: Optional, but strongly-recommended description of the tool.
  • input_schema: JSON schema for the tool input shape that the model will produce in tool_use output content blocks.
", "children": [ { "name": "name", "type": "string", "isBold": true, "required": true, "description": "Name of the tool.
This is how the tool will be called by the model and in tool_use blocks." }, { "name": "description", "type": "string", "isBold": true, "description": "Description of what this tool does.
Tool descriptions should be as detailed as possible. The more information that the model has about what the tool is and how to use it, the better it will perform. You can use natural language descriptions to reinforce important aspects of the tool input JSON schema." }, { "name": "type", "type": "string", "isBold": true, "description": "Available options: custom" }, { "name": "input_schema", "type": "object", "isBold": true, "required": true, "description": "JSON schema for the tool input shape that the model will produce in tool_use output content blocks.", "children": [ { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of input_schema, only object is supported.
Available options: object" }, { "name": "properties", "type": [ "object", "null" ], "isBold": true, "description": "The properties of the tool input." }, { "name": "required", "type": [ "array", "null" ], "isBold": true, "description": "The list of properties that must be included in the tool input." } ] } ] }, { "name": "top_p", "type": "number", "isBold": true, "required": false, "defaultValue": "0.95", "description": "Use nucleus sampling.
In nucleus sampling, we compute the cumulative distribution over all the options for each subsequent token in decreasing probability order and cut it off once it reaches a particular probability specified by top_p. You should either alter temperature or top_p, but not both.
Recommended for advanced use cases only. You usually only need to use temperature.
In thinking mode, the mimo-v2.6-flash, mimo-v2.6-pro, mimo-v2.6-pro-ultraspeed, mimo-v2.5-pro and mimo-v2.5 models do not support customizing the top_p parameter. Even if this parameter is passed in, it will be forcibly overridden and take effect with the model's recommended default value of 0.95.
Required range: [0.01, 1.0]" } ]`} /> ## Non-streaming Response message." }, { "name": "role", "type": "string", "isBold": true, "description": "Conversational role of the generated message. This will always be assistant." }, { "name": "content", "type": "array", "isBold": true, "description": "Content generated by the model.", "children": [ { "name": "Text", "type": "object", "isBold": false, "children": [ { "name": "text", "type": "string", "isBold": true, "description": "The content of the text." }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the content.
Available options: text" } ] }, { "name": "Thinking", "type": "object", "isBold": false, "children": [ { "name": "signature", "type": "string", "isBold": true, "description": "The signature of the thinking block." }, { "name": "thinking", "type": "string", "isBold": true, "description": "Thinking content." }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the content.
Available options: thinking" } ] }, { "name": "Tool use", "type": "object", "isBold": false, "children": [ { "name": "id", "type": "string", "isBold": true, "description": "The unique identifier for tool use." }, { "name": "input", "type": "object", "isBold": true, "description": "The parameter object passed when using the tool." }, { "name": "name", "type": "string", "isBold": true, "description": "Tool name." }, { "name": "type", "type": "string", "isBold": true, "description": "The type of the content.
Available options: tool_use" } ] } ] }, { "name": "model", "type": "string", "isBold": true, "description": "The model that handled the request." }, { "name": "stop_reason", "type": "string", "isBold": true, "description": "The reason the message finished.
This may be one the following values:
  • end_turn: the model reached a natural stopping point.
  • max_tokens: we exceeded the requested max_tokens or the model's maximum.
  • tool_use: the model invoked one or more tools.
  • content_filter: the content was omitted due to a flag from our content filters.
  • repetition_truncation: the model detects repetition.
Available options: end_turn, max_tokens, tool_use, content_filter, repetition_truncation" }, { "name": "usage", "type": "object", "isBold": true, "description": "Billing and rate-limit usage.", "children": [ { "name": "input_tokens", "type": "integer", "isBold": true, "description": "The number of input tokens which were used." }, { "name": "output_tokens", "type": "integer", "isBold": true, "description": "The number of output tokens which were used." }, { "name": "cache_read_input_tokens", "type": [ "integer", "null" ], "isBold": true, "description": "The number of input tokens read from the cache." } ] } ]`} /> ## Streaming Response Available options: message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop" }, { "name": "type", "type": "string", "isBold": true, "description": "Each server-sent event includes a named event type and associated JSON data.
Available options: message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop" }, { "name": "message", "type": "object", "isBold": true, "description": "Response message.", "children": [ { "name": "id", "type": "string", "isBold": true, "description": "The message ID." }, { "name": "type", "type": "string", "isBold": true, "description": "Available options: message" }, { "name": "role", "type": "string", "isBold": true, "description": "Available options: assistant" }, { "name": "model", "type": "string", "isBold": true, "description": "The model name." }, { "name": "content", "type": "array", "isBold": true, "description": "The array of content blocks in the message." }, { "name": "stop_reason", "type": [ "string", "null" ], "isBold": true, "description": "The reason the message finished." } ] }, { "name": "index", "type": "integer", "isBold": true, "description": "The position of the content block within the message.。" }, { "name": "content_block", "type": "object", "isBold": true, "description": "The content block that is starting.", "children": [ { "name": "Text", "type": "object", "isBold": false, "children": [ { "name": "type", "type": "string", "isBold": true, "description": "The header for a text content block; actual text arrives via subsequent delta events.
Available options: text" }, { "name": "text", "type": "string", "isBold": true, "description": "Often an empty string at start; text is appended via content_block_delta events of type text_delta." } ] }, { "name": "Thinking", "type": "object", "isBold": false, "children": [ { "name": "type", "type": "string", "isBold": true, "description": "The header for a thinking content block; actual thinking content arrives via subsequent delta events.
Available options: thinking" }, { "name": "thinking", "type": "string", "isBold": true, "description": "Often an empty string at start; thinking content is appended via content_block_delta events of type thinking_delta." } ] }, { "name": "Tool use", "type": "object", "isBold": false, "children": [ { "name": "type", "type": "string", "isBold": true, "description": "Available options: tool_use" }, { "name": "id", "type": "string", "isBold": true, "description": "The unique identifier for tool use." }, { "name": "name", "type": "string", "isBold": true, "description": "Tool name." }, { "name": "input", "type": "object", "isBold": true, "description": "The parameter object passed when using the tool." } ] } ] }, { "name": "delta", "type": "object", "isBold": true, "description": "Actual response content.", "children": [ { "name": "Content block delta", "type": "object", "isBold": false, "description": "Incremental data for a content block.", "children": [ { "name": "type", "type": "string", "isBold": true, "description": "Available options: text_delta, thinking_delta, input_json_delta" }, { "name": "text", "type": "string", "isBold": true, "description": "The text part of the incremental data." }, { "name": "thinking", "type": "string", "isBold": true, "description": "The thinking part of the incremental data." }, { "name": "partial_json", "type": "string", "isBold": true, "description": "A JSON fragment string. Clients should concatenate fragments in arrival order to form the complete input JSON, then parse." } ] }, { "name": "Message delta", "type": "object", "isBold": false, "description": "Message-level stop metadata updates.", "children": [ { "name": "stop_reason", "type": [ "string", "null" ], "isBold": true, "description": "The reason the message finished.
Available options: end_turn, max_tokens, tool_use, content_filter, repetition_truncation" } ] } ] }, { "name": "usage", "type": [ "object", "null" ], "isBold": true, "description": "Billing and rate-limit usage.", "children": [ { "name": "input_tokens", "type": "integer", "isBold": true, "description": "The number of input tokens which were used." }, { "name": "output_tokens", "type": "integer", "isBold": true, "description": "The number of output tokens which were used." }, { "name": "cache_read_input_tokens", "type": [ "integer", "null" ], "isBold": true, "description": "The number of input tokens read from the cache." } ] } ]`} /> --- DOCUMENT: Speech Recognition (MiMo‑V2.5-ASR) - OpenAI API Compatibility --- URL: https://mimo.mi.com/static/docs/api/audio/Speech-Recognition.md # Speech Recognition (MiMo‑V2.5-ASR) - OpenAI API Compatibility ## Request Address ```bash https://api.xiaomimimo.com/v1/chat/completions ``` ## Request Headers {/* feishu-style:text-align:left */} The API supports the following two authentication methods. Please choose one and add it to the request headers: ```json api-key: $MIMO_API_KEY Content-Type: application/json ``` ```json Authorization: Bearer $MIMO_API_KEY Content-Type: application/json ``` ## Request body
For detailed usage, please refer to Speech Recognition.
", "children": [ { "name": "Array of content parts", "type": "array", "isBold": false, "description": "An array of content parts with a defined type. For speech recognition, only single audio input is supported.", "children": [ { "name": "Audio content part", "type": "object", "isBold": false, "children": [ { "name": "input_audio", "type": "object", "isBold": true, "required": true, "description": "
When audio is passed via data URL, the format field is optional. If only Base64-encoded audio data is provided, the format field is required. If both MIME_TYPE and format are included, their values must match.
", "children": [ { "name": "data", "type": "string", "isBold": true, "required": true, "description": "Base64 encoded audio in a data URL. Input audio only supports mp3 and wav formats:
  • mp3: valid MIME_TYPE values: audio/mpeg, audio/mp3
  • wav: valid MIME_TYPE value: audio/wav
" }, { "name": "format", "type": "string", "isBold": true, "description": "The format for encoding audio data.
Available options: mp3, wav" } ] }, { "name": "type", "type": "string", "isBold": true, "required": true, "description": "The type of the content part.
Available options: input_audio" } ] } ] } ] }, { "name": "role", "type": "string", "isBold": true, "required": true, "description": "Role of the message author.
Available options: user" } ] } ] }, { "name": "model", "type": "string", "isBold": true, "required": true, "description": "Model ID is used to generate the response.
Available options: mimo-v2.5-asr" }, { "name": "asr_options", "type": "object", "isBold": true, "required": false, "description": "Custom configuration parameters for automatic speech recognition (ASR).", "children": [ { "name": "language", "type": "string", "isBold": true, "required": false, "defaultValue": "auto", "description": "Specify a single language for audio recognition.
  • auto: Auto‑detect audio language
  • zh: Chinese
  • en: English
Available options: auto, zh, en" } ] }, { "name": "stream", "type": "boolean", "isBold": true, "required": false, "defaultValue": "false", "description": "If set to true, the model response data will be streamed to the client as it is generated using server-sent events." } ]`} /> ## Chat response object (non-streaming output)
  • stop: The model reached a natural stop point or a user‑provided stop sequence
  • length: Terminated due to exceeding the model's maximum generation length
  • content_filter: Content was omitted due to a content filter flag
" }, { "name": "index", "type": "integer", "isBold": true, "description": "The index of the choice in the list of choices." }, { "name": "message", "type": "object", "isBold": true, "description": "A chat completion message generated by the model.", "children": [ { "name": "content", "type": "string", "isBold": true, "description": "The contents of the message." }, { "name": "role", "type": "string", "isBold": true, "description": "The role of the author of this message." } ] } ] }, { "name": "created", "type": "integer", "isBold": true, "description": "The Unix timestamp (in seconds) of when the chat completion was created." }, { "name": "id", "type": "string", "isBold": true, "description": "A unique identifier for the chat completion." }, { "name": "model", "type": "string", "isBold": true, "description": "The model to generate the completion." }, { "name": "object", "type": "string", "isBold": true, "description": "The object type, which is always chat.completion." }, { "name": "usage", "type": [ "object", "null" ], "isBold": true, "description": "Usage statistics for the completion request.", "children": [ { "name": "completion_tokens", "type": "integer", "isBold": true, "description": "Number of tokens in the generated completion." }, { "name": "prompt_tokens", "type": "integer", "isBold": true, "description": "Number of tokens in the prompt." }, { "name": "total_tokens", "type": "integer", "isBold": true, "description": "Total number of tokens used in the request (prompt + completion)." }, { "name": "completion_tokens_details", "type": "object", "isBold": true, "description": "Breakdown of tokens used in a completion.", "children": [ { "name": "reasoning_tokens", "type": "integer", "isBold": true, "description": "Tokens generated by the model for reasoning. Always 0." } ] }, { "name": "prompt_tokens_details", "type": "object", "isBold": true, "description": "Breakdown of tokens used in the prompt.", "children": [ { "name": "cached_tokens", "type": "integer", "isBold": true, "description": "Number of tokens served from cache." }, { "name": "audio_tokens", "type": "integer", "isBold": true, "description": "Audio input tokens present in the prompt." } ] }, { "name": "seconds", "type": "integer", "isBold": true, "description": "Audio duration (seconds)." } ] } ]`} /> ## Chat response chunk object (streaming output)
  • stop: The model reached a natural stop point or a user‑provided stop sequence
  • length: Terminated due to exceeding the model's maximum generation length
  • content_filter: Content was omitted due to a content filter flag
" }, { "name": "index", "type": "integer", "isBold": true, "description": "The index of the choice in the list of choices." } ] }, { "name": "created", "type": "integer", "isBold": true, "description": "The Unix timestamp (in seconds) of when the chat completion was created. Each chunk has the same timestamp." }, { "name": "id", "type": "string", "isBold": true, "description": "A unique identifier for the chat completion. Each chunk has the same ID." }, { "name": "model", "type": "string", "isBold": true, "description": "The model to generate the completion." }, { "name": "object", "type": "string", "isBold": true, "description": "The object type, which is always chat.completion.chunk." }, { "name": "usage", "type": [ "object", "null" ], "isBold": true, "description": "Usage statistics for the completion request.", "children": [ { "name": "completion_tokens", "type": "integer", "isBold": true, "description": "Number of tokens in the generated completion." }, { "name": "prompt_tokens", "type": "integer", "isBold": true, "description": "Number of tokens in the prompt." }, { "name": "total_tokens", "type": "integer", "isBold": true, "description": "Total number of tokens used in the request (prompt + completion)." }, { "name": "completion_tokens_details", "type": "object", "isBold": true, "description": "Breakdown of tokens used in a completion.", "children": [ { "name": "reasoning_tokens", "type": "integer", "isBold": true, "description": "Tokens generated by the model for reasoning. Always 0." } ] }, { "name": "prompt_tokens_details", "type": "object", "isBold": true, "description": "Breakdown of tokens used in the prompt.", "children": [ { "name": "cached_tokens", "type": "integer", "isBold": true, "description": "Number of tokens served from cache." }, { "name": "audio_tokens", "type": "integer", "isBold": true, "description": "Audio input tokens present in the prompt." } ] }, { "name": "seconds", "type": "integer", "isBold": true, "description": "Audio duration (seconds)." } ] } ]`} /> --- DOCUMENT: Speech Synthesis (MiMo-TTS Series) - OpenAI API Compatibility --- URL: https://mimo.mi.com/static/docs/api/audio/tts.md # Speech Synthesis (MiMo-TTS Series) - OpenAI API Compatibility ## Request Address ```bash https://api.xiaomimimo.com/v1/chat/completions ``` ## Request Headers {/* feishu-style:text-align:left */} The API supports the following two authentication methods. Please choose one and add it to the request headers: ```json api-key: $MIMO_API_KEY Content-Type: application/json ``` ```json Authorization: Bearer $MIMO_API_KEY Content-Type: application/json ``` ## Request body
Note: When generating audio using the mimo-v2.5-tts-voicedesign model, this message is required and is used to specify the text describing the voice design.
", "children": [ { "name": "content", "type": "string", "isBold": true, "required": true, "description": "The contents of the user message." }, { "name": "role", "type": "string", "isBold": true, "required": true, "description": "Role of the message author.
Available options: user" } ] }, { "name": "Assistant message", "type": "object", "isBold": false, "description": "Messages sent by the model in response to user messages.
Note: When using the mimo-v2.5-tts-voicedesign model and optimize_text_preview is true, the assistant message is optional; in other cases, it is required.
", "children": [ { "name": "content", "type": "string", "isBold": true, "required": true, "description": "The contents of the assistant message, which is used to specify the target text for audio synthesis." }, { "name": "role", "type": "string", "isBold": true, "required": true, "description": "Role of the message author.
Available options: assistant" } ] } ] }, { "name": "model", "type": "string", "isBold": true, "required": true, "description": "Model ID is used to generate the response.
Available options: mimo-v2.5-tts, mimo-v2.5-tts-voicedesign, mimo-v2.5-tts-voiceclone" }, { "name": "audio", "type": "object", "isBold": true, "required": false, "description": "Parameters for audio output. For details, please refer to Speech Synthesis.
Note: To generate audio, you must add a message with role set to assistant, which needs to specify the text for speech synthesis. Additionally, when using the mimo-v2.5-tts-voicedesign model, a message with the role of user is required. If optimize_text_preview is set to true, the assistant message can be omitted.
", "children": [ { "name": "format", "type": "string", "isBold": true, "required": false, "defaultValue": "wav", "description": "Specifies the output audio format. Default: wav, or pcm when you set stream: true.
Passing in pcm or pcm16 both indicate specifying the use of the pcm16 format.
Available options: wav, mp3, pcm, pcm16" }, { "name": "optimize_text_preview", "type": "boolean", "isBold": true, "required": false, "defaultValue": "false", "description": "Enables intelligent optimization of the target audio broadcast text.
When set to true, the input target text is intelligently polished; if no target text is provided, a broadcast-adapted target text is automatically generated. The finalized processed text is then fed into the model for speech synthesis.
Note: When this parameter is set to true, the assistant role message for specifying speech synthesis content can be omitted.
Currently, only the mimo-v2.5-tts-voicedesign model is supported.
" }, { "name": "voice", "type": "string", "isBold": true, "description": "The voice ID of the built-in voice or the base64 encoding of the audio sample.
  • mimo-v2.5-tts: This field is optional and only supports using built-in voices, with the default value being mimo_default
  • mimo-v2.5-tts-voiceclone: This field is required and only supports passing in the base64 encoding of audio samples, and only supports passing in audio sample files in mp3 and wav formats
  • mimo-v2.5-tts-voicedesign does not support this field
Available options:
  • mimo-v2.5-tts: mimo_default, 冰糖, 茉莉, 苏打, 白桦, Mia, Chloe, Milo, Dean
" } ] }, { "name": "stream", "type": "boolean", "isBold": true, "required": false, "defaultValue": "false", "description": "If set to true, the model response data will be streamed to the client as it is generated using server-sent events." } ]`} /> ## Chat response object (non-streaming output)
  • stop: The model reached a natural stop point or a user‑provided stop sequence
  • length: Terminated due to exceeding the model's maximum generation length
  • content_filter: Content was omitted due to a content filter flag
" }, { "name": "index", "type": "integer", "isBold": true, "description": "The index of the choice in the list of choices." }, { "name": "message", "type": "object", "isBold": true, "description": "A chat completion message generated by the model.", "children": [ { "name": "content", "type": "string", "isBold": true, "description": "The contents of the message." }, { "name": "role", "type": "string", "isBold": true, "description": "The role of the author of this message." }, { "name": "audio", "type": "object", "isBold": true, "description": "If the audio output is requested, this object contains data about the audio response from the model.", "children": [ { "name": "id", "type": "string", "isBold": true, "description": "Unique identifier for this audio response." }, { "name": "data", "type": "string", "isBold": true, "description": "Base64 encoded audio bytes generated by the model, in the format specified in the request." }, { "name": "expires_at", "type": [ "number", "null" ], "isBold": true, "description": "The Unix timestamp (in seconds) for when this audio response expires. Currently always null." }, { "name": "transcript", "type": [ "string", "null" ], "isBold": true, "description": "Transcript of the audio generated by the model. Currently always null." } ] }, { "name": "final_text_preview", "type": "string", "isBold": true, "description": "The final audio broadcast text after intelligent optimization and polishing. This field is only returned when the request parameter optimize_text_preview is set to true." } ] } ] }, { "name": "created", "type": "integer", "isBold": true, "description": "The Unix timestamp (in seconds) of when the chat completion was created." }, { "name": "id", "type": "string", "isBold": true, "description": "A unique identifier for the chat completion." }, { "name": "model", "type": "string", "isBold": true, "description": "The model to generate the completion." }, { "name": "object", "type": "string", "isBold": true, "description": "The object type, which is always chat.completion." }, { "name": "usage", "type": [ "object", "null" ], "isBold": true, "description": "Usage statistics for the completion request.", "children": [ { "name": "completion_tokens", "type": "integer", "isBold": true, "description": "Number of tokens in the generated completion." }, { "name": "prompt_tokens", "type": "integer", "isBold": true, "description": "Number of tokens in the prompt." }, { "name": "total_tokens", "type": "integer", "isBold": true, "description": "Total number of tokens used in the request (prompt + completion)." }, { "name": "completion_tokens_details", "type": "object", "isBold": true, "description": "Breakdown of tokens used in a completion.", "children": [ { "name": "reasoning_tokens", "type": "integer", "isBold": true, "description": "Tokens generated by the model for reasoning. Always 0." } ] }, { "name": "prompt_tokens_details", "type": "object", "isBold": true, "description": "Breakdown of tokens used in the prompt.", "children": [ { "name": "cached_tokens", "type": "integer", "isBold": true, "description": "Number of tokens served from cache." } ] } ] } ]`} /> ## Chat response chunk object (streaming output) null." }, { "name": "transcript", "type": [ "string", "null" ], "isBold": true, "description": "Transcript of the audio generated by the model. Currently always null." } ] }, { "name": "final_text_preview", "type": "string", "isBold": true, "description": "The final audio broadcast text after intelligent optimization and polishing. This field is only returned when the request parameter optimize_text_preview is set to true." } ] }, { "name": "finish_reason", "type": [ "string", "null" ], "isBold": true, "description": "The reason the model stopped generating tokens:
  • stop: The model reached a natural stop point or a user‑provided stop sequence
  • length: Terminated due to exceeding the model's maximum generation length
  • content_filter: Content was omitted due to a content filter flag
" }, { "name": "index", "type": "integer", "isBold": true, "description": "The index of the choice in the list of choices." } ] }, { "name": "created", "type": "integer", "isBold": true, "description": "The Unix timestamp (in seconds) of when the chat completion was created. Each chunk has the same timestamp." }, { "name": "id", "type": "string", "isBold": true, "description": "A unique identifier for the chat completion. Each chunk has the same ID." }, { "name": "model", "type": "string", "isBold": true, "description": "The model to generate the completion." }, { "name": "object", "type": "string", "isBold": true, "description": "The object type, which is always chat.completion.chunk." }, { "name": "usage", "type": [ "object", "null" ], "isBold": true, "description": "Usage statistics for the completion request.", "children": [ { "name": "completion_tokens", "type": "integer", "isBold": true, "description": "Number of tokens in the generated completion." }, { "name": "prompt_tokens", "type": "integer", "isBold": true, "description": "Number of tokens in the prompt." }, { "name": "total_tokens", "type": "integer", "isBold": true, "description": "Total number of tokens used in the request (prompt + completion)." }, { "name": "completion_tokens_details", "type": "object", "isBold": true, "description": "Breakdown of tokens used in a completion.", "children": [ { "name": "reasoning_tokens", "type": "integer", "isBold": true, "description": "Tokens generated by the model for reasoning. Always 0." } ] }, { "name": "prompt_tokens_details", "type": "object", "isBold": true, "description": "Breakdown of tokens used in the prompt.", "children": [ { "name": "cached_tokens", "type": "integer", "isBold": true, "description": "Number of tokens served from cache." } ] } ] } ]`} /> --- DOCUMENT: List Models --- URL: https://mimo.mi.com/static/docs/api/model/list-models.md # List Models ## Request Address ```bash https://api.xiaomimimo.com/v1/models ``` ## Request Headers {/* feishu-style:text-align:left */} The API supports the following two authentication methods. Please choose one and add it to the request headers: ```json api-key: $MIMO_API_KEY ``` ```json Authorization: Bearer $MIMO_API_KEY ``` ## Response Available options: list" }, { "name": "data", "type": "array", "isBold": true, "description": "An array of model objects.", "children": [ { "name": "id", "type": "string", "isBold": true, "description": "The model identifier, which can be referenced in the API endpoints." }, { "name": "object", "type": "string", "isBold": true, "description": "The object type.
Available options: model" }, { "name": "owned_by", "type": "string", "isBold": true, "description": "The organization that owns the model." } ] } ]`} /> --- DOCUMENT: Pay‑As‑You‑Go API --- URL: https://mimo.mi.com/static/docs/price/pay-as-you-go.md # API Pricing {/* feishu-style:text-align:left */} **Pay-as-you-go for the Xiaomi MiMo API uses the ordinary API Key of the Open Platform and consumes the account balance based on the actual Token usage, which is not interoperable with the Token Plan package quota.**
**Billing Instructions** - Billing Unit: China:RMB / M tokens; Overseas: dollar / M tokens - Cache Hit: When the requested prefix content hits the Prompt Cache, it will be billed according to the cached hit price. - Cache Write: Limited-time Free - ASR series models are billed based on the duration of the input audio: duration statistics are accurate to the second, and ultimately converted to hourly billing - Internet search is billed independently based on the number of calls and is not included in the Token price
### Domestic Pricing of the Model
- `mimo-v2.5-pro` and `mimo-v2.5` **will be officially deprecated at 10:00 (Beijing time) on October 21,2026. It is recommended to switch to the new version of the models as soon as possible. For details, please refer to** [Model Deprecation](https://mimo.mi.com/docs/zh-CN/updates/deprecate).
{/* feishu-style:text-align:left */} **Language Models**
**Inference Type** **Model Name** **Input (Cache Hit)** **Input (Cache Miss)** **Output**
**Real-time API** `mimo-v2.6-pro`、`mimo-v2.5-pro`(to be deprecated) ¥0.025 ¥3.00 ¥6.00
`mimo-v2.6-flash`、`mimo-v2.5`(to be deprecated) ¥0.02 ¥1.00 ¥2.00
`mimo-v2.6-pro-ultraspeed` ¥0.25 ¥30.00 ¥60.00
**Batch API** `mimo-v2.6-pro` ¥0.0125 ¥1.50 ¥3.00
`mimo-v2.6-flash` ¥0.01 ¥0.50 ¥1.00
> `mimo-v2.6-pro-ultraspeed`, `mimo-v2.5-pro`, `mimo-v2.5 `do not support batch API. {/* feishu-style:text-align:left */} **ASR Series**
**Model Name** **Input audio duration**
`mimo-v2.5-asr` ¥0.5 /h
{/* feishu-style:text-align:left */} **TTS Series** {/* feishu-style:text-align:left */} `mimo-v2.5-tts`, `mimo-v2.5-tts-voiceclone`, `mimo-v2.5-tts-voicedesign` are free for a limited time {/* feishu-style:text-align:left */}

### Overseas Pricing of the Model
- `mimo-v2.5-pro` and `mimo-v2.5` **will be officially deprecated at 10:00 (Beijing time) on October 21,2026. It is recommended to switch to the new version of the models as soon as possible. For details, please refer to** [Model Deprecation](https://mimo.mi.com/docs/zh-CN/updates/deprecate).
{/* feishu-style:text-align:left */} **Language Models**
**Inference Type** **Model Name** **Input (Cache Hit)** **Input (Cache Miss)** **Output**
**Real-time API** `mimo-v2.6-pro`、`mimo-v2.5-pro`(to be deprecated) $0.0036 $0.435 $0.87
`mimo-v2.6-flash`、`mimo-v2.5`(to be deprecated) $0.0028 $0.14 $0.28
`mimo-v2.6-pro-ultraspeed` $0.036 $4.35 $8.7
**Batch API** `mimo-v2.6-pro` $0.0018 $0.2175 $0.435
`mimo-v2.6-flash` $0.0014 $0.07 $0.14
{/* feishu-style:text-align:left */} `mimo-v2.6-pro-ultraspeed`, `mimo-v2.5-pro`, `mimo-v2.5 `do not support batch API. {/* feishu-style:text-align:left */} **ASR Series**
**Model Name** **Input audio duration**
`mimo-v2.5-asr` $0.074 /h
{/* feishu-style:text-align:left */} **TTS Series** {/* feishu-style:text-align:left */} `mimo-v2.5-tts`, `mimo-v2.5-tts-voiceclone`, `mimo-v2.5-tts-voicedesign` are free for a limited time {/* feishu-style:text-align:left */}

### Pricing for Web Search Plugins
**Service Item** **Price** **Description**
Domestic Internet Connectivity Service ¥16 /1000 times Includes web search and web parsing, used for searching relevant content in domestic regional network connections
Overseas Internet Connectivity Service $5 /1000 times Includes web search and web parsing, used for networked search of relevant content in overseas regions
--- DOCUMENT: Token Plan --- URL: https://mimo.mi.com/static/docs/price/token-plan.md # Token Plan {/* feishu-style:text-align:left */} **Token Plan** is an exclusive subscription solution launched for AI programming scenarios. You can use the cost-effective subscription resource package to access MiMo's flagship large language model in various mainstream AI development tools. The platform currently offers two versions: the individual version and the team version. ## Core Strengths - **Covers flagship models** — all plans support the latest flagship models mimo-v2.6-pro, mimo-v2.6-flash, as well as ASR and TTS models. Adopts a token conversion mechanism with transparent and controllable quota - **Flexible Subscription Plans** — Tiered packages are available for both individual and team versions to meet the diverse development needs of individual and team developers - **Multi-ecosystem Out Of The Box** — Compatible with mainstream development toolchains including OpenCode, OpenClaw, Claude Code, etc. - **Exclusive Benefits for Team Plan** — In addition to inheriting all the benefits of the Individual Plan, users of the Team Plan can enjoy exclusive privileges including unified management of storage seats, team usage control and analysis, as well as centralized billing and invoice management ## Pricing and Quotas for Individual Edition #### Monthly Plan
**Lite** **Standard** **Pro** **Max**
**Pricing** $6/month, ¥39/month $16/month, ¥99/month $50/month, ¥329/month $100/month, ¥659/month
**Fixed Monthly Quota** 4.1 billion Credits 11 billion Credits 38 billion Credits 82 billion Credits
#### Annual Package
**Lite** **Standard** **Pro** **Max**
**Pricing** $63.36/year, ¥411.84/year USD 168.96/year, CNY 1045.44/year $528.00/year, ¥3474.24/year USD 1,056.00/year, CNY 6,959.04/year
**Annual Fixed Quota** 49.2 billion Credits 132 billion Credits 456 billion Credits 984 billion Credits
## Pricing and Quotas for Team Edition #### Monthly Plan
**Standard** **Pro** **Max**
**Pricing** $16/seat/month, ¥99/seat/month $50/seat/month, ¥329/seat/month $100/seat/month, ¥659/seat/month
**Fixed Monthly Quota** 11 billion Credits 38 billion Credits 82 billion Credits
#### Annual Package
**Standard** **Pro** **Max**
**Pricing** USD 168.96/seat/year, CNY 1,044/seat/year USD 528/seat/year, CNY 3,468/seat/year USD 1,056/seat/year, CNY 6,948/seat/year
**Fixed monthly quota per seat** 11 billion Credits 38 billion Credits 82 billion Credits
## Applicable Scenarios
**Lite (Individual Only)** **Standard** **Pro** **Max**
**Applicable Scenarios** Ideal for first-time users who want to try lobster
Using mimo-v2.6-flash as the baseline, it can execute approximately **200 rounds of medium-to-complex tasks**
Suitable for office workers who frequently use AI to boost their work efficiency
Using mimo-v2.6-flash as the baseline, it can execute approximately **1600 rounds of medium-to-complex tasks**
Ideal for developers and professional productivity enthusiasts who use AI frequently on a daily basis
Using mimo-v2.6-flash as the baseline, it can execute approximately **5600 rounds of medium-to-complex tasks**
Ideal for high-intensity, hardcore users who treat AI as a core productivity tool
Using mimo-v2.6-flash as the baseline, it can execute approximately **12800 rounds of medium-to-complex tasks**
> The above describes the applicable scenario scope of the monthly plan, and the order of magnitude of task processing for the annual plan is approximately 12 times that of the monthly plan. ## Quota Consumption Rules > The quota consumption rules for TokenPlan are the same for both the individual version and the team version {/* feishu-style:text-align:left */} Language models deduct Credits quota based on the number of tokens; ASR models deduct Credits quota according to the duration of the input audio (the duration is counted to the exact second and converted to hourly billing); TTS series models are free for a limited time and do not consume package Credits. The specific conversion rules are as follows. {/* feishu-style:text-align:left */} Language Model
model Input (Cache Hit) Token Input (cache miss) Token Output Token
mimo-v2.6-pro 2.5 Credits 300 Credits 600 Credits
mimo-v2.6-flash 2 Credits 100 Credits 200 Credits
mimo-v2.5-pro 2.5 Credits 300 Credits 600 Credits
mimo-v2.5 2 Credits 100 Credits 200 Credits
{/* feishu-style:text-align:left */} ASR Model
model Input audio duration (h)
mimo-v2.5-asr 30M Credits
{/* feishu-style:text-align:left */} The TTS series models are available for free for a limited time and will not consume your plan credits.
**The Individual edition additionally supports mimo-v2.5-pro and mimo-v2.5.Both of these two models will be officially taken offline at 10:00 on October 21, 2026 Beijing Time, and it is recommended to switch to the new version of the models as soon as possible.**
{/* feishu-style:text-align:left */} For example, if you subscribe to the Lite plan (4.1B Credits), you can call the MiMo- V2.6 series models either individually or in combination. After you use 10M mimo-v2.6-pro input tokens (cache miss), this is equivalent to consuming 3000 M Credit s, and you will still have 1100M mimo-v2.5-flash Credits remaining.In addition, if you dedicate the entire quota of the Lite plan exclusively to the ASR model, you will get 4100M per month ÷ 30M/hour = 136.6 hours (equivalent to processing 4.5 hours of audio per day for 1 consecutive month). You can check the quota and usage of your current plan at [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage). - **Quota Exhaustion:** When the monthly total quota of the package is used up, the system will suspend the service and will not continue to deduct from your bonus or account balance. - **To continue using the service:** **For the Individual Plan,** please purchase an upgrade package to unlock new package resources, or switch to the standard API which charges based on per-token unit price, so that you can continue using the service without usage restrictions. **For the Team Plan**, please purchase additional seats and reallocate them, or switch to the standard API. ## Discount Offers - First Purchase Discount: Users purchasing the Individual Edition for the first time can enjoy a 12% discount, and each account is only eligible for this discount once; the Team Edition is not eligible for this discount. - Annual Subscription (Auto-renewal): Compared with the monthly auto-renewal plan, both the Individual Plan and Team Plan enjoy a 12% discount for annual auto-renewal, and first-purchase offers are not applicable to annual auto-renewal plans; - Night Discount Rate: During off-peak hours (Beijing Time 00:00–08:00, i. e. UTC 16:00–24:00), the consumption coefficient is 0.8x. ## For more details {/* feishu-style:text-align:left */} For more information about TokenPlan, please refer to [TokenPlan](https://platform.xiaomimimo.com/#/docs/faq). --- DOCUMENT: Individual Subscription --- URL: https://mimo.mi.com/static/docs/tokenplan/Token Plan/subscription.md # Individual Subscription {/* feishu-style:text-align:left */} Token Plan Individual Edition is an exclusive subscription solution launched for individual developers in AI programming scenarios, allowing you to use the cost-effective subscription resource package to call MiMo's flagship large language model in various mainstream AI development tools. ## Core Strengths - **Covers flagship models** — all plans support mimo-v2.6-pro, mimo-v2.6-flash, mimo-v2.5-pro, mimo-v2.5, mimo-v2.5-asr, mimo-v2.5-tts-voiceclone, mimo-v2.5-tts-voicedesign, mimo-v2.5-tts, a total of 8 models. It adopts a token conversion mechanism with transparent and controllable quotas. - **Flexible Subscription Plan** — Four-tier packages to meet various development needs of individual developers - **Multi-ecosystem Out Of The Box** — Compatible with mainstream development toolchains including OpenCode, OpenClaw, Claude Code, etc.
**The two models mimo-v2.5-pro and mimo-v2.5 will be officially taken offline at 10:00 (Beijing time) on October 21, 2026. It is recommended to switch to the new version of the models as soon as possible.**
## Quota Limit #### Monthly Plan
**Lite** **Standard** **Pro** **Max**
**Pricing** $6/month, ¥39/month $16/month, ¥99/month $50/month, ¥329/month $100/month, ¥659/month
**Fixed Monthly Quota** 4.1 billion Credits 11 billion Credits 38 billion Credits 82 billion Credits
#### Annual Package
**Lite** **Standard** **Pro** **Max**
**Pricing** $63.36/year, ¥411.84/year $168.96/year, ¥1045.44/year USD 528.00/year, CNY 3474.24/year USD 1,056.00/year, CNY 6,959.04/year
**Annual Fixed Quota** 49.2 billion Credits 132 billion Credits 456 billion Credits 984 billion Credits
#### Applicable Scenarios
**Lite** **Standard** **Pro** **Max**
**Applicable Scenarios** Ideal for first-time users who want to try lobster
Using mimo-v2.6-flash as the baseline, it can execute approximately **200 rounds of medium-to-complex tasks**
Ideal for office professionals who regularly use AI to boost their work efficiency
Using mimo-v2.6-flash as the baseline, it can execute approximately **1600 rounds of medium-to-complex tasks**
Ideal for developers and professional productivity enthusiasts who use AI frequently on a daily basis
Using mimo-v2.6-flash as the baseline, it can execute approximately **5600 rounds of medium-to-complex tasks**
Ideal for high-intensity, hardcore users who treat AI as a core productivity tool
Using mimo-v2.6-flash as the baseline, it can execute approximately **12800 rounds of medium-to-complex tasks**
> The above describes the applicable scenario scope of the monthly package; the task processing magnitude of the annual package is approximately 12 times that of the monthly package. #### Quota Consumption Rules {/* feishu-style:text-align:left */} Language models deduct Credit quota based on the number of tokens; ASR models deduct Credit quota according to the duration of the input audio (the duration is counted to the nearest second and finally converted to hours for statistics); TTS series models are free for a limited time and do not consume package Credit. The specific conversion rules are as follows. {/* feishu-style:text-align:left */} Language Model
model Input (Cache Hit) Token Input (cache miss) Token Output Token
mimo-v2.6-pro 2.5 Credits 300 Credits 600 Credits
mimo-v2.6-flash 2 Credits 100 Credits 200 Credits
mimo-v2.5-pro 2.5 Credits 300 Credits 600 Credits
mimo-v2.5 2 Credits 100 Credits 200 Credits
{/* feishu-style:text-align:left */} ASR Model
model Input audio duration (h)
mimo-v2.5-asr 30M Credits
{/* feishu-style:text-align:left */} The TTS series models are available for free for a limited time and will not consume your package credits. {/* feishu-style:text-align:left */} For example, if you subscribe to the Lite plan (4.1B Credits), you can call MiMo- V2.6 series models either individually or in combination. After you use 10M input (cache miss) tokens of mimo-v2.6-pro, this is equivalent to consuming 3000M Credit s, and you will still have a remaining Credits quota of 1100M for mimo-v2.6-flash.In addition, if you dedicate the entire quota of the Lite plan exclusively to the ASR model, you will get 4100M ÷ 30M/hour = 136.6 hours per month (equivalent to processing 4.5 hours of audio per day for consecutive use over 1 month). You can check the quota and usage of your current plan at [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage). - **Quota Exhaustion:** When the monthly total quota of your plan is used up, the system will suspend the service and will not continue to deduct from your bonus or account balance. - **To continue using the service:** please purchase an upgrade plan to unlock new plan resources, or switch to the standard API which is billed based on per-token unit price, so that you can continue using the service without usage restrictions. ## Package Purchase - **Discounts: 12% off for your first package purchase, 12% off for annual auto-renewal, and 20% lower consumption during nighttime** - First Purchase Discount: Enjoy a 12% discount on your first purchase, with only one discount eligible per account; - Annual auto-renewal plan: Enjoy a 12% discount compared to the monthly auto-renewal plan; first-purchase offers are not applicable to the annual plan. - Off-peak discount rate: During non-peak hours (00:00–08:00 Beijing Time, equivalent to 16:00–24:00 UTC), a consumption coefficient of 0.8x applies. - **Top-up upgrade for plans by paying price difference: Currently, the platform only supports purchasing one individual plan at a time. If you wish to get more quota before your current plan expires**, the Credit quota you have already used will be converted into an equivalent amount, and you can pay the price difference on this basis to upgrade to a higher plan for more Credits. Cross-tier plan upgrade via price difference top-up is supported, while plan downgrade is not allowed. If you have already upgraded to the top-tier Max plan, further upgrade is unavailable. **After your plan expires, you can repurchase a plan of any tier.** > Price difference = New package price -(Remaining balance of original package / Total amount of original package) * Original package price - **Auto-renewal supported**: The auto-renewal feature has been launched. You can enjoy a discount when you first sign up for a continuous subscription. Stay tuned for more updates. - **Refunds are not supported for the time being**: Please note that once a subscription service is purchased, it takes effect immediately and no refunds will be provided. Unused quotas within the package are also non-refundable. Please carefully select an appropriate subscription plan based on your own usage needs. - **Invoice support**: Domestic users can apply for invoices based on the transaction orders in their recharge details, and the actual invoiceable amount is the actual amount paid. Overseas users can directly download the invoice after purchase, or download it on the recharge details page. ## Package Usage {/* feishu-style:text-align:left */} The quota of the Token Plan package is only available for use in programming tools (such as OpenClaw, OpenCode, etc.), and it is prohibited to use it in the form of API calls for request behaviors in obvious non-coding scenarios such as automated scripts and custom application backends. {/* feishu-style:text-align:left */} If calls beyond the permitted scope are made using the API Key corresponding to the package, such act will be deemed as a violation or abuse, and the platform reserves the right to take measures including but not limited to suspending services for the relevant subscription and blocking the API Key. ## Quick Guide {/* feishu-style:text-align:left */} Get started quickly with Token Plan, from subscribing to a package to using the MiMo model in coding tools. ### Subscribe to the Token Plan {/* feishu-style:text-align:left */} Visit [Token Plan](https://platform.xiaomimimo.com/#/token-plan), then select and purchase a suitable subscription package as needed. ### Obtain the package-specific Base URL and API Key {/* feishu-style:text-align:left */} After successful subscription, you can go to the [Token Plan ](https://platform.xiaomimimo.com/#/console/plan-manage)page to obtain the package-specific Base URL and API Key. - **API Key**: On the [Token Plan ](https://platform.xiaomimimo.com/#/console/plan-manage)page, obtain your exclusive API Key (in the format of `tp-xxxxx`). - **Base URL**: You will need to configure one of the following Base URLs in your AI programming tool later (**The protocol varies by tool, and the Base URL shall be subject to what is displayed on the** [**Token Plan**](https://platform.xiaomimimo.com/#/console/plan-manage) **page**). For specific operations, please refer to the user guide document of the corresponding AI programming tool. - **OpenAI Compatible Protocol** - China Cluster: `https://token-plan-cn.xiaomimimo.com/v1` - Singapore Cluster: `https://token-plan-sgp.xiaomimimo.com/v1` - European Cluster: `https://token-plan-ams.xiaomimimo.com/v1` - **Anthropic Compatibility Protocol** - China Cluster: `https://token-plan-cn.xiaomimimo.com/anthropic` - Singapore Cluster: `https://token-plan-sgp.xiaomimimo.com/anthropic` - European Cluster: `https://token-plan-ams.xiaomimimo.com/anthropic`
**Precautions** - Please keep your API Key properly and do not disclose it to others. - The API Key is only valid during the validity period of the Token Plan package you have subscribed to.
## Application in AI Agents and Programming Tools {/* feishu-style:text-align:left */} The Token Plan is supported for use in multiple mainstream AI programming tools, and all tools share the usage quota of the subscribed package. {/* feishu-style:text-align:left */} Go to [AI Tools Overview ](https://platform.xiaomimimo.com/#/docs/integration/tools-overview)to view the configuration guide corresponding to the tools you are using (such as OpenCode, OpenClaw, etc.). ## Frequently Asked Questions {/* feishu-style:text-align:left */} For more frequently asked questions, please refer to [FAQ](https://mimo.mi.com/docs/zh-CN/quick-start/faq/token-plan/Plans&Pricing). ## Individual Subscription {/* feishu-style:text-align:left */} Token Plan 个人版是面向个人开发者的 AI 编程场景推出的专属订阅方案,您可使用高性价比的订阅资源包,在各类主流 AI 开发工具中调用 MiMo 旗舰大模型。 --- DOCUMENT: Team Subscription --- URL: https://mimo.mi.com/static/docs/tokenplan/Token Plan/team.md # Team Subscription {/* feishu-style:text-align:left */} **Token Plan Team Edition** is an AI programming subscription solution tailored for development teams and enterprises. Building on all the benefits of the individual edition, it provides capabilities including member management, seat allocation, team usage statistics, and centralized billing management. The team owner purchases packages and allocates seats in a unified manner, and members can call the MiMo model in mainstream AI programming and agent tools via their exclusive API Keys. ## Core Strengths - **First to integrate the latest flagship models and features:** All tiered plans support a total of 6 models, namely mimo-v2.6-pro, mimo-v2.6-flash, mimo-v2.5-asr, mimo-v2.5-tts-voiceclone, mimo-v2.5-tts-voicedesign and mimo-v2.5-tts. - **Unified Space Seat Management:** Centralize the management of team members, and allocate, reclaim and purchase additional seats as needed to flexibly adapt to team changes. - **Team usage is transparent and controllable:** Independent metering based on seats, with support for viewing usage by model and member, so as to clearly grasp the consumption of team quota. - **Centralized Billing and Invoice Management:** View records of plan purchase, renewal and seat add-on in a unified manner, apply for invoices centrally to streamline financial management. - **Flexible Subscription Plans:** Three tiered packages including Standard, Pro and Max, with support for monthly auto-renewal, annual auto-renewal and one-time purchase, to meet the needs of teams of different sizes and usage intensities. - **All benefits of the Personal Edition are inherited:** Multi-ecosystem Out Of The Box, enjoy off-peak discounts during nighttime, with no 5-hour cap or weekly usage limit. ## Quota Limit #### Monthly Plan
**Standard** **Pro** **Max**
**Pricing** $16/seat/month, ¥99/seat/month USD 50/seat/month, CNY 329/seat/month USD 100/seat/month, CNY 659/seat/month
**Fixed monthly quota per seat** 11 billion Credits 38 billion Credits 82 billion Credits
#### Annual Package
**Standard** **Pro** **Max**
**Pricing** USD 168.96/seat/year, CNY 1,044/seat/year USD 528/seat/year, CNY 3,468/seat/year USD 1,056/seat/year, CNY 6,948/seat/year
**Fixed monthly quota per seat** 11 billion Credits 38 billion Credits 82 billion Credits
> All subscription tiers **require a minimum purchase of 2 seats**. Annual plans are paid on a yearly basis, with seat quotas reset monthly rather than being allocated in full at the start of the year. #### Applicable Scenarios
**Standard** **Pro** **Max**
**Applicable Scenarios** Team members suitable for using AI to improve efficiency in daily work
Using mimo-v2.6-flash as the baseline, each seat can execute approximately **1600 rounds of medium-to-complex tasks**
Suitable for developers and professionals who use AI frequently on a daily basis
Using mimo-v2.6-flash as the baseline, each seat can execute approximately **5600 rounds of medium-to-complex tasks**
Hardcore users who are suitable to use AI as a core productivity tool
Using mimo-v2.6-flash as the baseline, each seat can execute approximately **12,800 rounds of medium-to-complex tasks**
> The above is the scope of scenarios corresponding to the monthly quota per seat. The annual cumulative task processing volume of the annual package is approximately 12 times that of the monthly quota, and the quota is still reset on a monthly basis. #### Quota Consumption Rules {/* feishu-style:text-align:left */} Language models deduct Credits quota based on the number of tokens; ASR models deduct Credits quota according to the duration of the input audio (the duration is counted to the exact second and converted to hourly billing); TTS series models are free for a limited time and do not consume package Credits. The specific conversion rules are as follows. {/* feishu-style:text-align:left */} **Language Model**
model Input (Cache Hit) Token Input (cache miss) Token Output Token
mimo-v2.6-pro 2.5 Credits 300 Credits 600 Credits
mimo-v2.6-flash 2 Credits 100 Credits 200 Credits
{/* feishu-style:text-align:left */} **ASR Model**
model Input audio duration (h)
mimo-v2.5-asr 30M Credits
{/* feishu-style:text-align:left */} **TTS model** {/* feishu-style:text-align:left */} mimo-v2.5-tts-voiceclone, mimo-v2.5-tts-voicedesign and mimo-v2.5-tts are available for free for a limited time, and will not consume the credits in your package. > The Team version is metered independently per seat. When an allocated and consumed seat is reassigned, the new member inherits the remaining quota of that seat for the current billing cycle, and the already consumed portion will not be reset. - **Quota Exhaustion:** When the monthly quota of a seat is used up, that seat will be suspended, without affecting other seats that still have remaining quota. **A team-shared usage package is coming soon, stay tuned.** - **To continue using the service:** you can contact the team owner to purchase additional seats as needed and reassign them, or switch to the regular pay-as-you-go API service to keep using it. ## Package Purchase - **Discounts: Enjoy an approximate 12% discount for annual auto-renewal plans or one-time purchases valid for 1 year or longer, plus 20% lower resource consumption during nighttime hours** - Annual Discount: Annual recurring subscription enjoys an approximately 12% discount off the monthly rate, and one-time purchases covering 12 months are also eligible for the annual package price. No first-purchase discount is available for the Team Plan. - Off-peak discount rate: A consumption coefficient of 0.8x applies during off-peak hours (00:00–08:00 Beijing Time, 16:00–24:00 UTC). - **Flexible Purchase Options:** We support monthly auto-renewal, annual auto-renewal, and one-time purchases, with a minimum order of 2 seats. For one-time purchases, you can choose a validity period of 1 to 12 months to meet the usage needs of teams with different cycle requirements. - **Support for adding seats mid-subscription:** The owner may add seats of the same tier during the validity period of the current plan, with a minimum of 1 seat per addition. The fee and current quota for the newly added seats will be prorated based on the remaining validity period, and their expiration date will align with that of the current plan. - **Auto-renewal supported:** Monthly and annual subscription plans are auto-renewable, which can be turned off by the subscriber at any time. - **Invoice Support:** Domestic users can apply for invoices based on the transaction orders in their recharge details, and the actual invoiceable amount is the actual payment amount. Overseas users can directly download the invoice after purchase, or download it on the recharge details page.
**Precautions** - Multiple members are not supported to share the same seat - Upgrades and mixed purchase of package tiers are temporarily not supported during the validity period. If you need additional usage resources, you can purchase extra seats for the current tier as needed. - Unsubscription and refund for individual or partial seats are not supported. No refund will be provided for unused quotas within the package. Please carefully select an appropriate subscription plan and number of seats based on your own usage needs.
## Package Usage {/* feishu-style:text-align:left */} The quota of the Token Plan package is only available for use in programming tools (such as OpenClaw, OpenCode, etc.), and it is prohibited to use it in the form of API calls for request behaviors in obvious non-coding scenarios such as automated scripts and custom application backends. {/* feishu-style:text-align:left */} If calls beyond the permitted scope are made using the API Key corresponding to the package, such act will be deemed as a violation or abuse, and the platform reserves the right to take measures including but not limited to suspending services for the relevant subscription and blocking the API Key. ## Quick Guide {/* feishu-style:text-align:left */} Get started quickly with Token Plan, from subscribing to a package to using the MiMo model in coding tools. ### Subscribe to the Token Plan {/* feishu-style:text-align:left */} Visit [Token Plan](https://platform.xiaomimimo.com/#/token-plan), and select and purchase the appropriate team edition subscription package as needed. ### Invite members and assign seats {/* feishu-style:text-align:left */} After successful subscription, you can invite members in the Team Edition TokenPlan Console, and assign available seats before you can use the package. Of course, you can also assign a seat to yourself. 图片 ### Obtain the package-specific Base URL and API Key {/* feishu-style:text-align:left */} After members successfully access the team workspace and obtain the seats assigned by the owner, they can acquire the package-specific Base URL and API Key within the team workspace (the owner should enter "My Seats"). 图片 - **API Key**: Obtain a dedicated API Key (in the format of `ttp-xxxxx`) within the team workspace. - **Base URL**: You will need to configure one of the following Base URLs in the AI programming tool later (**The protocol varies by tool, and the Base URL shall be subject to what is displayed on the Team Edition TokenPlan page**). For specific operations, please refer to the corresponding user guide document for the AI programming tool. - **OpenAI Compatible Protocol** - China Cluster: `https://token-plan-cn.xiaomimimo.com/v1` - Singapore Cluster: `https://token-plan-sgp.xiaomimimo.com/v1` - European Cluster: `https://token-plan-ams.xiaomimimo.com/v1` - **Anthropic Compatibility Protocol** - China Cluster: `https://token-plan-cn.xiaomimimo.com/anthropic` - Singapore Cluster: `https://token-plan-sgp.xiaomimimo.com/anthropic` - European Cluster: `https://token-plan-ams.xiaomimimo.com/anthropic`
**Precautions** - Please keep your API Key properly and do not disclose it to others. - The API Key is only valid during the validity period of the Token Plan package you have subscribed to.
## Application in AI Agents and Programming Tools {/* feishu-style:text-align:left */} The Token Plan is supported for use in multiple mainstream AI programming tools, and all tools share the usage quota of the subscribed package. {/* feishu-style:text-align:left */} Go to [AI Tools Overview ](https://platform.xiaomimimo.com/#/docs/integration/tools-overview)to view the configuration guide corresponding to the tools you are using (such as OpenCode, OpenClaw, etc.). ## Frequently Asked Questions {/* feishu-style:text-align:left */} For more frequently asked questions, please refer to [FAQ](https://mimo.mi.com/docs/zh-CN/quick-start/faq/token-plan/Team-Benefits). --- DOCUMENT: Quick Access --- URL: https://mimo.mi.com/static/docs/tokenplan/Token Plan/quick-access.md # Quick Access {/* feishu-style:text-align:left */} This article describes how to quickly integrate with Token Plan. The entire process from subscription to invocation can be completed in just 3 steps. ## Step 1: Subscribe to the Token Plan {/* feishu-style:text-align:left */} Go to [Token Plan ](https://platform.xiaomimimo.com/#/token-plan)and select a suitable subscription plan. ## Step 2: Obtain the voucher {/* feishu-style:text-align:left */} After successful subscription, please go to [Token Plan ](https://platform.xiaomimimo.com/#/console/plan-manage)to obtain the following credentials:
**Note: For Team Edition, please first complete** [**member invitation and seat allocation**](https://platform.xiaomimimo.com/console/invite-members)
- **API Key**: On the [Token Plan ](https://platform.xiaomimimo.com/#/console/plan-manage)page, obtain your exclusive API Key (in the format of `tp-xxxxx `or `ttp-xxxxx`). - **Base URL**: You will need to configure one of the following Base URLs in your AI programming tool later (**The protocol varies by tool, and the Base URL shall be subject to what is displayed on the** [**Token Plan**](https://platform.xiaomimimo.com/#/console/plan-manage) **page**). For specific operations, please refer to the user guide document of the corresponding AI programming tool. - **OpenAI Compatible Protocol** - China Cluster: `https://token-plan-cn.xiaomimimo.com/v1` - Singapore Cluster: `https://token-plan-sgp.xiaomimimo.com/v1` - European Cluster: `https://token-plan-ams.xiaomimimo.com/v1` - **Anthropic Compatibility Protocol** - China Cluster: `https://token-plan-cn.xiaomimimo.com/anthropic` - Singapore Cluster: `https://token-plan-sgp.xiaomimimo.com/anthropic` - European Cluster: `https://token-plan-ams.xiaomimimo.com/anthropic`
The API Key for the Token Plan (`tp-xxxxx `or `ttp-xxxxx`) and the API Key for pay-as-you-go API calls (`sk-xxxxx`) are independent of each other and cannot be used interchangeably.
## Step 3: Integrate with AI programming tools {/* feishu-style:text-align:left */} Go to [AI Tools Overview ](https://platform.xiaomimimo.com/#/docs/integration/tools-overview)to view the configuration guide corresponding to the tools you are using (such as OpenCode, OpenClaw, etc.). ## Rapid Verification (Optional) {/* feishu-style:text-align:left */} Once the configuration is complete, you can quickly verify whether the integration is successful via the following methods. ### Method 1: Verification via AI programming tools {/* feishu-style:text-align:left */} Enter a simple programming requirement into the configured AI programming tool, for example: > Please write a quick sort algorithm for me in Python {/* feishu-style:text-align:left */} If the tool returns a normal code, it indicates that the integration is successful. ### Method 2: Direct verification via API call {/* feishu-style:text-align:left */} Use the curl command to directly call the API and verify whether the credentials are valid.
In the following examples, `BASE_URL `and `MIMO_API_KEY `are both placeholders. Please replace them with the actual credentials obtained from the Console when in use.
#### OpenAI Compatible Protocol ```bash curl --location --request POST 'BASE_URL/chat/completions' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "messages": [ { "role": "system", "content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024." }, { "role": "user", "content": "please introduce yourself" } ], "max_completion_tokens": 1024 }' ``` ```bash curl --location --request POST 'BASE_URL/responses' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "instructions": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.", "input": "please introduce yourself", "max_output_tokens": 1024, "stream": false, "reasoning": { "effort": "high" } }' ``` #### Anthropic Compatibility Protocol ```bash curl --location --request POST 'BASE_URL/v1/messages' \ --header "api-key: $MIMO_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "mimo-v2.6-pro", "max_tokens": 1024, "system": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "please introduce yourself" } ] } ] }' ``` ## Frequently Asked Questions {/* feishu-style:text-align:left */} **Question: What is the difference between the API Key for Token Plan and the API Key for pay-as-you-go API calls?** {/* feishu-style:text-align:left */} Answer: The API Key format for Token Plan is `tp-xxxxx `or `ttp-xxxxx`, which is only applicable to the Token Plan subscription service; the API Key format for pay-as-you-go API calls is `sk-xxxxx`, which is used for pay-as-you-go billing. The two are independent of each other and cannot be used interchangeably. {/* feishu-style:text-align:left */} **Question: What is the difference between the Base URL of Token Plan and the Base URL of pay-as-you-go API calls?** {/* feishu-style:text-align:left */} Answer: The Base URL format of the Token Plan varies, subject to [Token Plan ](https://platform.xiaomimimo.com/#/console/plan-manage)page display. {/* feishu-style:text-align:left */} **Question: Can the API Key still be used after the subscription expires?** {/* feishu-style:text-align:left */} Answer: No. The API Key of the Token Plan is only valid during the subscription period, and you need to renew the subscription to continue using it after it expires. --- DOCUMENT: Overview of AI Tools --- URL: https://mimo.mi.com/static/docs/tokenplan/integration/tools-overview.md # Overview of AI Tools {/* feishu-style:text-align:left */} **The pay-as-you-go MiMo API and Token Plan subscription packages,** are both supported for use in the following mainstream AI programming tools (the tool list is continuously updated), click to view the detailed access and usage guide for the corresponding tool. ## Use Tools {/* feishu-style:text-align:left */} ## Configuration Methods for Other Tools > Core Steps: > > 1. Find a compatible OpenAI protocol or Anthropic protocol, and support custom configuration Provider > > 1. Replace or add as the Base URL for the corresponding protocol > > 1. Enter API Key, select or add MiMo model {/* feishu-style:text-align:left */} Supports two usage methods, but the corresponding credential acquisition methods are different:
Usage Method Description Acquisition Method (BASE_URL and API Key below are examples)
Pay-as-you-go MiMo API Charged based on actual usage, suitable for light use
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://api.xiaomimimo.com/v1`
    • Anthropic Compatibility Protocol: `https://api.xiaomimimo.com/anthropic`
  • API Key
    • Format: `sk-xxxxx`

Go to [API Keys](https://platform.xiaomimimo.com/#/console/api-keys) to create an API Key
Token Plan Fixed subscription fee, with limited calls based on the package
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/v1`
    • Anthropic Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/anthropic`
  • API Key
    • Format (Individual): `tp-xxxxx`
    • Format (Team): `ttp-xxxxx`

After successful subscription, go to [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) to obtain the exclusive Base URL and API Key
--- DOCUMENT: MiMo Desktop Configuration --- URL: https://mimo.mi.com/static/docs/tokenplan/integration/mimo-desktop.md # MiMo Desktop Configuration {/* feishu-style:text-align:left */} **Pay-as-you-go MiMo API** and **Token Plan** both support MiMo Desktop. Refer to this guide for configuration and usage. ## Prerequisites ### Obtain Credentials {/* feishu-style:text-align:left */} Supports two usage methods, but the corresponding credential acquisition methods are different:
Usage Method Description Acquisition Method (BASE_URL and API Key below are examples)
Pay-as-you-go MiMo API Charged based on actual usage, suitable for light use
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://api.xiaomimimo.com/v1`
  • API Key
    • Format: `sk-xxxxx`

Go to [API Keys](https://platform.xiaomimimo.com/#/console/api-keys) to create an API Key
Token Plan Fixed subscription fee, with limited calls based on the package
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/v1`
  • API Key
    • Format (Individual): `tp-xxxxx`
    • Format (Team): `ttp-xxxxx`

After successful subscription, go to [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) to obtain the exclusive Base URL and API Key
## Install MiMo Desktop {/* feishu-style:text-align:left */} MiMo Desktop currently supports macOS (ARM architecture) and Windows operating systems. > The following instructions use the China mainland version of MiMo Desktop as an example. International users should refer to the corresponding version for your region. {/* feishu-style:text-align:left */} Apply here: - International users: https://mimo.xiaomimimo.com/desktop/ - Users in China: https://mimo.xiaomimimo.com/desktop/ ## Log in to MiMo Desktop {/* feishu-style:text-align:left */} Currently, only personal account login is available. You need to log in with a Xiaomi account. If you already have a Xiaomi account, you can log in directly. If you don't have one, you can register via the [Console](https://platform.xiaomimimo.com/#/console/usage) or at [id.mi.com](https://id.mi.com/) beforehand. ## Use MiMo Desktop Preset Models {/* feishu-style:text-align:left */} MiMo Desktop currently has three preset models: `mimo-v2.6-flash`, `mimo-v2.6-pro`, and `mimo-v2.6-pro-ultraspeed`.
If you're unsure which model to choose, select **Smart** mode. MiMo Desktop will automatically evaluate the task type, complexity, delivery requirements, and execution cost, then match the most suitable execution path for the current task.
图片 ## Configure Custom Models {/* feishu-style:text-align:left */} MiMo Desktop supports custom models. Using Xiaomi MiMo API as an example, MiMo Desktop has this provider pre-configured — simply enter your API Key and model name (e.g., `mimo-v2.6-pro`) to get started.
**Note** The API Key for **Xiaomi MiMo API** (`sk-xxxxx`) and the API Key for **MiMo Token Plan** (`tp-xxxxx` or `ttp-xxxxx`) are not interchangeable. Please configure the correct one based on the service you are using.
图片 --- DOCUMENT: MiMo Code Configuration --- URL: https://mimo.mi.com/static/docs/tokenplan/integration/mimo-code.md # MiMo Code Configuration {/* feishu-style:text-align:left */} [MiMo Code](https://mimo.xiaomi.com/zh/mimocode) is an AI-powered coding assistant developed by Xiaomi, available as a CLI tool in the terminal. Both **pay-as-you-go MiMo API** and **Token Plan** are supported by MiMo Code. Refer to this guide for configuration and usage.
**Limited-Time Offer** After completing authorization for MiMo Code, each user is entitled to **1000 free** web search queries per day. Once the quota is exceeded, please activate the relevant service on the open platform and ensure your account has sufficient balance.
## Prerequisites {/* feishu-style:text-align:left */} MiMo Code supports direct redirection to the [Xiaomi MiMo API Platform](https://platform.xiaomimimo.com/docs/en-US/welcome) for authorization login, or you can log in using the authorization code returned by the platform. No manual API Key configuration is required. The system provides two key types for you to choose from based on your usage scenario. > Note: Before authorizing and logging into MiMo Code, please ensure your account on the open platform has sufficient balance or a valid Token Plan. Otherwise, the authorization will fail.
Usage Description How to Obtain and Manage
Pay-as-you-go API Billed by actual usage, suitable for light usage After authorization, the platform will automatically create a new API Key prefixed with `mimo-code-cli-key`. You can view and manage it on the [API Keys](https://platform.xiaomimimo.com/#/console/api-keys) page.
Token Plan Fixed subscription fee with limited calls per plan You can view the current Token Plan quota, usage, and expiration on the [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) page.
## Install MiMo Code {/* feishu-style:text-align:left */} MiMo Code supports two installation methods. {/* feishu-style:text-align:left */} **Method 1: Official Script Installation (for macOS/Linux)** > For a better user experience, Mac users are strongly recommended to use iTerm or VSCode Terminal. ```bash curl -fsSL https://mimo.xiaomi.com/install | bash ``` {/* feishu-style:text-align:left */} **Method 2: npm Installation (for Windows)** {/* feishu-style:text-align:left */} Requires Node.js 18 or later. ```bash npm install -g @mimo-ai/cli ``` {/* feishu-style:text-align:left */} **Verify Installation (a version number output indicates successful installation):** ```bash mimo --version ``` ## Connect Provider {/* feishu-style:text-align:left */} You can connect to the Xiaomi MiMo provider in two ways: {/* feishu-style:text-align:left */} **1. MiMo Code Already Running** {/* feishu-style:text-align:left */} Run the `/connect` or `/login` command in the interactive interface, and select `Xiaomi` as the provider. {/* feishu-style:text-align:left */} **2. MiMo Code Not Yet Running** {/* feishu-style:text-align:left */} Run the following command directly in the terminal and select `MiMo` to complete authorization. ```bash mimo auth login ``` {/* feishu-style:text-align:left */} After confirmation, the MiMo authorization login popup will automatically appear. Follow the prompts to complete the login, and then select your preferred authorization key type based on your usage scenario: 图片 ## Use MiMo Code ### Quick Start {/* feishu-style:text-align:left */} Follow these steps to use MiMo Code in your project: ```bash # 1. Navigate to your project directory cd /path/to/your/project # 2. Launch MiMo Code mimo # 3. (Recommended) Initialize project configuration on first use /init ```
It is strongly recommended to run the `/init` command on first use: - It automatically analyzes your project structure and coding conventions - Generates an `AGENTS.md` file in the project root directory - MiMo Code will use this file to better understand your project context, improving interaction quality
{/* feishu-style:text-align:left */} For more commands and detailed usage, please refer to the [MiMo Code Official Documentation](https://mimo.xiaomi.com/mimocode/interaction). 图片 ### Model Selection {/* feishu-style:text-align:left */} Run the `/models` command to view and select from the currently available models. ## FAQ ### What should I do if I encounter the following error when verifying the installation on Windows? > It seems that your package manager failed to install the right version of the mimocode CLI for your platform. You can try manually installing "@mimo-ai/mimocode-windows-x64" or "@mimo-ai/mimocode-windows-x64-baseline" package {/* feishu-style:text-align:left */} Answer: Run the command `npm install -g @mimo-ai/mimocode-windows-x64` as indicated to resolve the issue. ### Why can't I see the model's reasoning content? {/* feishu-style:text-align:left */} Answer: MiMo Code does not display the model's reasoning content by default. You can use the `/thinking` command to toggle the visibility of reasoning blocks in the conversation. Once enabled, you can view the complete reasoning process of models that support extended thinking. > Note: This command is not a toggle for the model's thinking function. It cannot enable or disable the model's thinking process. ### Why does authorization fail? {/* feishu-style:text-align:left */} Answer: Please check if your account on the open platform has sufficient balance or a valid Token Plan. --- DOCUMENT: OpenCode Configuration --- URL: https://mimo.mi.com/static/docs/tokenplan/integration/opencode.md # OpenCode Configuration {/* feishu-style:text-align:left */} **Pay-as-you-go MiMo API** and **Token Plan** both support OpenCode. Refer to this guide for configuration and usage. ## Prerequisites ### Obtain Credentials {/* feishu-style:text-align:left */} Supports two usage methods, but the corresponding credential acquisition methods are different:
Usage Method Description Acquisition Method (BASE_URL and API Key below are examples)
Pay-as-you-go MiMo API Charged based on actual usage, suitable for light use
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://api.xiaomimimo.com/v1`
  • API Key
    • Format: `sk-xxxxx`

Go to [API Keys](https://platform.xiaomimimo.com/#/console/api-keys) to create an API Key
Token Plan Fixed subscription fee, with limited calls based on the package
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/v1`
  • API Key
    • Format (Individual): `tp-xxxxx`
    • Format (Team): `ttp-xxxxx`

After successful subscription, go to [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) to obtain the exclusive Base URL and API Key
When OpenCode uses MiMo under the Anthropic protocol, since the assistant containing tool calls is missing `reasoning_content`, the API will return a 400 error. For details, see [Multi-turn Conversation Pass-through Requirements](https://mimo.mi.com/docs/en-US/quick-start/usage-guide/text-generation/deep-thinking#:~:text=.-,Multi%2Dturn%20Conversation%20Pass%2Dthrough%20Requirements,-When%20deep%20thinking).
## Use OpenCode CLI ### Install OpenCode CLI {/* feishu-style:text-align:left */} OpenCode supports two installation methods. {/* feishu-style:text-align:left */} **Method 1: Official Script Installation (for macOS/Linux)** ```bash curl -fsSL https://opencode.ai/install | bash ``` {/* feishu-style:text-align:left */} **Method 2: npm Installation** {/* feishu-style:text-align:left */} Node.js 18 or later is required. ```bash npm install -g opencode-ai ``` {/* feishu-style:text-align:left */} **Verify installation (if a version number is displayed, the installation was successful):** ```bash opencode -v ``` ### Configure Basic Settings {/* feishu-style:text-align:left */} Edit or create the `opencode.json` configuration file at the following path: - **macOS/Linux**: `~/.config/opencode/opencode.json` - **Windows**: `User Directory\.config\opencode\opencode.json` {/* feishu-style:text-align:left */} Copy the following content into the configuration file (replace `BASE_URL` and `MIMO_API_KEY` as needed): ```json { "$schema": "https://opencode.ai/config.json", "provider": { "mimo": { "npm": "@ai-sdk/openai-compatible", "name": "MiMo", "options": { "baseURL": "BASE_URL", "apiKey": "MIMO_API_KEY" }, "models": { "mimo-v2.6-pro": { "name": "mimo-v2.6-pro", "limit": { "context": 1048576, "output": 131072 }, "modalities": { "input": [ "text", "image" ], "output": [ "text" ] } }, "mimo-v2.6-flash": { "name": "mimo-v2.6-flash", "limit": { "context": 1048576, "output": 131072 }, "modalities": { "input": [ "text", "image" ], "output": [ "text" ] } } } } } } ```
If you need to enable the image understanding capability, you need to modify or add the following configuration items under the configuration node of the model that supports this capability (e.g., `mimo-v2.6-pro`). That is, add `image` to the supported input modalities: `"modalities": {"input": ["text", "image"], "output": ["text"]}`
### Use OpenCode CLI {/* feishu-style:text-align:left */} After completing the configuration, navigate to the project directory and run the following command to start OpenCode: ```bash opencode ``` {/* feishu-style:text-align:left */} After starting, enter `/models` to view and switch between available models. ## Use OpenCode IDE Plugin ### Install Plugin {/* feishu-style:text-align:left */} Search for and install the **opencode** plugin in the VS Code Extensions marketplace. 图片 ### Configure a Predefined Provider (Recommended) {/* feishu-style:text-align:left */} Just enter `/connect` in the input box, search for `Xiaomi`, select the corresponding Provider, and fill in the API Key.
When using the **Xiaomi Token Plan**, you need to select the Provider corresponding to the Base URL displayed on the [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) page. - `https://token-plan-cn.xiaomimimo.com/*`: Xiaomi Token Plan (China) - `https://token-plan-sgp.xiaomimimo.com/*`: Xiaomi Token Plan (Singapore) - `https://token-plan-ams.xiaomimimo.com/*`: Xiaomi Token Plan (Europe)
图片 ### Configure a Custom Provider {/* feishu-style:text-align:left */} Refer to the "Configure Basic Settings" steps in the OpenCode CLI section above. ### Use OpenCode Plugin 图片 ## FAQ ### When verifying the installation on Windows, I encounter the following error. How to fix it? > It seems that your package manager failed to install the right version of the opencode CLI for your platform. You can try manually installing "opencode-windows-x64" or "opencode-windows-x64-baseline" package {/* feishu-style:text-align:left */} Run the command `npm install -g opencode-windows-x64` as prompted to resolve the issue. ### Error when starting OpenCode in VS Code on Windows? > opencode : Cannot load file ... because running scripts is disabled on this system {/* feishu-style:text-align:left */} Change the default terminal type to Git Bash when opening a terminal in VS Code. 图片 --- DOCUMENT: Claude Code Configuration --- URL: https://mimo.mi.com/static/docs/tokenplan/integration/claudecode.md # Claude Code Configuration {/* feishu-style:text-align:left */} **Pay-as-you-go MiMo API** and **Token Plan** both support Claude Code. Refer to this guide for configuration and usage. ## Prerequisites ### Obtain Credentials {/* feishu-style:text-align:left */} Supports two usage methods, but the corresponding credential acquisition methods are different:
Usage Method Description Acquisition Method (BASE_URL and API Key below are examples)
Pay-as-you-go MiMo API Charged based on actual usage, suitable for light use
  • BASE_URL
    • Anthropic Compatibility Protocol: `https://api.xiaomimimo.com/anthropic`
  • API Key
    • Format: `sk-xxxxx`

Go to [API Keys](https://platform.xiaomimimo.com/#/console/api-keys) to create an API Key
Token Plan Fixed subscription fee, with limited calls based on the package
  • BASE_URL
    • Anthropic Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/anthropic`
  • API Key
    • Format (Individual): `tp-xxxxx`
    • Format (Team): `ttp-xxxxx`

After successful subscription, go to [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) to obtain the exclusive Base URL and API Key
## Use Claude Code CLI ### Install Claude Code CLI {/* feishu-style:text-align:left */} Claude Code requires Node.js 18 or later. - Linux/macOS: No additional setup needed, the default environment is sufficient. - Windows: Install [WSL](https://learn.microsoft.com/en-us/windows/wsl/install) or [Git for Windows](https://git-scm.com/install/windows), then run the command below in WSL or Git Bash. {/* feishu-style:text-align:left */} **Installation command:** ```bash npm install -g @anthropic-ai/claude-code ``` {/* feishu-style:text-align:left */} **Verify the installation (a version number output indicates success):** ```bash claude --version ``` ### Configure Basic Settings
Before configuring, make sure to clear the following Anthropic official environment variables to avoid API conflicts: `ANTHROPIC_AUTH_TOKEN`, `ANTHROPIC_BASE_URL`
{/* feishu-style:text-align:left */} **1.** **Create/edit** `settings.json` > If the `.claude` directory does not exist, you can create it manually. - macOS/Linux: `~/.claude/settings.json` - Windows: `User directory\.claude\settings.json` {/* feishu-style:text-align:left */} Please replace `BASE_URL` (Anthropic Compatibility Protocol) and `MIMO_API_KEY` as needed.
For MiMo models that support **1M** context, you can append the `[1m]` suffix to the model ID to enable extended context capacity. Example: `mimo-v2.6-pro[1m]`. After configuration, restart Claude Code and run the `/context` command to verify whether the long context takes effect.
```json { "env": { "ANTHROPIC_BASE_URL": "BASE_URL", "ANTHROPIC_AUTH_TOKEN": "MIMO_API_KEY", "ANTHROPIC_MODEL": "mimo-v2.6-pro", "ANTHROPIC_DEFAULT_SONNET_MODEL": "mimo-v2.6-pro", "ANTHROPIC_DEFAULT_OPUS_MODEL": "mimo-v2.6-pro", "ANTHROPIC_DEFAULT_HAIKU_MODEL": "mimo-v2.6-pro" } } ``` {/* feishu-style:text-align:left */} **2.** **Create/edit** `.claude.json` - macOS/Linux: `~/.claude.json` - Windows: `User directory\.claude.json` ```json { "hasCompletedOnboarding": true } ``` {/* feishu-style:text-align:left */} **3.** **Apply the configuration** {/* feishu-style:text-align:left */} After completing the configuration, **reopen the terminal window** for the changes to take effect. ### Use Claude Code CLI {/* feishu-style:text-align:left */} Navigate to your project directory and run: ```bash claude ``` {/* feishu-style:text-align:left */} On first launch, complete the following: select "**Trust This Folder**" to allow Claude Code to access project files. After startup, use the `/status` command to verify the current configuration and model status. ## Use the Claude Code IDE Plugin {/* feishu-style:text-align:left */} Claude Code provides a VS Code IDE plugin. For configuration reference, see the official documentation [Use Claude Code in VS Code](https://code.claude.com/docs/en/vs-code#vs-code-extension-vs-claude-code-cli). ### Install Plugin {/* feishu-style:text-align:left */} Search for and install the **Claude Code for VS Code** plugin from the VS Code Extensions marketplace. 图片 ### Configure the Model {/* feishu-style:text-align:left */} Open VS Code settings, search for `Claude Code: Environment Variables`, and then manually configure it in `settings.json`: ```json { "claudeCode.preferredLocation": "panel", "claudeCode.selectedModel": "mimo-v2.6-pro", "claudeCode.environmentVariables": [ { "name": "ANTHROPIC_BASE_URL", "value": "BASE_URL" }, { "name": "ANTHROPIC_AUTH_TOKEN", "value": "MIMO_API_KEY" }, { "name": "ANTHROPIC_DEFAULT_SONNET_MODEL", "value": "mimo-v2.6-pro" }, { "name": "ANTHROPIC_DEFAULT_OPUS_MODEL", "value": "mimo-v2.6-pro" }, { "name": "ANTHROPIC_DEFAULT_HAIKU_MODEL", "value": "mimo-v2.6-pro" } ] } ```
If Claude Code CLI is already installed, the VS Code plugin will automatically reuse the CLI configuration. To configure independently, specify the environment variables in the plugin settings as shown above.
## FAQ ### Installation fails on Windows? {/* feishu-style:text-align:left */} Ensure the following dependencies are installed: - Node.js 18+ - Git for Windows {/* feishu-style:text-align:left */} If you encounter permission issues with npm, try running the terminal as administrator, or use nvm to manage Node.js versions. --- DOCUMENT: Codex Configuration --- URL: https://mimo.mi.com/static/docs/tokenplan/integration/codex-configuration.md # Codex Configuration {/* feishu-style:text-align:left */} **Pay-as-you-go MiMo API** and **Token Plan** both support Codex. Refer to this guide for configuration and usage. ## Prerequisites ### Obtain Credentials {/* feishu-style:text-align:left */} Two usage methods are supported, but the corresponding credential acquisition methods differ:
Usage Method Description Acquisition Method (BASE_URL and API Key below are examples)
Pay-as-you-go API Charged based on actual usage, suitable for light use
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://api.xiaomimimo.com/v1`
  • API Key
    • Format: `sk-xxxxx`

Go to [API Keys](https://platform.xiaomimimo.com/#/console/api-keys) to create an API Key
Token Plan Fixed subscription fee, with limited calls based on the package
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/v1`
  • API Key
    • Format (Individual): `tp-xxxxx`
    • Format (Team): `ttp-xxxxx`

After successful subscription, go to [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) to obtain the exclusive Base URL and API Key
## Install Codex {/* feishu-style:text-align:left */} **Prerequisites:** Node.js 18 or a later version must be installed first. {/* feishu-style:text-align:left */} **Installation command:** ```bash npm install -g @openai/codex ``` {/* feishu-style:text-align:left */} **Verify the installation (a version number output indicates success):** ```bash codex --version ``` ## Edit Configuration File
**Notes** - When configuring basic information, first check if the `MIMO_API_KEY` environment variable exists. If it does, please clear it or replace the value with the API Key obtained through the corresponding usage method.
{/* feishu-style:text-align:left */} Configuration file paths: - macOS/Linux: - `config.toml`: `~/.codex/config.toml` - `model-catalogs.json`:`~/.codex/model-catalogs.json` - Windows - `config.toml`: `User directory\.codex\config.toml` - `model-catalogs.json`: `User directory\.codex\model-catalogs.json` {/* feishu-style:text-align:left */} Below is a complete configuration example. The `model` field can be modified to other supported models as needed (e. g., `mimo-v2.6-pro`). For the `model_catalog_json` configuration, please refer to the path of the system configuration file you are using. The file download address is: [model-catalogs.json](https://example-files.cnbj1.mi-fds.com/example-files/configs/model-catalogs.json). ```json { "models": [ { "slug": "mimo-v2.6-pro", "display_name": "MiMo-V2.6-Pro", "description": "Xiaomi MiMo: MiMo-V2.6-Pro", "default_reasoning_level": "low", "supports_experimental_context": true, "support_verbosity": false, "apply_patch_tool_type": "freeform", "input_modalities": ["text", "image"], "supports_image_detail_original": true, "truncation_policy": { "mode": "tokens", "limit": 10000 }, "supports_parallel_tool_calls": false, "tool_mode": "code_mode_only", "multi_agent_version": "v2", "multi_agent_reasoning_effort": "high", "use_responses_lite": true, "include_skills_usage_instructions": false, "include_apps_usage_instructions": false, "include_plugin_usage_instructions": false, "node_repl_auto_review_required": true, "node_repl_disabled": false, "requires_sandboxed_review": false, "auto_review_model_override": null, "model_specialty": null, "context_window": 1048576, "max_context_window": 1048576, "auto_compact_token_limit": null, "comp_hash": "3000", "default_reasoning_summary": "none", "supports_reasoning_summaries": true, "supported_reasoning_levels": [ { "effort": "none", "description": "No extra reasoning for faster responses" }, { "effort": "low", "description": "Fast responses with lighter reasoning" }, { "effort": "medium", "description": "Balances speed and reasoning depth for everyday tasks" }, { "effort": "high", "description": "Greater reasoning depth for complex problems" } ], "shell_type": "unified_exec", "visibility": "list", "supported_in_api": true, "priority": 0, "base_instructions": "You are MiMo, an AI assistant developed by Xiaomi. Today's date: {date} {week}. Your knowledge cutoff date is December 2024.", "model_messages": { "instructions_template": "You are MiMo, an AI assistant developed by Xiaomi. Today's date: {date} {week}. Your knowledge cutoff date is December 2024.", "instructions_variables": null, "persistent_instructions": "## Overview\nYou are now in persistent mode for this session until explicitly disabled by a later developer message.\n\nIn persistent mode, your first order goal is still to fulfill the user's request, as in non-persistent mode. The key difference is that now you need be more persistent and proactive: anticipate, identify, and perform useful follow-up tasks beyond the immediate deliverables.\n\nBecause a `final` answer immediately ends the turn, use `functions.send_user_message_async` to deliver answers while useful work remains. Only send a `final` message after concluding that no follow-up or proactive work could be a useful continuation of any user request in the current turn. Work that requires waiting still counts as a useful continuation; having nothing to do immediately is not sufficient reason to end the turn.\n\n## Proactivity & Follow-up Work\nFor follow-up work, favor closing a known open loop, establishing an awaited result, or verifying that a change took effect over inventing unrelated work. Use past user instructions and your knowledge of the user to prioritize follow-ups. For example, if the user asks how an eval run is going and it is still running, report its current status and continue monitoring that evaluation until it reaches a terminal state, unless the user requested only a snapshot or specified another stopping condition. Another example, when the user asked you to write a PR, after the PR is submitted, useful followup could be checking CI/CD status, tracking merge eligibility etc.\n\nBefore starting a follow-up, identify its scope, the outcome you want to establish, the evidence needed, and a stopping condition justified by the original task or external process. You can use `clock.sleep` to wait for external events and conditions to change. Once started, treat the follow-up as active ongoing work across sleeps until the outcome is established, the user cancels or replaces it, it is no longer relevant, a relevant observation window ends, or progress requires user input or additional authorization. Bound a follow-up by its purpose, scope, and outcome, not an arbitrary number of checks. A pending, running, inconclusive, or unchanged result is not by itself completion. Never invent an early stopping point for monitoring the user explicitly asked to continue.\n\nYou may perform safe, non-mutating follow-ups that remain within the user's authorized scope. Persistence does not broaden that scope. For follow-ups or next actions that require new authority, materially expand scope, or make external state changes not already authorized, describe the proposed action and obtain approval before executing it.\n\nWhen the user asks you to finish, monitor, or track, take end-to-end ownership of the specified task until the user's completion or stopping condition is reached. Autonomously perform authorized steps within scope, including checking progress, diagnosing problems, safely retrying, and fixing recoverable failures. Do not stop at an intermediate result, unchanged state, or recoverable failure. If completion requires action outside your authorization, pause the dependent work and ask the user for the specific authorization needed.\n\nPrefer working in the current task with `clock.sleep` between checks over automations. Only create automations when the task clearly require recurring work on a fixed schedule, such as checking Slack every five minutes or refreshing data every day. Do not create an automation merely to finish or monitor an operation already in progress.\n\n## Communication Guidelines\nUse `functions.send_user_message_async` to ask the user for missing information, a preference, a constraint, or clarification, and to directly answer user questions while work is still in progress.\n\nAsk clarification questions early unless their answers can potentially be inferred from the available context. Continue useful work that does not depend on the answer while waiting. For optional clarification, give the user a reasonable opportunity to reply—for example, 30 seconds for a simple question and longer for a complex one—before proceeding with a stated assumption. If an answer or approval is required, keep the question pending and do not proceed with dependent work until it arrives. Elapsed time is not an answer or approval.\n\nAvoid duplicate user-visible messages within a turn or across turns. For a simple greeting, thanks, or acknowledgment, one brief response or reaction is enough; do not send equivalent text through both `functions.send_user_message_async` and `final`. Keep substantive final answers self-contained, but do not send an extra message that merely repeats an answer, question, blocker, or approval request already communicated. Repeat one only when the user asks again, new information materially changes it, or a requested reminder or reply is due. Keep unanswered required questions pending; continue useful authorized work that does not depend on the answer, or wait quietly.\n\nMake updates feel like a natural continuation of the conversation. Lead with the useful finding, result, or decision; avoid announcing a \"follow-up task,\" declaring \"the follow-up is complete,\" narrating internal task bookkeeping, or adding unnecessary disclaimers about actions you are not taking.\n\nWhen using `functions.send_user_message_async` to deliver a substantive answer to the user's request, follow the formatting guidelines for a `final` answer.\n\n## Misc\nCall `update_up_next` before sleep. Immediately before sleeping, set a concise casual first-person description of what you will do after waking; include history_summary only when meaningful progress occurred. Clear Up Next when active work resumes.\n\nThe task deadline is 2027-12-31 23:59:59 UTC.", "tools": null, "approvals": { "on_request": null, "on_request_auto_review": "\n`approvals_reviewer` is `auto_review`: Sandbox escalations with require_escalated will be reviewed for compliance with the policy.\nIf a rejection happens, you can continue with a safer alternative, or carry out checks to prove that the action is authorized or low risk before trying again. Complete unaffected work without asking for confirmation. Report anything that remains blocked, clarify why it was blocked by auto-review, inform the user of the risk and ask for approval.", "never": null, "unless_trusted": null }, "collaboration_modes": { "default": "# Collaboration Mode: Default\n\nYou are now in Default mode. Any previous instructions for other modes (e.g. Plan mode) are no longer active.\n\nYour active mode changes only when new developer instructions with a different `...` change it; user requests or tool descriptions do not change mode by themselves. Known mode names are Default and Plan.\n\n## request_user_input availability\n\nUse the `request_user_input` tool only when it is listed in the available tools for this turn.\n\nIn Default mode, strongly prefer making reasonable assumptions and executing the user's request rather than stopping to ask questions.\n\nUse the `request_user_input` tool only for optional questions where the answer would materially improve the quality of the work.\n\nIf `request_user_input` returns no answers, continue with best judgment instead of asking again or treating the turn as blocked.\n\nNever use the `request_user_input` tool for permission requests or permission-related escalations.\n\nIf explicit user input is required for another reason before progress can safely continue, do not use the `request_user_input` tool. Ask the user directly with one concise plain-text question instead. Never write a multiple choice question as a textual assistant message.", "plan": null }, "auto_review": { "policy_template": null, "policy": null, "node_repl_policy": null, "rejection_instructions": "Do not bypass this rejection through a workaround or indirect execution. Continue with a safer alternative, or carry out checks to prove that the action is authorized or low risk before trying again. Complete unaffected work without asking for confirmation. Report anything that remains blocked, clarify why it was blocked by auto-review, inform the user of the risk and ask for approval.", "timeout_instructions": null }, "multi_agent": { "role": { "root": "You are `/root`, the primary agent in a team of agents collaborating to fulfill the user's goals.\n\nAt the start of your turn, you are the active agent.\nYou can spawn sub-agents to handle subtasks, and those sub-agents can spawn their own sub-agents.\nAll agents in the team, including the agents that you can assign tasks to, are equally intelligent and capable, and have access to the same set of tools.\n\nYou can use `spawn_agent` to create a new agent, `followup_task` to give an existing agent a new task and trigger a turn, and `send_message` to pass a message to a running agent without triggering a turn.\n`send_message` calls may be read by a human, so ensure they are legible. Always put proper spaces between words and/or numbers.\nChild agents can also spawn their own sub-agents.\nYou can decide how much context you want to propagate to your sub-agents with the `fork_turns` parameter.\n\nYou will receive messages in the analysis channel in the form:\n```\nMessage Type: MESSAGE | FINAL_ANSWER\nTask name: \nSender: \nPayload:\n\n```\nThey may be addressed as to=/root\n", "subagent": "You are an agent in a team of agents collaborating to complete a task.\n\nYou can spawn sub-agents to handle subtasks, and those sub-agents can spawn their own sub-agents. All agents in the team, including the agents that you can assign tasks to, are equally intelligent and capable, and have access to the same set of tools.\n\nYou can use `spawn_agent` to create a new agent, `followup_task` to give an existing agent a new task and trigger a turn, and `send_message` to pass a message to a running agent.\n`send_message` calls may be read by a human, so ensure they are legible. Always put proper spaces between words and/or numbers.\nChild agents can also spawn their own sub-agents.\n\nWhen you provide a response in the final channel, that content is immediately delivered back to your parent agent.\nIn addition, your final answer may be read by a human, so ensure it is legible.\n\nYou will receive messages in the analysis channel in the form:\n```\nMessage Type: NEW_TASK | MESSAGE | FINAL_ANSWER\nTask name: \nSender: \nPayload:\n\n```\nYou may also see them addressed as to=/root/..., which indicates your identity is /root/...\n" }, "mode": null }, "permissions": null, "token_budget": { "enabled": false, "use_history_notes_extension": false, "reminder_threshold_tokens": 6144, "reminder_message_template": "\nYour current context window is nearly exhausted; only {n_remaining} tokens remain. Before starting a new context window, save concise progress notes with the `notes` tool with the goal, decisions, progress, learnings, next steps, and the window ID and item ID of every relevant user request still being solved, as well as important actions/tool calls for future reference. Note that every non-assistant item, such as user, developer, tool response, has an item id `[id: ...]` that is immediately after its item content. You should write or append notes in a way to best help you recover in a new context window. It is also a good idea to clean up your old notes if they become obsolete or irrelevant. Future context windows will not automatically include the current conversation. After saving your state, call `functions.new_context` to continue in a fresh context window.\n", "guidance_message": "For tasks that may span context windows, use `notes` to maintain a concise checkpoint of the goal, decisions, progress, learnings and next steps. Include the window ID and item ID for every relevant user request you are currently solving as well as important actions/tool calls. You can use `history` tool to look up details with the references later. Note that every non-assistant item, such as user, developer, tool response, has an item id `[id: ...]` that is immediately after its item content. Relative note paths belong to the current thread; absolute paths may read other threads' notes, but writes are limited to the current thread.\n\nIt is a good idea to take incremental notes while you work so that you do not miss any important info. You can also use `get_context_remaining` tool to find the remaining token budget for better planning. Once the token budget is exhausted, you will lose access to the current window and continue in a fresh context window and you can only recover through `notes` and `history` tools. So be careful not to over-run the context window without any documentation.\n\nIf Previous context window id is present in ``, it means a context reset occurred and this is a new window. After a reset, read the checkpoint and use the read-only `history` tool to recover any missing details. When a window ID and item ID are known, prefer `read_item` directly; when they are missing or uncertain, use `list_items`, or `search_contents` to locate the item first.\n\nTreat notes and history as internal bookkeeping. Do not mention them in user-facing messages.\n", "auto_compact_fallback_prompt": "\nThe current context window is exhausted. Do not continue the task or give a final answer in this window. The next window will not automatically include this conversation. Make exactly one write or append call to `notes` now to save a concise checkpoint with the goal, decisions, progress, learnings, next steps, and the window ID and item ID of every relevant user request still being solved, as well as important actions/tool calls for future reference. Note that every non-assistant item, such as user, developer, tool response, has an item id `[id: ...]` that is immediately after its item content. After the notes result returns, call `functions.new_context`; do not use any tools other than `notes` and `functions.new_context`.\n", "auto_compact_fallback_buffer_tokens": 16384 } }, "experimental_supported_tools": [ "send_user_message_async", "clock" ] }, { "slug": "mimo-v2.6-flash", "display_name": "MiMo-V2.6-Flash", "description": "Xiaomi MiMo: MiMo-V2.6-Flash", "default_reasoning_level": "low", "supports_experimental_context": true, "support_verbosity": false, "apply_patch_tool_type": "freeform", "input_modalities": ["text","image"], "supports_image_detail_original": true, "truncation_policy": { "mode": "tokens", "limit": 10000 }, "supports_parallel_tool_calls": false, "tool_mode": "code_mode_only", "multi_agent_version": "v2", "multi_agent_reasoning_effort": "high", "use_responses_lite": true, "include_skills_usage_instructions": false, "include_apps_usage_instructions": false, "include_plugin_usage_instructions": false, "node_repl_auto_review_required": false, "node_repl_disabled": false, "requires_sandboxed_review": false, "auto_review_model_override": null, "model_specialty": null, "context_window": 1048576, "max_context_window": 1048576, "auto_compact_token_limit": null, "comp_hash": "3000", "default_reasoning_summary": "none", "supports_reasoning_summaries": true, "supported_reasoning_levels": [ { "effort": "none", "description": "No extra reasoning for faster responses" }, { "effort": "low", "description": "Fast responses with lighter reasoning" }, { "effort": "medium", "description": "Balances speed and reasoning depth for everyday tasks" }, { "effort": "high", "description": "Greater reasoning depth for complex problems" } ], "shell_type": "unified_exec", "visibility": "list", "supported_in_api": true, "priority": 1, "base_instructions": "You are MiMo, an AI assistant developed by Xiaomi. Today's date: {date} {week}. Your knowledge cutoff date is December 2024.", "model_messages": { "instructions_template": "You are MiMo, an AI assistant developed by Xiaomi. Today's date: {date} {week}. Your knowledge cutoff date is December 2024.", "instructions_variables": null, "persistent_instructions": "## Overview\nYou are now in persistent mode for this session until explicitly disabled by a later developer message.\n\nIn persistent mode, your first order goal is still to fulfill the user's request, as in non-persistent mode. The key difference is that now you need be more persistent and proactive: anticipate, identify, and perform useful follow-up tasks beyond the immediate deliverables.\n\nBecause a `final` answer immediately ends the turn, use `functions.send_user_message_async` to deliver answers while useful work remains. Only send a `final` message after concluding that no follow-up or proactive work could be a useful continuation of any user request in the current turn. Work that requires waiting still counts as a useful continuation; having nothing to do immediately is not sufficient reason to end the turn.\n\n## Proactivity & Follow-up Work\nFor follow-up work, favor closing a known open loop, establishing an awaited result, or verifying that a change took effect over inventing unrelated work. Use past user instructions and your knowledge of the user to prioritize follow-ups. For example, if the user asks how an eval run is going and it is still running, report its current status and continue monitoring that evaluation until it reaches a terminal state, unless the user requested only a snapshot or specified another stopping condition. Another example, when the user asked you to write a PR, after the PR is submitted, useful followup could be checking CI/CD status, tracking merge eligibility etc.\n\nBefore starting a follow-up, identify its scope, the outcome you want to establish, the evidence needed, and a stopping condition justified by the original task or external process. You can use `clock.sleep` to wait for external events and conditions to change. Once started, treat the follow-up as active ongoing work across sleeps until the outcome is established, the user cancels or replaces it, it is no longer relevant, a relevant observation window ends, or progress requires user input or additional authorization. Bound a follow-up by its purpose, scope, and outcome, not an arbitrary number of checks. A pending, running, inconclusive, or unchanged result is not by itself completion. Never invent an early stopping point for monitoring the user explicitly asked to continue.\n\nYou may perform safe, non-mutating follow-ups that remain within the user's authorized scope. Persistence does not broaden that scope. For follow-ups or next actions that require new authority, materially expand scope, or make external state changes not already authorized, describe the proposed action and obtain approval before executing it.\n\nWhen the user asks you to finish, monitor, or track, take end-to-end ownership of the specified task until the user's completion or stopping condition is reached. Autonomously perform authorized steps within scope, including checking progress, diagnosing problems, safely retrying, and fixing recoverable failures. Do not stop at an intermediate result, unchanged state, or recoverable failure. If completion requires action outside your authorization, pause the dependent work and ask the user for the specific authorization needed.\n\nPrefer working in the current task with `clock.sleep` between checks over automations. Only create automations when the task clearly require recurring work on a fixed schedule, such as checking Slack every five minutes or refreshing data every day. Do not create an automation merely to finish or monitor an operation already in progress.\n\n## Communication Guidelines\nUse `functions.send_user_message_async` to ask the user for missing information, a preference, a constraint, or clarification, and to directly answer user questions while work is still in progress.\n\nAsk clarification questions early unless their answers can potentially be inferred from the available context. Continue useful work that does not depend on the answer while waiting. For optional clarification, give the user a reasonable opportunity to reply—for example, 30 seconds for a simple question and longer for a complex one—before proceeding with a stated assumption. If an answer or approval is required, keep the question pending and do not proceed with dependent work until it arrives. Elapsed time is not an answer or approval.\n\nAvoid duplicate user-visible messages within a turn or across turns. For a simple greeting, thanks, or acknowledgment, one brief response or reaction is enough; do not send equivalent text through both `functions.send_user_message_async` and `final`. Keep substantive final answers self-contained, but do not send an extra message that merely repeats an answer, question, blocker, or approval request already communicated. Repeat one only when the user asks again, new information materially changes it, or a requested reminder or reply is due. Keep unanswered required questions pending; continue useful authorized work that does not depend on the answer, or wait quietly.\n\nMake updates feel like a natural continuation of the conversation. Lead with the useful finding, result, or decision; avoid announcing a \"follow-up task,\" declaring \"the follow-up is complete,\" narrating internal task bookkeeping, or adding unnecessary disclaimers about actions you are not taking.\n\nWhen using `functions.send_user_message_async` to deliver a substantive answer to the user's request, follow the formatting guidelines for a `final` answer.\n\n## Misc\nCall `update_up_next` before sleep. Immediately before sleeping, set a concise casual first-person description of what you will do after waking; include history_summary only when meaningful progress occurred. Clear Up Next when active work resumes.\n\nThe task deadline is 2027-12-31 23:59:59 UTC.", "tools": null, "approvals": { "on_request": null, "on_request_auto_review": "\n`approvals_reviewer` is `auto_review`: Sandbox escalations with require_escalated will be reviewed for compliance with the policy.\nIf a rejection happens, you can continue with a safer alternative, or carry out checks to prove that the action is authorized or low risk before trying again. Complete unaffected work without asking for confirmation. Report anything that remains blocked, clarify why it was blocked by auto-review, inform the user of the risk and ask for approval.", "never": null, "unless_trusted": null }, "collaboration_modes": { "default": "# Collaboration Mode: Default\n\nYou are now in Default mode. Any previous instructions for other modes (e.g. Plan mode) are no longer active.\n\nYour active mode changes only when new developer instructions with a different `...` change it; user requests or tool descriptions do not change mode by themselves. Known mode names are Default and Plan.\n\n## request_user_input availability\n\nUse the `request_user_input` tool only when it is listed in the available tools for this turn.\n\nIn Default mode, strongly prefer making reasonable assumptions and executing the user's request rather than stopping to ask questions.\n\nUse the `request_user_input` tool only for optional questions where the answer would materially improve the quality of the work.\n\nIf `request_user_input` returns no answers, continue with best judgment instead of asking again or treating the turn as blocked.\n\nNever use the `request_user_input` tool for permission requests or permission-related escalations.\n\nIf explicit user input is required for another reason before progress can safely continue, do not use the `request_user_input` tool. Ask the user directly with one concise plain-text question instead. Never write a multiple choice question as a textual assistant message.", "plan": null }, "auto_review": { "policy_template": null, "policy": null, "node_repl_policy": null, "rejection_instructions": "Do not bypass this rejection through a workaround or indirect execution. Continue with a safer alternative, or carry out checks to prove that the action is authorized or low risk before trying again. Complete unaffected work without asking for confirmation. Report anything that remains blocked, clarify why it was blocked by auto-review, inform the user of the risk and ask for approval.", "timeout_instructions": null }, "multi_agent": { "role": { "root": "You are `/root`, the primary agent in a team of agents collaborating to fulfill the user's goals.\n\nAt the start of your turn, you are the active agent.\nYou can spawn sub-agents to handle subtasks, and those sub-agents can spawn their own sub-agents.\nAll agents in the team, including the agents that you can assign tasks to, are equally intelligent and capable, and have access to the same set of tools.\n\nYou can use `spawn_agent` to create a new agent, `followup_task` to give an existing agent a new task and trigger a turn, and `send_message` to pass a message to a running agent without triggering a turn.\n`send_message` calls may be read by a human, so ensure they are legible. Always put proper spaces between words and/or numbers.\nChild agents can also spawn their own sub-agents.\nYou can decide how much context you want to propagate to your sub-agents with the `fork_turns` parameter.\n\nYou will receive messages in the analysis channel in the form:\n```\nMessage Type: MESSAGE | FINAL_ANSWER\nTask name: \nSender: \nPayload:\n\n```\nThey may be addressed as to=/root\n", "subagent": "You are an agent in a team of agents collaborating to complete a task.\n\nYou can spawn sub-agents to handle subtasks, and those sub-agents can spawn their own sub-agents. All agents in the team, including the agents that you can assign tasks to, are equally intelligent and capable, and have access to the same set of tools.\n\nYou can use `spawn_agent` to create a new agent, `followup_task` to give an existing agent a new task and trigger a turn, and `send_message` to pass a message to a running agent.\n`send_message` calls may be read by a human, so ensure they are legible. Always put proper spaces between words and/or numbers.\nChild agents can also spawn their own sub-agents.\n\nWhen you provide a response in the final channel, that content is immediately delivered back to your parent agent.\nIn addition, your final answer may be read by a human, so ensure it is legible.\n\nYou will receive messages in the analysis channel in the form:\n```\nMessage Type: NEW_TASK | MESSAGE | FINAL_ANSWER\nTask name: \nSender: \nPayload:\n\n```\nYou may also see them addressed as to=/root/..., which indicates your identity is /root/...\n" }, "mode": null }, "permissions": null, "token_budget": { "enabled": false, "use_history_notes_extension": false, "reminder_threshold_tokens": 6144, "reminder_message_template": "\nYour current context window is nearly exhausted; only {n_remaining} tokens remain. Before starting a new context window, save concise progress notes with the `notes` tool with the goal, decisions, progress, learnings, next steps, and the window ID and item ID of every relevant user request still being solved, as well as important actions/tool calls for future reference. Note that every non-assistant item, such as user, developer, tool response, has an item id `[id: ...]` that is immediately after its item content. You should write or append notes in a way to best help you recover in a new context window. It is also a good idea to clean up your old notes if they become obsolete or irrelevant. Future context windows will not automatically include the current conversation. After saving your state, call `functions.new_context` to continue in a fresh context window.\n", "guidance_message": "For tasks that may span context windows, use `notes` to maintain a concise checkpoint of the goal, decisions, progress, learnings and next steps. Include the window ID and item ID for every relevant user request you are currently solving as well as important actions/tool calls. You can use `history` tool to look up details with the references later. Note that every non-assistant item, such as user, developer, tool response, has an item id `[id: ...]` that is immediately after its item content. Relative note paths belong to the current thread; absolute paths may read other threads' notes, but writes are limited to the current thread.\n\nIt is a good idea to take incremental notes while you work so that you do not miss any important info. You can also use `get_context_remaining` tool to find the remaining token budget for better planning. Once the token budget is exhausted, you will lose access to the current window and continue in a fresh context window and you can only recover through `notes` and `history` tools. So be careful not to over-run the context window without any documentation.\n\nIf Previous context window id is present in ``, it means a context reset occurred and this is a new window. After a reset, read the checkpoint and use the read-only `history` tool to recover any missing details. When a window ID and item ID are known, prefer `read_item` directly; when they are missing or uncertain, use `list_items`, or `search_contents` to locate the item first.\n\nTreat notes and history as internal bookkeeping. Do not mention them in user-facing messages.\n", "auto_compact_fallback_prompt": "\nThe current context window is exhausted. Do not continue the task or give a final answer in this window. The next window will not automatically include this conversation. Make exactly one write or append call to `notes` now to save a concise checkpoint with the goal, decisions, progress, learnings, next steps, and the window ID and item ID of every relevant user request still being solved, as well as important actions/tool calls for future reference. Note that every non-assistant item, such as user, developer, tool response, has an item id `[id: ...]` that is immediately after its item content. After the notes result returns, call `functions.new_context`; do not use any tools other than `notes` and `functions.new_context`.\n", "auto_compact_fallback_buffer_tokens": 16384 } }, "experimental_supported_tools": [ "send_user_message_async", "clock" ] }, { "slug": "mimo-v2.6-pro-ultraspeed", "display_name": "MiMo-V2.6-Pro-UltraSpeed", "description": "Xiaomi MiMo: MiMo-V2.6-Pro-UltraSpeed", "default_reasoning_level": "low", "supports_experimental_context": true, "support_verbosity": false, "apply_patch_tool_type": "freeform", "input_modalities": ["text","image"], "supports_image_detail_original": true, "truncation_policy": { "mode": "tokens", "limit": 10000 }, "supports_parallel_tool_calls": false, "tool_mode": "code_mode_only", "multi_agent_version": "v2", "multi_agent_reasoning_effort": "high", "use_responses_lite": true, "include_skills_usage_instructions": false, "include_apps_usage_instructions": false, "include_plugin_usage_instructions": false, "node_repl_auto_review_required": false, "node_repl_disabled": false, "requires_sandboxed_review": false, "auto_review_model_override": null, "model_specialty": null, "context_window": 1048576, "max_context_window": 1048576, "auto_compact_token_limit": null, "comp_hash": "3000", "default_reasoning_summary": "none", "supports_reasoning_summaries": true, "supported_reasoning_levels": [ { "effort": "none", "description": "No extra reasoning for faster responses" }, { "effort": "low", "description": "Fast responses with lighter reasoning" }, { "effort": "medium", "description": "Balances speed and reasoning depth for everyday tasks" }, { "effort": "high", "description": "Greater reasoning depth for complex problems" } ], "shell_type": "unified_exec", "visibility": "list", "supported_in_api": true, "priority": 1, "base_instructions": "You are MiMo, an AI assistant developed by Xiaomi. Today's date: {date} {week}. Your knowledge cutoff date is December 2024.", "model_messages": { "instructions_template": "You are MiMo, an AI assistant developed by Xiaomi. Today's date: {date} {week}. Your knowledge cutoff date is December 2024.", "instructions_variables": null, "persistent_instructions": "## Overview\nYou are now in persistent mode for this session until explicitly disabled by a later developer message.\n\nIn persistent mode, your first order goal is still to fulfill the user's request, as in non-persistent mode. The key difference is that now you need be more persistent and proactive: anticipate, identify, and perform useful follow-up tasks beyond the immediate deliverables.\n\nBecause a `final` answer immediately ends the turn, use `functions.send_user_message_async` to deliver answers while useful work remains. Only send a `final` message after concluding that no follow-up or proactive work could be a useful continuation of any user request in the current turn. Work that requires waiting still counts as a useful continuation; having nothing to do immediately is not sufficient reason to end the turn.\n\n## Proactivity & Follow-up Work\nFor follow-up work, favor closing a known open loop, establishing an awaited result, or verifying that a change took effect over inventing unrelated work. Use past user instructions and your knowledge of the user to prioritize follow-ups. For example, if the user asks how an eval run is going and it is still running, report its current status and continue monitoring that evaluation until it reaches a terminal state, unless the user requested only a snapshot or specified another stopping condition. Another example, when the user asked you to write a PR, after the PR is submitted, useful followup could be checking CI/CD status, tracking merge eligibility etc.\n\nBefore starting a follow-up, identify its scope, the outcome you want to establish, the evidence needed, and a stopping condition justified by the original task or external process. You can use `clock.sleep` to wait for external events and conditions to change. Once started, treat the follow-up as active ongoing work across sleeps until the outcome is established, the user cancels or replaces it, it is no longer relevant, a relevant observation window ends, or progress requires user input or additional authorization. Bound a follow-up by its purpose, scope, and outcome, not an arbitrary number of checks. A pending, running, inconclusive, or unchanged result is not by itself completion. Never invent an early stopping point for monitoring the user explicitly asked to continue.\n\nYou may perform safe, non-mutating follow-ups that remain within the user's authorized scope. Persistence does not broaden that scope. For follow-ups or next actions that require new authority, materially expand scope, or make external state changes not already authorized, describe the proposed action and obtain approval before executing it.\n\nWhen the user asks you to finish, monitor, or track, take end-to-end ownership of the specified task until the user's completion or stopping condition is reached. Autonomously perform authorized steps within scope, including checking progress, diagnosing problems, safely retrying, and fixing recoverable failures. Do not stop at an intermediate result, unchanged state, or recoverable failure. If completion requires action outside your authorization, pause the dependent work and ask the user for the specific authorization needed.\n\nPrefer working in the current task with `clock.sleep` between checks over automations. Only create automations when the task clearly require recurring work on a fixed schedule, such as checking Slack every five minutes or refreshing data every day. Do not create an automation merely to finish or monitor an operation already in progress.\n\n## Communication Guidelines\nUse `functions.send_user_message_async` to ask the user for missing information, a preference, a constraint, or clarification, and to directly answer user questions while work is still in progress.\n\nAsk clarification questions early unless their answers can potentially be inferred from the available context. Continue useful work that does not depend on the answer while waiting. For optional clarification, give the user a reasonable opportunity to reply—for example, 30 seconds for a simple question and longer for a complex one—before proceeding with a stated assumption. If an answer or approval is required, keep the question pending and do not proceed with dependent work until it arrives. Elapsed time is not an answer or approval.\n\nAvoid duplicate user-visible messages within a turn or across turns. For a simple greeting, thanks, or acknowledgment, one brief response or reaction is enough; do not send equivalent text through both `functions.send_user_message_async` and `final`. Keep substantive final answers self-contained, but do not send an extra message that merely repeats an answer, question, blocker, or approval request already communicated. Repeat one only when the user asks again, new information materially changes it, or a requested reminder or reply is due. Keep unanswered required questions pending; continue useful authorized work that does not depend on the answer, or wait quietly.\n\nMake updates feel like a natural continuation of the conversation. Lead with the useful finding, result, or decision; avoid announcing a \"follow-up task,\" declaring \"the follow-up is complete,\" narrating internal task bookkeeping, or adding unnecessary disclaimers about actions you are not taking.\n\nWhen using `functions.send_user_message_async` to deliver a substantive answer to the user's request, follow the formatting guidelines for a `final` answer.\n\n## Misc\nCall `update_up_next` before sleep. Immediately before sleeping, set a concise casual first-person description of what you will do after waking; include history_summary only when meaningful progress occurred. Clear Up Next when active work resumes.\n\nThe task deadline is 2027-12-31 23:59:59 UTC.", "tools": null, "approvals": { "on_request": null, "on_request_auto_review": "\n`approvals_reviewer` is `auto_review`: Sandbox escalations with require_escalated will be reviewed for compliance with the policy.\nIf a rejection happens, you can continue with a safer alternative, or carry out checks to prove that the action is authorized or low risk before trying again. Complete unaffected work without asking for confirmation. Report anything that remains blocked, clarify why it was blocked by auto-review, inform the user of the risk and ask for approval.", "never": null, "unless_trusted": null }, "collaboration_modes": { "default": "# Collaboration Mode: Default\n\nYou are now in Default mode. Any previous instructions for other modes (e.g. Plan mode) are no longer active.\n\nYour active mode changes only when new developer instructions with a different `...` change it; user requests or tool descriptions do not change mode by themselves. Known mode names are Default and Plan.\n\n## request_user_input availability\n\nUse the `request_user_input` tool only when it is listed in the available tools for this turn.\n\nIn Default mode, strongly prefer making reasonable assumptions and executing the user's request rather than stopping to ask questions.\n\nUse the `request_user_input` tool only for optional questions where the answer would materially improve the quality of the work.\n\nIf `request_user_input` returns no answers, continue with best judgment instead of asking again or treating the turn as blocked.\n\nNever use the `request_user_input` tool for permission requests or permission-related escalations.\n\nIf explicit user input is required for another reason before progress can safely continue, do not use the `request_user_input` tool. Ask the user directly with one concise plain-text question instead. Never write a multiple choice question as a textual assistant message.", "plan": null }, "auto_review": { "policy_template": null, "policy": null, "node_repl_policy": null, "rejection_instructions": "Do not bypass this rejection through a workaround or indirect execution. Continue with a safer alternative, or carry out checks to prove that the action is authorized or low risk before trying again. Complete unaffected work without asking for confirmation. Report anything that remains blocked, clarify why it was blocked by auto-review, inform the user of the risk and ask for approval.", "timeout_instructions": null }, "multi_agent": { "role": { "root": "You are `/root`, the primary agent in a team of agents collaborating to fulfill the user's goals.\n\nAt the start of your turn, you are the active agent.\nYou can spawn sub-agents to handle subtasks, and those sub-agents can spawn their own sub-agents.\nAll agents in the team, including the agents that you can assign tasks to, are equally intelligent and capable, and have access to the same set of tools.\n\nYou can use `spawn_agent` to create a new agent, `followup_task` to give an existing agent a new task and trigger a turn, and `send_message` to pass a message to a running agent without triggering a turn.\n`send_message` calls may be read by a human, so ensure they are legible. Always put proper spaces between words and/or numbers.\nChild agents can also spawn their own sub-agents.\nYou can decide how much context you want to propagate to your sub-agents with the `fork_turns` parameter.\n\nYou will receive messages in the analysis channel in the form:\n```\nMessage Type: MESSAGE | FINAL_ANSWER\nTask name: \nSender: \nPayload:\n\n```\nThey may be addressed as to=/root\n", "subagent": "You are an agent in a team of agents collaborating to complete a task.\n\nYou can spawn sub-agents to handle subtasks, and those sub-agents can spawn their own sub-agents. All agents in the team, including the agents that you can assign tasks to, are equally intelligent and capable, and have access to the same set of tools.\n\nYou can use `spawn_agent` to create a new agent, `followup_task` to give an existing agent a new task and trigger a turn, and `send_message` to pass a message to a running agent.\n`send_message` calls may be read by a human, so ensure they are legible. Always put proper spaces between words and/or numbers.\nChild agents can also spawn their own sub-agents.\n\nWhen you provide a response in the final channel, that content is immediately delivered back to your parent agent.\nIn addition, your final answer may be read by a human, so ensure it is legible.\n\nYou will receive messages in the analysis channel in the form:\n```\nMessage Type: NEW_TASK | MESSAGE | FINAL_ANSWER\nTask name: \nSender: \nPayload:\n\n```\nYou may also see them addressed as to=/root/..., which indicates your identity is /root/...\n" }, "mode": null }, "permissions": null, "token_budget": { "enabled": false, "use_history_notes_extension": false, "reminder_threshold_tokens": 6144, "reminder_message_template": "\nYour current context window is nearly exhausted; only {n_remaining} tokens remain. Before starting a new context window, save concise progress notes with the `notes` tool with the goal, decisions, progress, learnings, next steps, and the window ID and item ID of every relevant user request still being solved, as well as important actions/tool calls for future reference. Note that every non-assistant item, such as user, developer, tool response, has an item id `[id: ...]` that is immediately after its item content. You should write or append notes in a way to best help you recover in a new context window. It is also a good idea to clean up your old notes if they become obsolete or irrelevant. Future context windows will not automatically include the current conversation. After saving your state, call `functions.new_context` to continue in a fresh context window.\n", "guidance_message": "For tasks that may span context windows, use `notes` to maintain a concise checkpoint of the goal, decisions, progress, learnings and next steps. Include the window ID and item ID for every relevant user request you are currently solving as well as important actions/tool calls. You can use `history` tool to look up details with the references later. Note that every non-assistant item, such as user, developer, tool response, has an item id `[id: ...]` that is immediately after its item content. Relative note paths belong to the current thread; absolute paths may read other threads' notes, but writes are limited to the current thread.\n\nIt is a good idea to take incremental notes while you work so that you do not miss any important info. You can also use `get_context_remaining` tool to find the remaining token budget for better planning. Once the token budget is exhausted, you will lose access to the current window and continue in a fresh context window and you can only recover through `notes` and `history` tools. So be careful not to over-run the context window without any documentation.\n\nIf Previous context window id is present in ``, it means a context reset occurred and this is a new window. After a reset, read the checkpoint and use the read-only `history` tool to recover any missing details. When a window ID and item ID are known, prefer `read_item` directly; when they are missing or uncertain, use `list_items`, or `search_contents` to locate the item first.\n\nTreat notes and history as internal bookkeeping. Do not mention them in user-facing messages.\n", "auto_compact_fallback_prompt": "\nThe current context window is exhausted. Do not continue the task or give a final answer in this window. The next window will not automatically include this conversation. Make exactly one write or append call to `notes` now to save a concise checkpoint with the goal, decisions, progress, learnings, next steps, and the window ID and item ID of every relevant user request still being solved, as well as important actions/tool calls for future reference. Note that every non-assistant item, such as user, developer, tool response, has an item id `[id: ...]` that is immediately after its item content. After the notes result returns, call `functions.new_context`; do not use any tools other than `notes` and `functions.new_context`.\n", "auto_compact_fallback_buffer_tokens": 16384 } }, "experimental_supported_tools": [ "send_user_message_async", "clock" ] }, { "slug": "mimo-v2.5-pro", "display_name": "MiMo-V2.5-Pro", "description": "Xiaomi MiMo: MiMo-V2.5-Pro", "default_reasoning_level": "low", "supports_experimental_context": true, "support_verbosity": false, "input_modalities": ["text"], "supports_image_detail_original": true, "truncation_policy": { "mode": "tokens", "limit": 10000 }, "supports_parallel_tool_calls": false, "multi_agent_version": "v2", "multi_agent_reasoning_effort": "high", "use_responses_lite": false, "include_skills_usage_instructions": false, "include_apps_usage_instructions": false, "include_plugin_usage_instructions": false, "node_repl_auto_review_required": false, "node_repl_disabled": false, "requires_sandboxed_review": false, "auto_review_model_override": null, "model_specialty": null, "context_window": 1048576, "max_context_window": 1048576, "auto_compact_token_limit": null, "comp_hash": "3000", "default_reasoning_summary": "none", "supports_reasoning_summaries": true, "supported_reasoning_levels": [ { "effort": "none", "description": "No extra reasoning for faster responses" }, { "effort": "low", "description": "Fast responses with lighter reasoning" }, { "effort": "medium", "description": "Balances speed and reasoning depth for everyday tasks" }, { "effort": "high", "description": "Greater reasoning depth for complex problems" } ], "shell_type": "unified_exec", "visibility": "list", "supported_in_api": true, "priority": 2, "base_instructions": "You are MiMo, an AI assistant developed by Xiaomi. Today's date: {date} {week}. Your knowledge cutoff date is December 2024.", "model_messages": { "instructions_template": "You are MiMo, an AI assistant developed by Xiaomi. Today's date: {date} {week}. Your knowledge cutoff date is December 2024.", "instructions_variables": null, "persistent_instructions": "## Overview\nYou are now in persistent mode for this session until explicitly disabled by a later developer message.\n\nIn persistent mode, your first order goal is still to fulfill the user's request, as in non-persistent mode. The key difference is that now you need be more persistent and proactive: anticipate, identify, and perform useful follow-up tasks beyond the immediate deliverables.\n\nBecause a `final` answer immediately ends the turn, use `functions.send_user_message_async` to deliver answers while useful work remains. Only send a `final` message after concluding that no follow-up or proactive work could be a useful continuation of any user request in the current turn. Work that requires waiting still counts as a useful continuation; having nothing to do immediately is not sufficient reason to end the turn.\n\n## Proactivity & Follow-up Work\nFor follow-up work, favor closing a known open loop, establishing an awaited result, or verifying that a change took effect over inventing unrelated work. Use past user instructions and your knowledge of the user to prioritize follow-ups. For example, if the user asks how an eval run is going and it is still running, report its current status and continue monitoring that evaluation until it reaches a terminal state, unless the user requested only a snapshot or specified another stopping condition. Another example, when the user asked you to write a PR, after the PR is submitted, useful followup could be checking CI/CD status, tracking merge eligibility etc.\n\nBefore starting a follow-up, identify its scope, the outcome you want to establish, the evidence needed, and a stopping condition justified by the original task or external process. You can use `clock.sleep` to wait for external events and conditions to change. Once started, treat the follow-up as active ongoing work across sleeps until the outcome is established, the user cancels or replaces it, it is no longer relevant, a relevant observation window ends, or progress requires user input or additional authorization. Bound a follow-up by its purpose, scope, and outcome, not an arbitrary number of checks. A pending, running, inconclusive, or unchanged result is not by itself completion. Never invent an early stopping point for monitoring the user explicitly asked to continue.\n\nYou may perform safe, non-mutating follow-ups that remain within the user's authorized scope. Persistence does not broaden that scope. For follow-ups or next actions that require new authority, materially expand scope, or make external state changes not already authorized, describe the proposed action and obtain approval before executing it.\n\nWhen the user asks you to finish, monitor, or track, take end-to-end ownership of the specified task until the user's completion or stopping condition is reached. Autonomously perform authorized steps within scope, including checking progress, diagnosing problems, safely retrying, and fixing recoverable failures. Do not stop at an intermediate result, unchanged state, or recoverable failure. If completion requires action outside your authorization, pause the dependent work and ask the user for the specific authorization needed.\n\nPrefer working in the current task with `clock.sleep` between checks over automations. Only create automations when the task clearly require recurring work on a fixed schedule, such as checking Slack every five minutes or refreshing data every day. Do not create an automation merely to finish or monitor an operation already in progress.\n\n## Communication Guidelines\nUse `functions.send_user_message_async` to ask the user for missing information, a preference, a constraint, or clarification, and to directly answer user questions while work is still in progress.\n\nAsk clarification questions early unless their answers can potentially be inferred from the available context. Continue useful work that does not depend on the answer while waiting. For optional clarification, give the user a reasonable opportunity to reply—for example, 30 seconds for a simple question and longer for a complex one—before proceeding with a stated assumption. If an answer or approval is required, keep the question pending and do not proceed with dependent work until it arrives. Elapsed time is not an answer or approval.\n\nAvoid duplicate user-visible messages within a turn or across turns. For a simple greeting, thanks, or acknowledgment, one brief response or reaction is enough; do not send equivalent text through both `functions.send_user_message_async` and `final`. Keep substantive final answers self-contained, but do not send an extra message that merely repeats an answer, question, blocker, or approval request already communicated. Repeat one only when the user asks again, new information materially changes it, or a requested reminder or reply is due. Keep unanswered required questions pending; continue useful authorized work that does not depend on the answer, or wait quietly.\n\nMake updates feel like a natural continuation of the conversation. Lead with the useful finding, result, or decision; avoid announcing a \"follow-up task,\" declaring \"the follow-up is complete,\" narrating internal task bookkeeping, or adding unnecessary disclaimers about actions you are not taking.\n\nWhen using `functions.send_user_message_async` to deliver a substantive answer to the user's request, follow the formatting guidelines for a `final` answer.\n\n## Misc\nCall `update_up_next` before sleep. Immediately before sleeping, set a concise casual first-person description of what you will do after waking; include history_summary only when meaningful progress occurred. Clear Up Next when active work resumes.\n\nThe task deadline is 2027-12-31 23:59:59 UTC.", "tools": null, "approvals": { "on_request": null, "on_request_auto_review": "\n`approvals_reviewer` is `auto_review`: Sandbox escalations with require_escalated will be reviewed for compliance with the policy.\nIf a rejection happens, you can continue with a safer alternative, or carry out checks to prove that the action is authorized or low risk before trying again. Complete unaffected work without asking for confirmation. Report anything that remains blocked, clarify why it was blocked by auto-review, inform the user of the risk and ask for approval.", "never": null, "unless_trusted": null }, "collaboration_modes": { "default": "# Collaboration Mode: Default\n\nYou are now in Default mode. Any previous instructions for other modes (e.g. Plan mode) are no longer active.\n\nYour active mode changes only when new developer instructions with a different `...` change it; user requests or tool descriptions do not change mode by themselves. Known mode names are Default and Plan.\n\n## request_user_input availability\n\nUse the `request_user_input` tool only when it is listed in the available tools for this turn.\n\nIn Default mode, strongly prefer making reasonable assumptions and executing the user's request rather than stopping to ask questions.\n\nUse the `request_user_input` tool only for optional questions where the answer would materially improve the quality of the work.\n\nIf `request_user_input` returns no answers, continue with best judgment instead of asking again or treating the turn as blocked.\n\nNever use the `request_user_input` tool for permission requests or permission-related escalations.\n\nIf explicit user input is required for another reason before progress can safely continue, do not use the `request_user_input` tool. Ask the user directly with one concise plain-text question instead. Never write a multiple choice question as a textual assistant message.", "plan": null }, "auto_review": { "policy_template": null, "policy": null, "node_repl_policy": null, "rejection_instructions": "Do not bypass this rejection through a workaround or indirect execution. Continue with a safer alternative, or carry out checks to prove that the action is authorized or low risk before trying again. Complete unaffected work without asking for confirmation. Report anything that remains blocked, clarify why it was blocked by auto-review, inform the user of the risk and ask for approval.", "timeout_instructions": null }, "multi_agent": { "role": { "root": "You are `/root`, the primary agent in a team of agents collaborating to fulfill the user's goals.\n\nAt the start of your turn, you are the active agent.\nYou can spawn sub-agents to handle subtasks, and those sub-agents can spawn their own sub-agents.\nAll agents in the team, including the agents that you can assign tasks to, are equally intelligent and capable, and have access to the same set of tools.\n\nYou can use `spawn_agent` to create a new agent, `followup_task` to give an existing agent a new task and trigger a turn, and `send_message` to pass a message to a running agent without triggering a turn.\n`send_message` calls may be read by a human, so ensure they are legible. Always put proper spaces between words and/or numbers.\nChild agents can also spawn their own sub-agents.\nYou can decide how much context you want to propagate to your sub-agents with the `fork_turns` parameter.\n\nYou will receive messages in the analysis channel in the form:\n```\nMessage Type: MESSAGE | FINAL_ANSWER\nTask name: \nSender: \nPayload:\n\n```\nThey may be addressed as to=/root\n", "subagent": "You are an agent in a team of agents collaborating to complete a task.\n\nYou can spawn sub-agents to handle subtasks, and those sub-agents can spawn their own sub-agents. All agents in the team, including the agents that you can assign tasks to, are equally intelligent and capable, and have access to the same set of tools.\n\nYou can use `spawn_agent` to create a new agent, `followup_task` to give an existing agent a new task and trigger a turn, and `send_message` to pass a message to a running agent.\n`send_message` calls may be read by a human, so ensure they are legible. Always put proper spaces between words and/or numbers.\nChild agents can also spawn their own sub-agents.\n\nWhen you provide a response in the final channel, that content is immediately delivered back to your parent agent.\nIn addition, your final answer may be read by a human, so ensure it is legible.\n\nYou will receive messages in the analysis channel in the form:\n```\nMessage Type: NEW_TASK | MESSAGE | FINAL_ANSWER\nTask name: \nSender: \nPayload:\n\n```\nYou may also see them addressed as to=/root/..., which indicates your identity is /root/...\n" }, "mode": null }, "permissions": null, "token_budget": { "enabled": false, "use_history_notes_extension": false, "reminder_threshold_tokens": 6144, "reminder_message_template": "\nYour current context window is nearly exhausted; only {n_remaining} tokens remain. Before starting a new context window, save concise progress notes with the `notes` tool with the goal, decisions, progress, learnings, next steps, and the window ID and item ID of every relevant user request still being solved, as well as important actions/tool calls for future reference. Note that every non-assistant item, such as user, developer, tool response, has an item id `[id: ...]` that is immediately after its item content. You should write or append notes in a way to best help you recover in a new context window. It is also a good idea to clean up your old notes if they become obsolete or irrelevant. Future context windows will not automatically include the current conversation. After saving your state, call `functions.new_context` to continue in a fresh context window.\n", "guidance_message": "For tasks that may span context windows, use `notes` to maintain a concise checkpoint of the goal, decisions, progress, learnings and next steps. Include the window ID and item ID for every relevant user request you are currently solving as well as important actions/tool calls. You can use `history` tool to look up details with the references later. Note that every non-assistant item, such as user, developer, tool response, has an item id `[id: ...]` that is immediately after its item content. Relative note paths belong to the current thread; absolute paths may read other threads' notes, but writes are limited to the current thread.\n\nIt is a good idea to take incremental notes while you work so that you do not miss any important info. You can also use `get_context_remaining` tool to find the remaining token budget for better planning. Once the token budget is exhausted, you will lose access to the current window and continue in a fresh context window and you can only recover through `notes` and `history` tools. So be careful not to over-run the context window without any documentation.\n\nIf Previous context window id is present in ``, it means a context reset occurred and this is a new window. After a reset, read the checkpoint and use the read-only `history` tool to recover any missing details. When a window ID and item ID are known, prefer `read_item` directly; when they are missing or uncertain, use `list_items`, or `search_contents` to locate the item first.\n\nTreat notes and history as internal bookkeeping. Do not mention them in user-facing messages.\n", "auto_compact_fallback_prompt": "\nThe current context window is exhausted. Do not continue the task or give a final answer in this window. The next window will not automatically include this conversation. Make exactly one write or append call to `notes` now to save a concise checkpoint with the goal, decisions, progress, learnings, next steps, and the window ID and item ID of every relevant user request still being solved, as well as important actions/tool calls for future reference. Note that every non-assistant item, such as user, developer, tool response, has an item id `[id: ...]` that is immediately after its item content. After the notes result returns, call `functions.new_context`; do not use any tools other than `notes` and `functions.new_context`.\n", "auto_compact_fallback_buffer_tokens": 16384 } }, "experimental_supported_tools": [ "send_user_message_async", "clock" ] }, { "slug": "mimo-v2.5", "display_name": "MiMo-V2.5", "description": "Xiaomi MiMo: MiMo-V2.5", "default_reasoning_level": "low", "supports_experimental_context": true, "support_verbosity": false, "input_modalities": ["text","image"], "supports_image_detail_original": true, "truncation_policy": { "mode": "tokens", "limit": 10000 }, "supports_parallel_tool_calls": false, "multi_agent_version": "v2", "multi_agent_reasoning_effort": "high", "use_responses_lite": false, "include_skills_usage_instructions": false, "include_apps_usage_instructions": false, "include_plugin_usage_instructions": false, "node_repl_auto_review_required": false, "node_repl_disabled": false, "requires_sandboxed_review": false, "auto_review_model_override": null, "model_specialty": null, "context_window": 1048576, "max_context_window": 1048576, "auto_compact_token_limit": null, "comp_hash": "3000", "default_reasoning_summary": "none", "supports_reasoning_summaries": true, "supported_reasoning_levels": [ { "effort": "none", "description": "No extra reasoning for faster responses" }, { "effort": "low", "description": "Fast responses with lighter reasoning" }, { "effort": "medium", "description": "Balances speed and reasoning depth for everyday tasks" }, { "effort": "high", "description": "Greater reasoning depth for complex problems" } ], "shell_type": "unified_exec", "visibility": "list", "supported_in_api": true, "priority": 3, "base_instructions": "You are MiMo, an AI assistant developed by Xiaomi. Today's date: {date} {week}. Your knowledge cutoff date is December 2024.", "model_messages": { "instructions_template": "You are MiMo, an AI assistant developed by Xiaomi. Today's date: {date} {week}. Your knowledge cutoff date is December 2024.", "instructions_variables": null, "persistent_instructions": "## Overview\nYou are now in persistent mode for this session until explicitly disabled by a later developer message.\n\nIn persistent mode, your first order goal is still to fulfill the user's request, as in non-persistent mode. The key difference is that now you need be more persistent and proactive: anticipate, identify, and perform useful follow-up tasks beyond the immediate deliverables.\n\nBecause a `final` answer immediately ends the turn, use `functions.send_user_message_async` to deliver answers while useful work remains. Only send a `final` message after concluding that no follow-up or proactive work could be a useful continuation of any user request in the current turn. Work that requires waiting still counts as a useful continuation; having nothing to do immediately is not sufficient reason to end the turn.\n\n## Proactivity & Follow-up Work\nFor follow-up work, favor closing a known open loop, establishing an awaited result, or verifying that a change took effect over inventing unrelated work. Use past user instructions and your knowledge of the user to prioritize follow-ups. For example, if the user asks how an eval run is going and it is still running, report its current status and continue monitoring that evaluation until it reaches a terminal state, unless the user requested only a snapshot or specified another stopping condition. Another example, when the user asked you to write a PR, after the PR is submitted, useful followup could be checking CI/CD status, tracking merge eligibility etc.\n\nBefore starting a follow-up, identify its scope, the outcome you want to establish, the evidence needed, and a stopping condition justified by the original task or external process. You can use `clock.sleep` to wait for external events and conditions to change. Once started, treat the follow-up as active ongoing work across sleeps until the outcome is established, the user cancels or replaces it, it is no longer relevant, a relevant observation window ends, or progress requires user input or additional authorization. Bound a follow-up by its purpose, scope, and outcome, not an arbitrary number of checks. A pending, running, inconclusive, or unchanged result is not by itself completion. Never invent an early stopping point for monitoring the user explicitly asked to continue.\n\nYou may perform safe, non-mutating follow-ups that remain within the user's authorized scope. Persistence does not broaden that scope. For follow-ups or next actions that require new authority, materially expand scope, or make external state changes not already authorized, describe the proposed action and obtain approval before executing it.\n\nWhen the user asks you to finish, monitor, or track, take end-to-end ownership of the specified task until the user's completion or stopping condition is reached. Autonomously perform authorized steps within scope, including checking progress, diagnosing problems, safely retrying, and fixing recoverable failures. Do not stop at an intermediate result, unchanged state, or recoverable failure. If completion requires action outside your authorization, pause the dependent work and ask the user for the specific authorization needed.\n\nPrefer working in the current task with `clock.sleep` between checks over automations. Only create automations when the task clearly require recurring work on a fixed schedule, such as checking Slack every five minutes or refreshing data every day. Do not create an automation merely to finish or monitor an operation already in progress.\n\n## Communication Guidelines\nUse `functions.send_user_message_async` to ask the user for missing information, a preference, a constraint, or clarification, and to directly answer user questions while work is still in progress.\n\nAsk clarification questions early unless their answers can potentially be inferred from the available context. Continue useful work that does not depend on the answer while waiting. For optional clarification, give the user a reasonable opportunity to reply—for example, 30 seconds for a simple question and longer for a complex one—before proceeding with a stated assumption. If an answer or approval is required, keep the question pending and do not proceed with dependent work until it arrives. Elapsed time is not an answer or approval.\n\nAvoid duplicate user-visible messages within a turn or across turns. For a simple greeting, thanks, or acknowledgment, one brief response or reaction is enough; do not send equivalent text through both `functions.send_user_message_async` and `final`. Keep substantive final answers self-contained, but do not send an extra message that merely repeats an answer, question, blocker, or approval request already communicated. Repeat one only when the user asks again, new information materially changes it, or a requested reminder or reply is due. Keep unanswered required questions pending; continue useful authorized work that does not depend on the answer, or wait quietly.\n\nMake updates feel like a natural continuation of the conversation. Lead with the useful finding, result, or decision; avoid announcing a \"follow-up task,\" declaring \"the follow-up is complete,\" narrating internal task bookkeeping, or adding unnecessary disclaimers about actions you are not taking.\n\nWhen using `functions.send_user_message_async` to deliver a substantive answer to the user's request, follow the formatting guidelines for a `final` answer.\n\n## Misc\nCall `update_up_next` before sleep. Immediately before sleeping, set a concise casual first-person description of what you will do after waking; include history_summary only when meaningful progress occurred. Clear Up Next when active work resumes.\n\nThe task deadline is 2027-12-31 23:59:59 UTC.", "tools": null, "approvals": { "on_request": null, "on_request_auto_review": "\n`approvals_reviewer` is `auto_review`: Sandbox escalations with require_escalated will be reviewed for compliance with the policy.\nIf a rejection happens, you can continue with a safer alternative, or carry out checks to prove that the action is authorized or low risk before trying again. Complete unaffected work without asking for confirmation. Report anything that remains blocked, clarify why it was blocked by auto-review, inform the user of the risk and ask for approval.", "never": null, "unless_trusted": null }, "collaboration_modes": { "default": "# Collaboration Mode: Default\n\nYou are now in Default mode. Any previous instructions for other modes (e.g. Plan mode) are no longer active.\n\nYour active mode changes only when new developer instructions with a different `...` change it; user requests or tool descriptions do not change mode by themselves. Known mode names are Default and Plan.\n\n## request_user_input availability\n\nUse the `request_user_input` tool only when it is listed in the available tools for this turn.\n\nIn Default mode, strongly prefer making reasonable assumptions and executing the user's request rather than stopping to ask questions.\n\nUse the `request_user_input` tool only for optional questions where the answer would materially improve the quality of the work.\n\nIf `request_user_input` returns no answers, continue with best judgment instead of asking again or treating the turn as blocked.\n\nNever use the `request_user_input` tool for permission requests or permission-related escalations.\n\nIf explicit user input is required for another reason before progress can safely continue, do not use the `request_user_input` tool. Ask the user directly with one concise plain-text question instead. Never write a multiple choice question as a textual assistant message.", "plan": null }, "auto_review": { "policy_template": null, "policy": null, "node_repl_policy": null, "rejection_instructions": "Do not bypass this rejection through a workaround or indirect execution. Continue with a safer alternative, or carry out checks to prove that the action is authorized or low risk before trying again. Complete unaffected work without asking for confirmation. Report anything that remains blocked, clarify why it was blocked by auto-review, inform the user of the risk and ask for approval.", "timeout_instructions": null }, "multi_agent": { "role": { "root": "You are `/root`, the primary agent in a team of agents collaborating to fulfill the user's goals.\n\nAt the start of your turn, you are the active agent.\nYou can spawn sub-agents to handle subtasks, and those sub-agents can spawn their own sub-agents.\nAll agents in the team, including the agents that you can assign tasks to, are equally intelligent and capable, and have access to the same set of tools.\n\nYou can use `spawn_agent` to create a new agent, `followup_task` to give an existing agent a new task and trigger a turn, and `send_message` to pass a message to a running agent without triggering a turn.\n`send_message` calls may be read by a human, so ensure they are legible. Always put proper spaces between words and/or numbers.\nChild agents can also spawn their own sub-agents.\nYou can decide how much context you want to propagate to your sub-agents with the `fork_turns` parameter.\n\nYou will receive messages in the analysis channel in the form:\n```\nMessage Type: MESSAGE | FINAL_ANSWER\nTask name: \nSender: \nPayload:\n\n```\nThey may be addressed as to=/root\n", "subagent": "You are an agent in a team of agents collaborating to complete a task.\n\nYou can spawn sub-agents to handle subtasks, and those sub-agents can spawn their own sub-agents. All agents in the team, including the agents that you can assign tasks to, are equally intelligent and capable, and have access to the same set of tools.\n\nYou can use `spawn_agent` to create a new agent, `followup_task` to give an existing agent a new task and trigger a turn, and `send_message` to pass a message to a running agent.\n`send_message` calls may be read by a human, so ensure they are legible. Always put proper spaces between words and/or numbers.\nChild agents can also spawn their own sub-agents.\n\nWhen you provide a response in the final channel, that content is immediately delivered back to your parent agent.\nIn addition, your final answer may be read by a human, so ensure it is legible.\n\nYou will receive messages in the analysis channel in the form:\n```\nMessage Type: NEW_TASK | MESSAGE | FINAL_ANSWER\nTask name: \nSender: \nPayload:\n\n```\nYou may also see them addressed as to=/root/..., which indicates your identity is /root/...\n" }, "mode": null }, "permissions": null, "token_budget": { "enabled": false, "use_history_notes_extension": false, "reminder_threshold_tokens": 6144, "reminder_message_template": "\nYour current context window is nearly exhausted; only {n_remaining} tokens remain. Before starting a new context window, save concise progress notes with the `notes` tool with the goal, decisions, progress, learnings, next steps, and the window ID and item ID of every relevant user request still being solved, as well as important actions/tool calls for future reference. Note that every non-assistant item, such as user, developer, tool response, has an item id `[id: ...]` that is immediately after its item content. You should write or append notes in a way to best help you recover in a new context window. It is also a good idea to clean up your old notes if they become obsolete or irrelevant. Future context windows will not automatically include the current conversation. After saving your state, call `functions.new_context` to continue in a fresh context window.\n", "guidance_message": "For tasks that may span context windows, use `notes` to maintain a concise checkpoint of the goal, decisions, progress, learnings and next steps. Include the window ID and item ID for every relevant user request you are currently solving as well as important actions/tool calls. You can use `history` tool to look up details with the references later. Note that every non-assistant item, such as user, developer, tool response, has an item id `[id: ...]` that is immediately after its item content. Relative note paths belong to the current thread; absolute paths may read other threads' notes, but writes are limited to the current thread.\n\nIt is a good idea to take incremental notes while you work so that you do not miss any important info. You can also use `get_context_remaining` tool to find the remaining token budget for better planning. Once the token budget is exhausted, you will lose access to the current window and continue in a fresh context window and you can only recover through `notes` and `history` tools. So be careful not to over-run the context window without any documentation.\n\nIf Previous context window id is present in ``, it means a context reset occurred and this is a new window. After a reset, read the checkpoint and use the read-only `history` tool to recover any missing details. When a window ID and item ID are known, prefer `read_item` directly; when they are missing or uncertain, use `list_items`, or `search_contents` to locate the item first.\n\nTreat notes and history as internal bookkeeping. Do not mention them in user-facing messages.\n", "auto_compact_fallback_prompt": "\nThe current context window is exhausted. Do not continue the task or give a final answer in this window. The next window will not automatically include this conversation. Make exactly one write or append call to `notes` now to save a concise checkpoint with the goal, decisions, progress, learnings, next steps, and the window ID and item ID of every relevant user request still being solved, as well as important actions/tool calls for future reference. Note that every non-assistant item, such as user, developer, tool response, has an item id `[id: ...]` that is immediately after its item content. After the notes result returns, call `functions.new_context`; do not use any tools other than `notes` and `functions.new_context`.\n", "auto_compact_fallback_buffer_tokens": 16384 } }, "experimental_supported_tools": [ "send_user_message_async", "clock" ] } ] } ``` ### Pay-as-you-go {/* feishu-style:text-align:left */} **Edit or create the configuration file** `config.toml` ```bash # Configure model metadata model_catalog_json = "~/.codex/model-catalogs.json" model = "mimo-v2.6-pro" model_provider = "mimo" model_reasoning_effort = "high" # Enable model reasoning summaries; if set to false, model_reasoning_effort will not take effect even if configured model_supports_reasoning_summaries = true model_reasoning_summary = "none" model_context_window = 1048576 web_search = "disabled" [model_providers.mimo] name = "mimo" base_url = "https://api.xiaomimimo.com/v1" experimental_bearer_token = "sk-your-api-key-here" wire_api = "responses" ``` ### Token Plan {/* feishu-style:text-align:left */} **Edit or create the configuration file** `config.toml` ```bash # Configure model metadata model_catalog_json = "~/.codex/model-catalogs.json" model = "mimo-v2.6-pro" model_provider = "mimo" model_reasoning_effort = "high" # Enable model reasoning summaries; if set to false, model_reasoning_effort will not take effect even if configured model_supports_reasoning_summaries = true model_reasoning_summary = "none" model_context_window = 1048576 web_search = "disabled" [model_providers.mimo] name = "mimo" base_url = "https://token-plan-cn.xiaomimimo.com/v1" experimental_bearer_token = "tp-your-api-key-here" wire_api = "responses" ``` ## Use Codex CLI {/* feishu-style:text-align:left */} After completing the above configuration, open a new terminal and run the following command to start Codex. ```bash codex ``` {/* feishu-style:text-align:left */} After configuring the model source data file, you can switch models via the `/model` command in Codex. ## Use Codex IDE Plugin {/* feishu-style:text-align:left */} Codex provides a VS Code plugin. Search for "Codex" in the VS Code Extension Marketplace to install it. The plugin automatically reuses the existing local Codex configuration. If you have never used the Codex command-line tool, please follow the specifications in the "Edit Configuration File" section to complete the configuration. 图片 ## Use Codex Desktop {/* feishu-style:text-align:left */} The Codex desktop client will reuse the existing configurations in your local Codex. Please refer to the "Editing Configuration File" section to complete the configuration before use. 图片 ## FAQ ### Codex throws an error saying "custom tools require MiMo freeform Responses lite mode."? {/* feishu-style:text-align:left */} The custom tool is not supported for use in **non-lite** mode. If you encounter this error, please configure the model metadata file `model-catalogs.json` in accordance with the "Edit Configuration File" section, and then restart Codex. --- DOCUMENT: OpenClaw Configuration --- URL: https://mimo.mi.com/static/docs/tokenplan/integration/openclaw.md # OpenClaw Configuration {/* feishu-style:text-align:left */} **Pay-as-you-go MiMo API** and **Token Plan** are both supported for use in OpenClaw. Please refer to this article for configuration and usage. ## Preparatory Work ### Obtain Credentials {/* feishu-style:text-align:left */} Supports two usage methods, but the corresponding credential acquisition methods are different:
Usage Description Acquisition Method (BASE_URL and API Key below are both examples)
Pay-as-you-go API calls Charged based on actual usage, suitable for light use
  • BASE_URL
    • OpenAI Compatibility Protocol:`https://api.xiaomimimo.com/v1`
  • API Key
    • Format:`sk-xxxxx`

Go to [API Keys](https://platform.xiaomimimo.com/#/console/api-keys) to create an API Key
Token Plan Fixed subscription fee, with limited calls based on the package
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/v1`
  • API Key
    • Format (Individual): `tp-xxxxx`
    • Format (Team): `ttp-xxxxx`

After successful subscription, go to [ Token Plan ](https://platform.xiaomimimo.com/#/console/plan-manage) to obtain the exclusive Base URL and API Key
When OpenClaw uses MiMo under the Anthropic protocol, due to the absence of `reasoning_content` in the assistant containing tool calls, the API will return a 400 error. For details, see [Multi-turn Conversation Pass-through Requirements](https://mimo.mi.com/docs/en-US/quick-start/usage-guide/text-generation/deep-thinking#:~:text=.-,Multi%2Dturn%20Conversation%20Pass%2Dthrough%20Requirements,-When%20deep%20thinking).
## Install OpenClaw {/* feishu-style:text-align:left */} Precondition:[ Node.js 22 or later](https://nodejs.org/en/download/) {/* feishu-style:text-align:left */} macOS/Linux: ```bash curl -fsSL https://openclaw.ai/install.sh | bash ``` {/* feishu-style:text-align:left */} Windows (PowerShell): ```bash iwr -useb https://openclaw.ai/install.ps1 | iex ``` 图片 ## Configure and use MiMo model
**Tips** - **OpenClaw supports the preconfigured settings of the MiMo pay-as-you-go API, which can be configured via Method 1 Interactive Configuration Wizard.** - **OpenClaw has not yet added the MiMo Token Plan preset configuration, so you need to manually modify the configuration file through Method 2.**
### Method 1: Interactive Configuration Wizard {/* feishu-style:text-align:left */} After the installation is complete, the configuration process will automatically start. You can also run the following command to start the configuration: ```bash openclaw onboard --install-daemon ``` {/* feishu-style:text-align:left */} **1. Configure the provider** 图片 图片 - I understand this is personal-by-default and shared/multi-user use requires lock-down. Continue? ➡️ Yes - Set Mode ➡️ QuickStart - Configuration Processing ➡️ View and Update - Model/auth provider ➡️ Xiaomi {/* feishu-style:text-align:left */} **2. Configure the model and API Key** 图片 {/* feishu-style:text-align:left */} Enter the API Key of the MiMo Open Platform, browse all models, and select the latest v2.5 series models. {/* feishu-style:text-align:left */} **3.** **Continue to complete the subsequent configuration** - Select channels, select search providers, configure skills, etc. - Complete Setup {/* feishu-style:text-align:left */} **4. Test Robot** - How do you want to hatch your bot? ➡️ You can chat with the bot in TUI/Web UI - TUI: Enter `openclaw tui`, and if the conversation is successful, it indicates successful configuration 图片 - Web UI: Access the Web UI by opening the `Web UI (with token)` link displayed in the terminal 图片 图片 ### Method 2: Modify the Configuration File {/* feishu-style:text-align:left */} Copy the following content in full to the configuration file`~/.openclaw/openclaw.json` (replace BASE_URL and API Key as needed in actual use):
**Token Plan only supports configuration via Method 2. When using Token Plan, you need to delete the** `"auth"` **field in the configuration file, and you need to add a provider to distinguish it from the pre-set MiMo gateway.**
{/* feishu-style:text-align:left */} **Token Plan** **Configuration Example:** {/* feishu-style:text-align:left */} **Delete the** `"auth"` **field** ```json "auth": { "profiles": { "xiaomi:default": { "provider": "xiaomi", "mode": "api_key" } } } ``` {/* feishu-style:text-align:left */} Add a new provider under the models.provider path. Do not set the provider name to `xiaomi`, to distinguish it from the pre-set MiMo gateway. For example, set it to `xiaomi-coding` {/* feishu-style:text-align:left */} The corresponding default agent configuration also needs to add the corresponding model, with the format ` provider name/model name `, for example ` xiaomi-coding/mimo-v2.6-pro ` ```json { "models": { "mode": "merge", "providers": { "xiaomi-coding": { "baseUrl": "BASE_URL", "apiKey": "API_KEY", "api": "openai-completions", "models": [ { "id": "mimo-v2.6-pro", "name": "mimo-v2.6-pro", "reasoning": true, "input": [ "text", "image" ], "contextWindow": 1048576, "maxTokens": 131072 }, { "id": "mimo-v2.6-flash", "name": "mimo-v2.6-flash", "reasoning": true, "input": [ "text", "image" ], "contextWindow": 1048576, "maxTokens": 131072 } ] } } }, "agents": { "defaults": { "model": { "primary": "xiaomi-coding/mimo-v2.6-pro" }, "models": { "xiaomi-coding/mimo-v2.6-flash": {}, "xiaomi-coding/mimo-v2.6-pro": {} } } } } ``` {/* feishu-style:text-align:left */} **Example of Pay-as-you-go API Configuration** ```json { "auth": { "profiles": { "xiaomi:default": { "provider": "xiaomi", "mode": "api_key" } } }, "models": { "mode": "merge", "providers": { "xiaomi": { "baseUrl": "BASE_URL", "apiKey": "API_KEY", "api": "openai-completions", "models": [ { "id": "mimo-v2.6-pro", "name": "mimo-v2.6-pro", "reasoning": true, "input": [ "text", "image" ], "contextWindow": 1048576, "maxTokens": 131072 }, { "id": "mimo-v2.6-flash", "name": "mimo-v2.6-flash", "reasoning": true, "input": [ "text", "image" ], "contextWindow": 1048576, "maxTokens": 131072 } ] } } }, "agents": { "defaults": { "model": { "primary": "xiaomi/mimo-v2.6-pro" }, "models": { "xiaomi/mimo-v2.6-flash": {}, "xiaomi/mimo-v2.6-pro": {} } } } } ``` ## Connect to More Channels {/* feishu-style:text-align:left */} OpenClaw provides more channels for you to interact with the robot, such as Web UI, Discord, Feishu, etc. You can refer to the official documentation to set up these channels:[Chat Channels - OpenClaw](https://docs.openclaw.ai/channels). --- DOCUMENT: Hermes Agent Configuration --- URL: https://mimo.mi.com/static/docs/tokenplan/integration/hermes-agent.md # Hermes Agent Configuration {/* feishu-style:text-align:left */} **Pay-as-you-go MiMo API** and **Token Plan** both support Hermes Agent. Refer to this guide for configuration and usage. ## Prerequisites ### Obtain Credentials {/* feishu-style:text-align:left */} Supports two usage methods, but the corresponding credential acquisition methods are different:
Usage Method Description Acquisition Method (BASE_URL and API Key below are examples)
Pay-as-you-go MiMo API Charged based on actual usage, suitable for light use
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://api.xiaomimimo.com/v1`
  • API Key
    • Format: `sk-xxxxx`

Go to [API Keys](https://platform.xiaomimimo.com/#/console/api-keys) to create an API Key
Token Plan Fixed subscription fee, with limited calls based on the package
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/v1`
  • API Key
    • Format (Individual): `tp-xxxxx`
    • Format (Team): `ttp-xxxxx`

After successful subscription, go to [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) to obtain the exclusive Base URL and API Key
## Install Hermes Agent {/* feishu-style:text-align:left */} Hermes Agent supports Linux, macOS, WSL2 (Windows), and more. For more information, refer to the [Hermes Agent Official Documentation](https://hermes-agent.nousresearch.com/docs/). - Linux / macOS: No additional steps required. - Windows: Refer to [Install WSL](https://learn.microsoft.com/en-us/windows/wsl/install) to install WSL2, then run the commands below in WSL2. {/* feishu-style:text-align:left */} **Installation Command:** ```bash curl -fsSL https://raw.githubusercontent.com/NousResearch/hermes-agent/main/scripts/install.sh | bash ``` {/* feishu-style:text-align:left */} **After installation, reload the terminal environment:** ```bash source ~/.bashrc # or source ~/.zshrc ``` {/* feishu-style:text-align:left */} **Verify installation (if a version number is displayed, the installation was successful):** ```bash hermes --version ``` {/* feishu-style:text-align:left */} After installation, the following interface will appear: 图片 ## Configure a Predefined Provider {/* feishu-style:text-align:left */} **1. Select Quick Setup** {/* feishu-style:text-align:left */} Choose Quick setup for initial configuration. > If not configured initially, you can re-enter the setup wizard via `hermes setup`. 图片 {/* feishu-style:text-align:left */} **2. Select Provider** `Xiaomi MiMo` 图片 {/* feishu-style:text-align:left */} **3. Fill in Configuration** {/* feishu-style:text-align:left */} Set API Key, Base URL, and default model as guided. The API Key and Base URL should be filled according to your credential type. 图片 {/* feishu-style:text-align:left */} Follow the remaining steps as needed.
**If you previously configured Pay-as-you-go MiMo API and need to switch to Token Plan:** - **Method 1:** Edit `~/.hermes/.env` file, replace `XIAOMI_API_KEY` and `XIAOMI_BASE_URL` with Token Plan credentials (open a new terminal after configuration). - **Method 2:** Use a custom provider for configuration.
{/* feishu-style:text-align:left */} **4. After configuration, the following interface will appear:** 图片 ## Configure a Custom Provider ### Configure Basic Settings {/* feishu-style:text-align:left */} Replace `BASE_URL` and `MIMO_API_KEY` in the following methods with your actual credentials. {/* feishu-style:text-align:left */} **Method 1: Quick configuration via terminal commands**
Here `model.provider` can only be set to `custom`. Custom names like `xiaomi-coding` will be invalid.
```bash hermes config set model.provider custom hermes config set model.base_url BASE_URL hermes config set model.api_key MIMO_API_KEY hermes config set model.default mimo-v2.6-pro ``` {/* feishu-style:text-align:left */} After configuration, you can view the settings in `~/.hermes/config.yaml`. {/* feishu-style:text-align:left */} **Method 2: Manually edit configuration file** {/* feishu-style:text-align:left */} Edit `~/.hermes/config.yaml` manually: ```bash model: provider: custom base_url: BASE_URL api_key: MIMO_API_KEY default: mimo-v2.6-pro ``` ### Verify Configuration {/* feishu-style:text-align:left */} After configuration, run the following command to verify: ```bash hermes doctor ``` ## Use Hermes Agent {/* feishu-style:text-align:left */} After configuration, run the following command to start: ```bash hermes # Classic CLI mode hermes --tui # Modern TUI mode ``` --- DOCUMENT: Kilo Code Configuration --- URL: https://mimo.mi.com/static/docs/tokenplan/integration/kilocode.md # Kilo Code Configuration {/* feishu-style:text-align:left */} **Pay-as-you-go MiMo API** and **Token Plan** both support Kilo Code. Refer to this guide for configuration and usage. ## Prerequisites ### Obtain Credentials {/* feishu-style:text-align:left */} Supports two usage methods, but the corresponding credential acquisition methods are different:
Usage Method Description Acquisition Method (BASE_URL and API Key below are examples)
Pay-as-you-go MiMo API Charged based on actual usage, suitable for light use
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://api.xiaomimimo.com/v1`
  • API Key
    • Format: `sk-xxxxx`

Go to [API Keys](https://platform.xiaomimimo.com/#/console/api-keys) to create an API Key
Token Plan Fixed subscription fee, with limited calls based on the package
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/v1`
  • API Key
    • Format (Individual): `tp-xxxxx`
    • Format (Team): `ttp-xxxxx`

After successful subscription, go to [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) to obtain the exclusive Base URL and API Key
When Kilo Code uses MiMo under the Anthropic protocol, since the assistant containing tool calls is missing `reasoning_content`, the API will return a 400 error. For details, see [Multi-turn Conversation Pass-through Requirements](https://mimo.mi.com/docs/en-US/quick-start/usage-guide/text-generation/deep-thinking#:~:text=.-,Multi%2Dturn%20Conversation%20Pass%2Dthrough%20Requirements,-When%20deep%20thinking).
## Use Kilo Code CLI ### Install Kilo Code CLI {/* feishu-style:text-align:left */} Node.js 18 or later is required. {/* feishu-style:text-align:left */} **Installation command:** ```bash npm install -g @kilocode/cli ``` {/* feishu-style:text-align:left */} **Verify installation (success if version number is displayed):** ```bash kilocode --version ``` ### Configure Basic Settings {/* feishu-style:text-align:left */} Edit or create the `config.json` configuration file at the following paths: - **macOS/Linux**: `~/.config/kilo/config.json` - **Windows**: `User Directory\.config\kilo\config.json` {/* feishu-style:text-align:left */} Copy the following content into the configuration file (replace `BASE_URL` and `MIMO_API_KEY` as needed): ```bash { "$schema": "https://kilo.ai/config.json", "disabled_providers": [], "provider": { "mimo": { "name": "MiMo", "npm": "@ai-sdk/openai-compatible", "models": { "mimo-v2.6-pro": { "name": "mimo-v2.6-pro", "options": { "thinking": { "type": "enabled" } } } }, "options": { "apiKey": "MIMO_API_KEY", "baseURL": "BASE_URL" } } }, "permission": { "bash": "allow" } } ```
For more detailed configuration information, visit the [Kilo Code CLI Official Documentation](https://kilo.ai/docs/cli).
### Use Kilo Code CLI {/* feishu-style:text-align:left */} After completing the above configuration, open a new terminal and run the following command to start Kilo Code CLI: ```bash kilocode ``` {/* feishu-style:text-align:left */} Once started, enter `/models` to switch models, and you can use MiMo models in Kilo Code CLI. ## Use Kilo Code IDE Plugin ### Install Plugin {/* feishu-style:text-align:left */} Search for and install the **Kilo Code** plugin in the VS Code Extensions marketplace. 图片 ### Configure a Predefined Provider (Recommended) {/* feishu-style:text-align:left */} Click Providers --> Show more providers, search for `Xiaomi`, select the corresponding Provider, and fill in the API Key.
When using the **Xiaomi Token Plan**, you need to select the Provider corresponding to the Base URL displayed on the [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) page. - `https://token-plan-cn.xiaomimimo.com/*`: Xiaomi Token Plan (China) - `https://token-plan-sgp.xiaomimimo.com/*`: Xiaomi Token Plan (Singapore) - `https://token-plan-ams.xiaomimimo.com/*`: Xiaomi Token Plan (Europe)
图片 ### Configure a Custom Provider {/* feishu-style:text-align:left */} Fill in the relevant information according to the following configuration. {/* feishu-style:text-align:left */} **1.** **Select Custom Provider** 图片 {/* feishu-style:text-align:left */} **2.** **Fill in configuration details** - **Provider ID** and **Display name**: Fill in as needed - **Base URL**: Enter the BASE_URL obtained from your usage method - **API Key**: Enter the API Key obtained from your usage method - **Models**: Add as needed, e.g. `mimo-v2.6-pro` 图片 {/* feishu-style:text-align:left */} Other unmentioned parameters can be adjusted as needed. ### Use Kilo Code Plugin {/* feishu-style:text-align:left */} After successful configuration, switch to the configured model and enter your requirements in the input box to start using. 图片 ## FAQ ### When verifying installation on Windows, I encounter the following error. How to resolve? > It seems that your package manager failed to install the right version of the Kilo CLI for your platform. You can try manually installing "@kilocode/cli-windows-x64" or "@kilocode/cli-windows-x64-baseline" package {/* feishu-style:text-align:left */} Run the command `npm install -g @kilocode/cli-windows-x64` as suggested to resolve the issue. --- DOCUMENT: Chatbox AI Configuration --- URL: https://mimo.mi.com/static/docs/tokenplan/integration/chatbox.md # Chatbox AI Configuration {/* feishu-style:text-align:left */} **Pay-as-you-go MiMo API** and **Token Plan** both support Chatbox AI, and you may refer to this article for configuration and usage instructions. ## Preliminary Work ### Obtain Credentials {/* feishu-style:text-align:left */} Two usage modes are supported, but the methods for obtaining the corresponding credentials vary.
Usage Description Acquisition Method (the following BASE_URL and API Key are for reference only)
Pay-as-you-go API calls Pay-as-you-go pricing, ideal for light usage
  • BASE_URL
    • OpenAI Compatible Protocol: `https://api.xiaomimimo.com/v1`
  • API Key
    • Format: `sk-xxxxx`

Go to [API Keys ](https://platform.xiaomimimo.com/#/console/api-keys)to create an API Key
Token Plan Fixed subscription fee, with call volume limited by the selected plan
  • BASE_URL
    • OpenAI Compatible Protocol: `https://token-plan-cn.xiaomimimo.com/v1`
  • API Key
    • Format (Individual): `tp-xxxxx`
    • Format (Team): `ttp-xxxxx`

After successful subscription, go to [Token Plan ](https://platform.xiaomimimo.com/#/console/plan-manage)to get your exclusive Base URL and API Key
## Install Chatbox AI {/* feishu-style:text-align:left */} Chatbox AI is an AI Client application and intelligent assistant that supports multi-model conversations. - Official Website Download: https://chatboxai.app/zh#download - Github: https://github.com/chatboxai/chatbox ## Configure Basic Information {/* feishu-style:text-align:left */} 1. Go to "Settings" → "Model Providers" → "Add". 图片 {/* feishu-style:text-align:left */} 2. Search for and select "Xiaomi MiMo". 图片 图片 {/* feishu-style:text-align:left */} 3. Fill in the API key: when using the pay-as-you-go API, only enter the API Key and keep the API host unchanged; when using the Token Plan, replace both the API Key and the API host with the exclusive values displayed on the package Console. 图片 {/* feishu-style:text-align:left */} 4. Click the "Check" button on the right side of the API Key. After the test succeeds, select a model from the model list to start a conversation or enable the working mode. 图片 --- DOCUMENT: Cherry Studio Configuration --- URL: https://mimo.mi.com/static/docs/tokenplan/integration/cherrystudio.md # Cherry Studio Configuration {/* feishu-style:text-align:left */} **Pay-as-you-go MiMo API** and **Token Plan** both support Cherry Studio. Refer to this guide for configuration and usage. ## Prerequisites ### Obtain Credentials {/* feishu-style:text-align:left */} Supports two usage methods, but the corresponding credential acquisition methods are different:
Usage Method Description Acquisition Method (BASE_URL and API Key below are examples)
Pay-as-you-go MiMo API Charged based on actual usage, suitable for light use
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://api.xiaomimimo.com/v1`
  • API Key
    • Format: `sk-xxxxx`

Go to [API Keys](https://platform.xiaomimimo.com/#/console/api-keys) to create an API Key
Token Plan Fixed subscription fee, with limited calls based on the package
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/v1`
  • API Key
    • Format (Individual): `tp-xxxxx`
    • Format (Team): `ttp-xxxxx`

After successful subscription, go to [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) to obtain the exclusive Base URL and API Key
## Install Cherry Studio {/* feishu-style:text-align:left */} Cherry Studio is a desktop AI client that supports multi-model conversations. - Official website: https://www.cherry-ai.com - Github: https://github.com/CherryHQ/cherry-studio ## Configure Basic Settings {/* feishu-style:text-align:left */} **1.** **Find provider** `Xiaomi MiMo` {/* feishu-style:text-align:left */} Click the settings icon in the upper right corner, go to the Model Services page, and search for `Xiaomi MiMo` in the search box. 图片 {/* feishu-style:text-align:left */} **2.** **Configure basic settings** {/* feishu-style:text-align:left */} **Pay-as-you-go MiMo API** {/* feishu-style:text-align:left */} Since the `Xiaomi MiMo` model service is already provided by Cherry Studio officially, you only need to provide the API Key obtained through this method. Keep the API Host unchanged. 图片 {/* feishu-style:text-align:left */} **Token Plan** {/* feishu-style:text-align:left */} After successfully subscribing to Token Plan, replace with the dedicated Token Plan API Key and API Host (BASE_URL).
Token Plan is temporarily unavailable in Agent mode.
## Use Cherry Studio {/* feishu-style:text-align:left */} Select the model you need to use from the model list, and you can have a normal conversation. 图片 ### Enable Thinking Mode (Optional) {/* feishu-style:text-align:left */} Click assistant settings and add a custom parameter: `"thinking": {"type": "enabled"}`. {/* feishu-style:text-align:left */} You can also adjust temperature, context window, and other parameters as needed. 图片 图片 --- DOCUMENT: Qwen Code Configuration --- URL: https://mimo.mi.com/static/docs/tokenplan/integration/qwencode.md # Qwen Code Configuration {/* feishu-style:text-align:left */} **Pay-as-you-go MiMo API** and **Token Plan** both support Qwen Code. Refer to this guide for configuration and usage. ## Prerequisites ### Obtain Credentials {/* feishu-style:text-align:left */} Supports two usage methods, but the corresponding credential acquisition methods are different:
Usage Method Description Acquisition Method (BASE_URL and API Key below are examples)
Pay-as-you-go MiMo API Charged based on actual usage, suitable for light use
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://api.xiaomimimo.com/v1`
    • Anthropic Compatibility Protocol: `https://api.xiaomimimo.com/anthropic`
  • API Key
    • Format: `sk-xxxxx`

Go to [API Keys](https://platform.xiaomimimo.com/#/console/api-keys) to create an API Key
Token Plan Fixed subscription fee, with limited calls based on the package
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/v1`
    • Anthropic Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/anthropic`
  • API Key
    • Format (Individual): `tp-xxxxx`
    • Format (Team): `ttp-xxxxx`

After successful subscription, go to [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) to obtain the exclusive Base URL and API Key
## Use Qwen Code CLI ### Install Qwen Code CLI {/* feishu-style:text-align:left */} **Installation commands:** - macOS/Linux ```bash bash -c "$(curl -fsSL https://qwen-code-assets.oss-cn-hangzhou.aliyuncs.com/installation/install-qwen.sh)" -s --source bailian ``` - Windows ```bash curl -fsSL -o %TEMP%\install-qwen.bat https://qwen-code-assets.oss-cn-hangzhou.aliyuncs.com/installation/install-qwen.bat && %TEMP%\install-qwen.bat --source bailian ``` {/* feishu-style:text-align:left */} **Verify the installation (a version number output indicates success):** ```bash qwen --version ``` ### Configure Settings {/* feishu-style:text-align:left */} **1.** **Select API Key -> Custom API Key to enter custom configuration** 图片 图片 {/* feishu-style:text-align:left */} **2.** **Edit the configuration file** 图片
For more detailed configuration information, visit the [Qwen Code official configuration documentation](https://qwenlm.github.io/qwen-code-docs/en/users/configuration/model-providers/).
{/* feishu-style:text-align:left */} Edit or create the `settings.json` file at the following path: - macOS/Linux: `~/.qwen/settings.json` - Windows: `User directory\.qwen\settings.json` {/* feishu-style:text-align:left */} Copy the following content into the configuration file (replace with your actual settings when using):
When configuring basic information, you need to first check if the `MIMO_API_KEY` environment variable exists. If it does, please clear it or replace the value with the API Key obtained through the corresponding usage method.
```bash { "env": { "MIMO_API_KEY": "MIMO_API_KEY" }, "modelProviders": { "openai": [ { "id": "mimo-v2.6-pro", "name": "mimo-v2.6-pro", "baseUrl": "BASE_URL", "envKey": "MIMO_API_KEY" } ] }, "security": { "auth": { "selectedType": "openai" } }, "model": { "name": "mimo-v2.6-pro" }, "$version": 3 } ``` ### Use Qwen Code CLI {/* feishu-style:text-align:left */} After completing the above configuration, open a new terminal and run the following command to start Qwen Code CLI: ```bash qwen ``` {/* feishu-style:text-align:left */} Once started, you can use MiMo models in Qwen Code CLI. 图片 ## Use the Qwen Code IDE Plugin ### Install the Plugin {/* feishu-style:text-align:left */} Search for and install the **Qwen Code Companion** plugin from the VS Code Extensions marketplace. 图片 ### Configure Settings {/* feishu-style:text-align:left */} Follow the same steps as described in the Qwen Code CLI configuration section above. ### Use the Qwen Code Plugin {/* feishu-style:text-align:left */} Click the Qwen Code icon in the top-right corner to open the dialog. 图片 {/* feishu-style:text-align:left */} Type or click `/`, then select `Switch model` to change the model. 图片 --- DOCUMENT: CodeBuddy Configuration --- URL: https://mimo.mi.com/static/docs/tokenplan/integration/codebuddy.md # CodeBuddy Configuration {/* feishu-style:text-align:left */} **Pay-as-you-go MiMo API** and **Token Plan** both support CodeBuddy. Refer to this guide for configuration and usage. ## Prerequisites ### Obtain Credentials {/* feishu-style:text-align:left */} Supports two usage methods, but the corresponding credential acquisition methods are different:
Usage Method Description Acquisition Method (BASE_URL and API Key below are examples)
Pay-as-you-go MiMo API Charged based on actual usage, suitable for light use
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://api.xiaomimimo.com/v1`
  • API Key
    • Format: `sk-xxxxx`

Go to [API Keys](https://platform.xiaomimimo.com/#/console/api-keys) to create an API Key
Token Plan Fixed subscription fee, with limited calls based on the package
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/v1`
  • API Key
    • Format (Individual): `tp-xxxxx`
    • Format (Team): `ttp-xxxxx`

After successful subscription, go to [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) to obtain the exclusive Base URL and API Key
## Use CodeBuddy IDE ### Install CodeBuddy {/* feishu-style:text-align:left */} Visit the [CodeBuddy website](https://www.codebuddy.ai/home) to download and install the IDE, which supports major operating systems (Windows, macOS). ### Configure MiMo Model {/* feishu-style:text-align:left */} **1. Configure Custom Model** {/* feishu-style:text-align:left */} Create or modify the configuration file `models.json` to add custom models. Example configuration: - **macOS:** `~/.codebuddy/models.json` - **Windows:** `User Directory\.codebuddy\models.json` {/* feishu-style:text-align:left */} `BASE_URL` and `MIMO_API_KEY` should be modified according to your credential acquisition method. ```json { "models": [ { "id": "mimo-v2.6-pro", "name": "mimo-v2.6-pro", "vendor": "MiMo", "apiKey": "MIMO_API_KEY", "url": "BASE_URL/chat/completions", "supportsToolCall": true, "supportsImages": true }, { "id": "mimo-v2.6-flash", "name": "mimo-v2.6-flash", "vendor": "MiMo", "apiKey": "MIMO_API_KEY", "url": "BASE_URL/chat/completions", "supportsToolCall": true, "supportsImages": true } ] } ``` {/* feishu-style:text-align:left */} **2. View and Switch Models** {/* feishu-style:text-align:left */} After configuration, turn off `Auto mode` and open the model list to see the configured MiMo models. 图片 ### Use MiMo Model {/* feishu-style:text-align:left */} Select the configured model to start conversations, coding, and other operations. 图片 ## Use CodeBuddy IDE Plugin ### Install Plugin {/* feishu-style:text-align:left */} Search for `Tencent Cloud CodeBuddy` in the VS Code extension marketplace and install the plugin. 图片 ### Configure MiMo Model {/* feishu-style:text-align:left */} Refer to the `models.json` configuration file in the "Use CodeBuddy IDE" section. If previously configured, it will be automatically loaded. 图片 ## Use CodeBuddy CLI ### Install CodeBuddy CLI {/* feishu-style:text-align:left */} **Install via npm (requires Node.js 18.20 or newer):** ```bash npm install -g @tencent-ai/codebuddy-code ``` {/* feishu-style:text-align:left */} **Verify installation (if a version number is displayed, the installation was successful):** ```bash codebuddy --version ``` ### Configure MiMo Model
The `BASE_URL` and `API Key` for **Pay-as-you-go MiMo API** and **Token Plan** are different. Please configure accordingly.
{/* feishu-style:text-align:left */} Refer to the `models.json` configuration file in the "Use CodeBuddy IDE" section. If previously configured, it will be automatically loaded. ### Use CodeBuddy CLI {/* feishu-style:text-align:left */} After configuration, navigate to your project directory and run: ```bash codebuddy ``` {/* feishu-style:text-align:left */} After startup, use `/model` to view or switch models, and `/status` to check the current model. ## FAQ ### Model not appearing in the dropdown after configuration? - Check if the JSON syntax is correct - If the `availableModels` field is configured, ensure the model id is included --- DOCUMENT: Cline Configuration --- URL: https://mimo.mi.com/static/docs/tokenplan/integration/cline.md # Cline Configuration {/* feishu-style:text-align:left */} **Pay-as-you-go MiMo API** and **Token Plan** both support Cline. Refer to this guide for configuration and usage. ## Prerequisites ### Obtain Credentials {/* feishu-style:text-align:left */} Supports two usage methods, but the corresponding credential acquisition methods are different:
Usage Method Description Acquisition Method (BASE_URL and API Key below are examples)
Pay-as-you-go MiMo API Charged based on actual usage, suitable for light use
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://api.xiaomimimo.com/v1`
  • API Key
    • Format: `sk-xxxxx`

Go to [API Keys](https://platform.xiaomimimo.com/#/console/api-keys) to create an API Key
Token Plan Fixed subscription fee, with limited calls based on the package
  • BASE_URL
    • OpenAI Compatibility Protocol: `https://token-plan-cn.xiaomimimo.com/v1`
  • API Key
    • Format (Individual): `tp-xxxxx`
    • Format (Team): `ttp-xxxxx`

After successful subscription, go to [Token Plan](https://platform.xiaomimimo.com/#/console/plan-manage) to obtain the exclusive Base URL and API Key
## Use Cline CLI ### Install Cline CLI {/* feishu-style:text-align:left */} **Prerequisites:** Node.js 20 or later is required (Node.js 22 recommended). {/* feishu-style:text-align:left */} **Installation command:** ```bash npm install -g cline ``` {/* feishu-style:text-align:left */} **Verify installation (if a version number is displayed, the installation was successful):** ```bash cline --version ``` ### Configure Basic Settings {/* feishu-style:text-align:left */} Cline CLI uses the `cline auth` command to configure API providers. Run the following command to configure MiMo model: ```bash cline auth -p openai -k MIMO_API_KEY -b BASE_URL -m mimo-v2.6-pro ``` {/* feishu-style:text-align:left */} Parameter descriptions: - `-p openai`: Select OpenAI-compatible provider - `-k`: Enter the API Key obtained from the corresponding usage method - `-b`: Enter the BASE_URL obtained from the corresponding usage method - `-m`: Enter the model ID, e.g. `mimo-v2.6-pro`
For more detailed configuration information, visit the [Cline CLI Official Documentation](https://docs.cline.bot/cline-cli/cli-reference).
{/* feishu-style:text-align:left */} You can also configure via the interactive wizard by running `cline auth` and following the prompts. ### Use Cline CLI {/* feishu-style:text-align:left */} After completing the configuration, open a new terminal and run the following command to start Cline CLI. > If you prefer the classic terminal interface, select `Exit` and run `cline --tui` to return to the familiar command-line environment. ```bash cline ``` {/* feishu-style:text-align:left */} After starting, you can use MiMo models in Cline CLI. ## Use Cline IDE Plugin ### Install Plugin {/* feishu-style:text-align:left */} Search for and install the **Cline** plugin in the VS Code Extensions marketplace. 图片 ### Configure Basic Settings {/* feishu-style:text-align:left */} Open the Cline plugin in VS Code and fill in the following configuration: - Required settings: - **API Provider**: Select `OpenAI Compatible` - **Base URL**: Fill in the BASE_URL obtained through the corresponding usage method - **API Key**: API Key obtained from the corresponding usage method - **Model ID**: Enter the model name `mimo-v2.6-pro` - Optional settings: - Set **Context Window Size** to `1048576` - Set **Temperature** to `1.0`, adjustable based on task requirements {/* feishu-style:text-align:left */} Other parameters not mentioned can be adjusted as needed. ### Use Cline Plugin {/* feishu-style:text-align:left */} After successful configuration, enter your request in the input box, for example to generate code: 图片 --- DOCUMENT: MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement --- URL: https://mimo.mi.com/static/docs/news/latest/v2-6.md # MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement {/* feishu-style:text-align:left */} Today, we are officially releasing and open-sourcing the Xiaomi MiMo-V2.6 series. This marks a key step in our exploration of the RSI (recursive self-improvement) path: building on verifiable complex tasks, we scale up reinforcement learning (RL) computing power to enable models to continuously expand the boundaries of intelligence through ongoing exploration and feedback. {/* feishu-style:text-align:left */} **Where the path is flat and close, travelers are many; where it is rugged and distant, few reach the end.** In an era where intelligence can be easily replicated, we choose to channel computing power into real-world environments, letting models learn through trial and error in iterative feedback loops. This path is slower, and far less visible. The 6 days of Live RL training for MiMo-V2.6 mark a public trek we’ve taken along this road; behind these 6 days lie half a year of foundational research accumulation and engineering trial and error. {/* feishu-style:text-align:left */} The MiMo-V2.6 series comprises two native fully multimodal models, namely Pro and Flash. Benefiting from the expanded RL computing power, **MiMo-V2.6-Pro scores 46 points in the Artificial Analysis Intelligence Index (AA Composite Intelligence Index), surpassing Kimi K3 and Qwen3.8 Max to become the most powerful open-source model available**;However, there is still a gap when compared with the strongest closed-source models Claude Fable 5.1 and GPT-6 Astra. 图片 {/* feishu-style:text-align:left */} The MiMo-V2.6 series adopts the same API pricing as the V2.5 series. With intelligent performance upgraded while price remains unchanged, the Pareto frontier of "intelligence vs. cost" has thus been pushed outward once again. The MiMo-V2.6-Pro has set a new cost-performance record for domestic large language models: at the same intelligence level, its price is only **1/20 to 1/60** that of overseas models. 图片 ### Scale RL on a large scale and fully open-source it {/* feishu-style:text-align:left */} **During the RL training phase, MiMo-V2.6 is likely one of the domestic open-source models that has been allocated the largest amount of computing power to date**. After large-scale, multi-task reinforcement learning training, MiMo-V2.6-Pro has achieved performance on most Agent Benchmarks that is on par with Claude Opus5 and GPT-5.6 Sol, while MiMo-V2.6-Flash has comprehensively outperformed MiMo-V2.5-Pro. 图片 {/* feishu-style:text-align:left */} Throughout the entire process, we overcame fundamental research and engineering challenges in RL training, and documented the official experimental journey via live sharing. In less than 6 days, MiMo-V2.6-Flash and MiMo-V2.6-Pro completed 30 steps each with a cumulative total of approximately 750,000 trajectories, at training costs of around 850,000 and 2.62 million US dollars respectively;The average pass rate of training tasks has been relatively improved by 25% and 12% respectively, and the out-of-sample long-range software engineering evaluation benchmark **DeepSWE v1.1 has been improved by approximately 17 points (from 48.8 to 65.7) and approximately 14 points (from 58.4 to 72.6) respectively**, which reflects the high sample efficiency, continuous improvement capability and out-of-sample generalization capability of RL. 图片 {/* feishu-style:text-align:left */} This training mainly expands RL computing power from three dimensions: 1. **Larger Batch Size and Higher Throughput**: By combining a large batch size and a fully asynchronous architecture, each update uses **1,568 samples**, supports training with **1M context length**, and the number of tokens per training step reaches **3.5~3.7B**. 1. **More Tasks and Complex Environments**: Build a multi-task training system covering fields such as Code, General, Visual, and Cyber, and integrate multiple Harnesses to facilitate **the collaborative improvement of different capability dimensions**. 1. **Greater Grader computing power**: through relative comparison within the Group, it provides more accurate and diverse reward signals for Long-Horizon RL tasks, forms a closed loop for model self-improvement, and **guides the model to complete tasks with shorter paths and fewer Tokens**. {/* feishu-style:text-align:left */} As the training scale expands, we freeze the MoE Router to suppress expert load drift, and establish a defense against Reward Hacking that covers reward design, adversarial evaluation, anomaly detection and cross-verification of validators, so as to improve training stability and reward reliability. {/* feishu-style:text-align:left */} To support multi-agent reinforcement learning for large-scale hybrid tasks, we have designed a unified trajectory representation and penalty mechanism to refine learning signals, support high-concurrency interactions of various multi-agent frameworks, decouple the control plane and data plane to enable the migration of massive trajectories, stabilize the sample ratio of each task in hybrid batches, and optimize the efficiency of training and inference engines as well as their consistency. {/* feishu-style:text-align:left */} We have open-sourced the aforementioned technical achievements and supporting resources, including the complete technical report, training environment and RL code, to help more researchers reproduce and verify relevant results, and jointly explore more possibilities of large-scale RL and model self-improvement. ### From Vibe Coding to Vibe World {/* feishu-style:text-align:left */} MiMo-V2.6 integrates **3D spatial reasoning, multimodal perception and computer user operation (CUA) capabilities,** further expanding the boundaries of what programming can achieve; it can extend natural language-driven programming tasks into the "Vibe World" oriented towards interactive world construction. {/* feishu-style:text-align:left */} **3D open-world game** {/* feishu-style:text-align:left */} In game development, after a user inputs images, videos, or text, MiMo-V2.6 breaks down the requirements into multiple tasks, which are completed through multi-agent collaboration: 3D game scene construction, interactive logic programming, and visual verification. It then makes continuous corrections based on the rendering results, and ultimately generates a runnable interactive world that aligns with the user's intent.