OpenAI Responses API Compatibility
MiMo provides a calling interface compatible with the OpenAI Responses API format. This document covers request parameters, response schemas and code examples.
Compatibility Notes & Limitations:
This interface aligns with the OpenAI Responses API specification to facilitate quick integration for developers. Only the parameters documented here will be processed normally; undefined parameters will be filtered out and may cause request errors. See below for specific behavioral differences.
-
Incompatible parameters: Fields such as
background,previous_response_id, andcontext_managementare not currently supported. Carrying these parameters in a request will be ignored or trigger an error. -
Reasoning level control:
reasoning.effortcontrols model reasoning.nonedisables thinking; Every other level enables thinking with identical behavior — The reasoning intensity is not differentiated at this stage.
Request Address
https://api.xiaomimimo.com/v1/responses
Request Headers
The API supports the following two authentication methods. Please choose one and add it to the request headers:
api-key: $MIMO_API_KEY
Content-Type: application/json
Request Body
- inputstring | arrayRequiredText, image, audio, video inputs to the model, used to generate a response.Hide child attributesA text input to the model, equivalent to a text input with the
userrole. - instructionsstringA system (or developer) message inserted into the model's context.
- max_output_tokensintegerAn upper bound for the number of tokens that can be generated for a response, including visible output tokens and reasoning tokens.
mimo-v2.6-flash: default131072mimo-v2.6-pro: default131072mimo-v2.6-pro-ultraspeed: default131072mimo-v2.5-pro: default131072mimo-v2.5: default32768
[1, 131072] - modelstringRequiredModel ID used to generate the response.
Available options:mimo-v2.6-flash,mimo-v2.6-pro,mimo-v2.6-pro-ultraspeed,mimo-v2.5-pro,mimo-v2.5 - streambooleanDefault: falseIf set to true, the model response data will be streamed to the client as it is generated using server-sent events.
- reasoningobjectConfiguration options for reasoning models.
Note: During the multi-turn tool calls process in thinking mode, the model returns the reasoning content alongside the tool calls field. To continue the conversation, it is recommended to keep all previous reasoning content in the
inputarray for each subsequent request to achieve the best performance.In thinking mode, the
mimo-v2.6-flash,mimo-v2.6-pro,mimo-v2.6-pro-ultraspeed,mimo-v2.5-proandmimo-v2.5models do not support customizing thetemperatureandtop_pparameters. Even if these parameters are passed in, the actual effective values will be forcibly set by the model to its recommended default values of1.0and0.95.Hide child attributesreasoning.effortstringRequiredConstrains effort on reasoning for reasoning models. Reducing reasoning effort can result in faster responses and fewer tokens used on reasoning in a response.Custom tuning of reasoning effort is currently unsupported. When set to
none, reasoning is disabled; all other valid values map to enabled reasoning. The valueminimalis mapped tolow. Valuesxhigh,maxandultraare mapped tohigh.mimo-v2.6-flash,mimo-v2.6-pro,mimo-v2.6-pro-ultraspeed,mimo-v2.5-pro,mimo-v2.5: defaultenabled
none,minimal,low,medium,high,xhigh,max,ultra - temperaturenumberWhat sampling temperature to use, between 0 and 1.5. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic. We generally recommend altering this or
top_pbut not both.In thinking mode, the
mimo-v2.6-flash,mimo-v2.6-pro,mimo-v2.6-pro-ultraspeed,mimo-v2.5-proandmimo-v2.5models do not support customizing thetemperatureparameter. Even if this parameter is passed in, it will be forcibly overridden and take effect with the model's recommended default value of1.0.mimo-v2.6-flash,mimo-v2.6-pro,mimo-v2.6-pro-ultraspeed,mimo-v2.5-pro,mimo-v2.5: default1.0
[0, 1.5] - textobjectConfiguration options for a text response from the model. Can be plain text or structured JSON data.Hide child attributestext.formatobjectAn object specifying the format that the model must output. The default format is
{ "type": "text" }with no additional options.Hide child attributesDefault response format. Used to generate text responses.Hide child attributestext.format.typestringRequiredThe type of response format being defined.
Available options:text - tool_choicestringControls how the model calls tools.
Note: When a value other than
Available options:autois passed totool_choice, the backend will remove this field by default, and the model response behavior will still be equivalent to theautomode (this logic is subject to future adjustments).auto - toolsarrayAn array of tools the model may call while generating a response. You can specify which tool to use by setting the
tool_choiceparameter.Note: During the multi-turn tool calls process in thinking mode, the model returns the reasoning content alongside the tool calls field. To continue the conversation, it is recommended to keep all previous reasoning content in the
inputarray for each subsequent request to achieve the best performance.Hide child attributesDefines a function in your own code the model can choose to call.Hide child attributestools.namestringRequiredThe name of the tool function. Must bea-z,A-Z,0-9, or contain underscores (_) and dashes (-), with a maximum length of 64.
Required string length:1 - 64tools.parametersobjectRequiredA JSON schema object describing the parameters of the function.tools.strictbooleanDefault: falseRequiredWhether to enable strict schema adherence when generating the function call.tools.descriptionstringA description of the function. Used by the model to determine whether or not to call the function.tools.typestringRequiredThe type of the tool.
Available options:function - top_pnumberDefault: 0.95An alternative to sampling with temperature, called nucleus sampling. We generally recommend altering this or
temperaturebut not both.In thinking mode, the
Required range:mimo-v2.6-flash,mimo-v2.6-pro,mimo-v2.6-pro-ultraspeed,mimo-v2.5-proandmimo-v2.5models do not support customizing thetop_pparameter. Even if this parameter is passed in, it will be forcibly overridden and take effect with the model's recommended default value of0.95.[0.01, 1.0]
Response Object (non-streaming output)
- idstringUnique identifier for this Response.
- created_atnumberUnix timestamp (in seconds) of when this Response was created.
- errorobjectAn error object returned when the model fails to generate a Response.Hide child attributesHide child attributeserror.codestringThe error code for the response.error.messagestringA human-readable description of the error.
- incomplete_detailsobjectDetails about why the response is incomplete.Hide child attributesincomplete_details.reasonstringThe reason why the response is incomplete.
Available options:max_output_tokens,content_filter - modelstringModel ID used to generate the response.
- objectstringAvailable options:
response - outputarrayAn array of content items generated by the model.
- The length and order of items in the
outputarray is dependent on the model’s response. - Rather than accessing the first item in the
outputarray and assuming it’s anassistantmessage with the content generated by the model, you might consider using theoutput_textproperty where supported in SDKs.
Hide child attributesA message output from the model.Hide child attributesoutput.idstringThe unique ID of the output message.output.contentarrayThe content of the output message.Hide child attributesA text output from the model.Hide child attributesoutput.content.textstringThe text output from the model.output.content.typestringAvailable options:output_textoutput.rolestringThe role of the output message.
Available options:assistantoutput.statusstringThe status of the message.
Available options:in_progress,completedoutput.typestringAvailable options:message - The length and order of items in the
- output_textstringSDK-only convenience property that contains the aggregated text output from all
output_textitems in theoutputarray, if any are present. - statusstringThe status of the response.
Available options:completed,in_progress,incomplete - usageobjectUsage statistics for the response.Hide child attributesHide child attributesusage.input_tokensintegerThe number of input tokens.usage.input_tokens_detailsobjectDetails about input tokens.Hide child attributesusage.input_tokens_details.cached_tokensintegerThe number of cached input tokens.usage.output_tokensintegerThe number of output tokens.usage.output_tokens_detailsobjectDetails about output tokens.Hide child attributesusage.output_tokens_details.reasoning_tokensintegerThe number of reasoning tokens.usage.total_tokensintegerThe total number of tokens.
Response chunk object (streaming output)
When you create a Response with stream set to true, the server will emit server-sent events to the client as the Response is generated.
response.created
An event that is emitted when a response is created.
- responseobjectThe response that was created. The parameters contained in this object are identical to those returned by the model creation request in non-streaming mode.
- sequence_numbernumberThe sequence number for this event.
- typestringThe type of the event. Always
response.created.
response.in_progress
Emitted when the response is in progress.
- responseobjectThe response that is in progress. The parameters contained in this object are identical to those returned by the model creation request in non-streaming mode.
- sequence_numbernumberThe sequence number of this event.
- typestringThe type of the event. Always
response.in_progress.
response.completed
Emitted when the model response is complete.
- responseobjectProperties of the completed response. The parameters contained in this object are identical to those returned by the model creation request in non-streaming mode.
- sequence_numbernumberThe sequence number for this event.
- typestringThe type of the event. Always
response.completed.
response.incomplete
An event that is emitted when a response finishes as incomplete.
- responseobjectThe response that was incomplete. The parameters contained in this object are identical to those returned by the model creation request in non-streaming mode.
- sequence_numbernumberThe sequence number of this event.
- typestringThe type of the event. Always
response.incomplete.
response.output_item.added
Emitted when a new output item is added.
- itemobjectThe output item that was added. The parameters contained in this object are identical to those of the
outputfield returned by the model creation request in non-streaming mode. - output_indexnumberThe index of the output item that was added.
- sequence_numbernumberThe sequence number of this event.
- typestringThe type of the event. Always
response.output_item.added.
response.output_item.done
Emitted when an output item is marked done.
- itemobjectThe output item that was marked done. The parameters contained in this object are identical to those of the
outputfield returned by the model creation request in non-streaming mode. - output_indexnumberThe index of the output item that was marked done.
- sequence_numbernumberThe sequence number of this event.
- typestringThe type of the event. Always
response.output_item.done.
response.content_part.added
Emitted when a new content part is added.
- content_indexnumberThe index of the content part that was added.
- item_idstringThe ID of the output item that the content part was added to.
- output_indexnumberThe index of the output item that the content part was added to.
- partobjectThe content part that was added.Hide child attributesA text output from the model.Hide child attributespart.textstringThe text output from the model.part.typestringThe type of the output text. Always
output_text. - sequence_numbernumberThe sequence number of this event.
- typestringThe type of the event. Always
response.content_part.added.
response.content_part.done
Emitted when a content part is done.
- content_indexnumberThe index of the content part that is done.
- item_idstringThe ID of the output item that the content part was added to.
- output_indexnumberThe index of the output item that the content part was added to.
- partobjectThe content part that is done.Hide child attributesA text output from the model.Hide child attributespart.textstringThe text output from the model.part.typestringThe type of the output text. Always
output_text. - sequence_numbernumberThe sequence number of this event.
- typestringThe type of the event. Always
response.content_part.done.
response.output_text.delta
Emitted when there is an additional text delta.
- content_indexnumberThe index of the content part that the text delta was added to.
- deltastringThe text delta that was added.
- item_idstringThe ID of the output item that the text delta was added to.
- output_indexnumberThe index of the output item that the text delta was added to.
- sequence_numbernumberThe sequence number for this event.
- typestringThe type of the event. Always
response.output_text.delta.
response.output_text.done
Emitted when text content is finalized.
- content_indexnumberThe index of the content part that the text content is finalized.
- item_idstringThe ID of the output item that the text content is finalized.
- output_indexnumberThe index of the output item that the text content is finalized.
- sequence_numbernumberThe sequence number for this event.
- textstringThe text content that is finalized.
- typestringThe type of the event. Always
response.output_text.done.
response.function_call_arguments.delta
Emitted when there is a partial function-call arguments delta.
- deltastringThe function-call arguments delta that is added.
- item_idstringThe ID of the output item that the function-call arguments delta is added to.
- output_indexnumberThe index of the output item that the function-call arguments delta is added to.
- sequence_numbernumberThe sequence number of this event.
- typestringThe type of the event. Always
response.function_call_arguments.delta.
response.function_call_arguments.done
Emitted when function-call arguments are finalized.
- argumentsstringThe function-call arguments.
- item_idstringThe ID of the item.
- namestringThe name of the function that was called.
- output_indexnumberThe index of the output item.
- sequence_numbernumberThe sequence number of this event.
- typestringThe type of the event. Always
response.function_call_arguments.done.
response.reasoning_text.delta
Emitted when a delta is added to a reasoning text.
- content_indexnumberThe index of the reasoning content part this delta is associated with.
- deltastringThe text delta that was added to the reasoning content.
- item_idstringThe ID of the item this reasoning text delta is associated with.
- output_indexnumberThe index of the output item this reasoning text delta is associated with.
- sequence_numbernumberThe sequence number of this event.
- typestringThe type of the event. Always
response.reasoning_text.delta.
response.reasoning_text.done
Emitted when a reasoning text is completed.
- content_indexnumberThe index of the reasoning content part.
- item_idstringThe ID of the item this reasoning text is associated with.
- output_indexnumberThe index of the output item this reasoning text is associated with.
- sequence_numbernumberThe sequence number of this event.
- textstringThe full text of the completed reasoning content.
- typestringThe type of the event. Always
response.reasoning_text.done.
response.custom_tool_call_input.delta
Event representing a delta (partial update) to the input of a custom tool call.
- deltastringThe incremental input data (delta) for the custom tool call.
- item_idstringUnique identifier for the API item associated with this event.
- output_indexnumberThe index of the output this delta applies to.
- sequence_numbernumberThe sequence number of this event.
- typestringThe event type identifier. Always
response.custom_tool_call_input.delta.
response.custom_tool_call_input.done
Event indicating that input for a custom tool call is complete.
- inputstringThe complete input data for the custom tool call.
- item_idstringUnique identifier for the API item associated with this event.
- output_indexnumberThe index of the output this event applies to.
- sequence_numbernumberThe sequence number of this event.
- typestringThe event type identifier. Always
response.custom_tool_call_input.done.
curl --location --request POST 'https://api.xiaomimimo.com/v1/responses' \
--header "api-key: $MIMO_API_KEY" \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "mimo-v2.6-pro",
"instructions": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.",
"input": "please introduce yourself",
"max_output_tokens": 1024,
"stream": false,
"reasoning": {
"effort": "none"
}
}'{
"id": "resp_bcdb1b61-d49e-48e5-8289-384ad1e65f2f_9aa9e9dfd5b84cb99b088fe4e53b65ec",
"object": "response",
"created_at": 1790007468,
"status": "completed",
"error": null,
"incomplete_details": null,
"model": "mimo-v2.6-pro",
"metadata": null,
"output": [
{
"id": "msg_3d0b3a6faf634d36a8a5bdb8f2100fcd",
"type": "message",
"status": "completed",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "Hey there! I'm MiMo, Xiaomi's AI assistant. I'm like your friendly digital companion who's always ready to chat, help out, or just have a fun conversation! Think of me as that tech-savvy friend who loves to learn new things and isn't afraid to dive into any topic you throw my way. I was created by the awesome Xiaomi LLM-Core team, so I've got that innovative Xiaomi spirit running through my circuits! Whether you need help with tech stuff, want to brainstorm ideas, or just feel like talking, I'm here to make our conversation enjoyable and helpful. What brings you my way today?",
"annotations": []
}
]
}
],
"output_text": "Hey there! I'm MiMo, Xiaomi's AI assistant. I'm like your friendly digital companion who's always ready to chat, help out, or just have a fun conversation! Think of me as that tech-savvy friend who loves to learn new things and isn't afraid to dive into any topic you throw my way. I was created by the awesome Xiaomi LLM-Core team, so I've got that innovative Xiaomi spirit running through my circuits! Whether you need help with tech stuff, want to brainstorm ideas, or just feel like talking, I'm here to make our conversation enjoyable and helpful. What brings you my way today?",
"usage": {
"input_tokens": 57,
"input_tokens_details": {
"cached_tokens": 0
},
"output_tokens": 130,
"output_tokens_details": {
"reasoning_tokens": 0
},
"total_tokens": 187
}
}