图片理解
图片理解模型可以根据您传入的图片进行回答,支持图片 URL 和 Base64 编码两种传入方式,适用于图片描述、分类等场景。
快速开始
获取 API Key 等准备工作,请参考 首次调用API。
通过图片 URL 方式传入模型快速体验图片理解效果,示例代码如下。
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ.get("MIMO_API_KEY"),
base_url="https://api.xiaomimimo.com/v1"
)
completion = client.chat.completions.create(
model="mimo-v2.6-pro",
messages=[
{
"role": "system",
"content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024."
},
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {
"url": "https://example-files.cnbj1.mi-fds.com/example-files/image/image_example.png"
}
},
{
"type": "text",
"text": "please describe the content of the image"
}
]
}
],
max_completion_tokens=1024
)
print(completion.model_dump_json())
响应结果
{
"id": "75a0045f-a28f-4b0f-8ddd-6c90c6ef6587_90e3b2e2dd1f410c8b5526a21526bdec",
"choices": [
{
"finish_reason": "stop",
"index": 0,
"message": {
"content": "This is a serene, enchanted-looking ancient forest scene, rich with detail:\n1. **Foreground:** At the bottom left, lush, large green ferns with delicate fronds grow on a bed of moss-covered smooth river stones. A shallow, clear stream flows gently over the stones, its surface glinting brightly as it catches golden sunlight. Mossy rocks and small, leafy green shrubs line the right bank, their leaves glowing in the warm light.\n2. **Midground:** The stream winds deeper into the woods, surrounded by thick, vibrant moss and lush low greenery. Dominating the left is a massive, gnarled ancient tree with thick, exposed roots gripping the earth, its dark, twisted trunk curving outward. Other tall, mature trees with textured dark bark fill the space, creating a dense, old-growth forest feel.\n3. **Light & Atmosphere:** Warm, golden sunbeams (crepuscular rays) filter down through the forest canopy, cutting through a soft mist that hangs in the air. Tiny glowing particles, like sun dust or fireflies, drift in the sunlit space, adding a magical, dreamlike quality. The distant forest fades into soft haze, giving the scene deep, peaceful atmosphere.\n4. **Overall feel:** The image evokes a sense of quiet, untouched wilderness, peaceful and full of life, with the gentle flow of the stream and warm, dappled light creating an intimate, enchanted natural space.",
"role": "assistant",
"tool_calls": null,
"reasoning_content": "Got it, let's break down this beautiful forest scene step by step.\n\nFirst, the overall mood: it's a misty, sun-dappled ancient forest, with a sense of quiet, magic (the little light specks feel like fireflies or sun particles), peaceful and lush.\n\nForeground: At the bottom left, large, vibrant green ferns with feathery fronds sit on moss-covered, smooth rounded stones. The stream here is shallow, with clear water flowing over the stones, catching the golden sunlight. Mossy rocks line the right side, with small leafy green bushes growing on them, their leaves catching the warm light.\n\nMidground: A gentle stream winds its way deeper into the forest, its surface glinting with the sunbeams. The banks are thick with moss, ferns, and low greenery, all lush and healthy. Large, ancient trees dominate here: on the left, a massive tree with thick, gnarled, exposed roots gripping the ground, its dark trunk curving outward. Other tall trees with dark, textured bark fill the space, their trunks disappearing into the mist.\n\nBackground & lighting: The scene is backlit by warm, golden sunbeams (crepuscular rays) that filter through the tree canopy, cutting through the mist that hangs in the air, creating a soft, hazy glow. The light catches tiny particles in the air, making them look like glowing embers or fireflies, adding a magical touch. The distant forest fades into soft mist, giving the scene depth.\n\nAll elements come together to feel like a serene, untouched, enchanted forest space, full of life and soft, warm light, with the sound of the gentle stream almost imaginable."
}
}
],
"created": 1790007893,
"model": "mimo-v2.6-pro",
"object": "chat.completion",
"usage": {
"completion_tokens": 656,
"prompt_tokens": 1085,
"total_tokens": 1741,
"completion_tokens_details": {
"reasoning_tokens": 351
},
"prompt_tokens_details": {
"cached_tokens": 1024,
"image_tokens": 1024
}
}
}
支持的模型列表
当前支持 mimo-v2.6-flash,mimo-v2.6-pro,mimo-v2.6-pro-ultraspeed 和 mimo-v2.5 模型。
图片传入方式
支持的图片传入方式如下:
-
图片 URL 传入:需提供公网可访问的图片 URL 地址。
-
Base64 编码传入:将图片转换为 Base64 编码字符串后再传入。
图片 URL 传入
通过公网可访问的图片 URL 地址直接传入图片,适用于图片已存储在公网可访问环境的场景。单张图片的文件大小不能超过 50 MB。
OpenAI Chat Completions API
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ.get("MIMO_API_KEY"),
base_url="https://api.xiaomimimo.com/v1"
)
completion = client.chat.completions.create(
model="mimo-v2.6-pro",
messages=[
{
"role": "system",
"content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024."
},
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {
"url": "https://example-files.cnbj1.mi-fds.com/example-files/image/image_example.png"
}
},
{
"type": "text",
"text": "please describe the content of the image"
}
]
}
],
max_completion_tokens=1024
)
print(completion.model_dump_json())
Anthropic Messages API
import os
from anthropic import Anthropic
client = Anthropic(
api_key=os.environ.get("MIMO_API_KEY"),
base_url="https://api.xiaomimimo.com/anthropic"
)
message = client.messages.create(
model="mimo-v2.6-pro",
max_tokens=1024,
system="You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.",
messages=[
{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "url",
"url": "https://example-files.cnbj1.mi-fds.com/example-files/image/image_example.png"
}
},
{
"type": "text",
"text": "please describe the content of the image"
}
]
}
]
)
print(message.content)
Base64 编码传入
将图片文件转换为 Base64 编码字符串后传入,适用于图片无法通过公网 URL 访问的场景。转换后的 Base64 编码的字符串大小不能超过 50 MB。
下述示例中的
{MIME_TYPE}与$BASE64_IMAGE均为占位符,运行前请替换为实际有效数据。
OpenAI Chat Completions API
请在 Base64 编码前携带前缀:data:{MIME_TYPE};base64,$BASE64_IMAGE
-
{MIME_TYPE}:图像的 MIME 类型(媒体类型),用于标识图像格式,需替换为实际图像对应的 MIME 值。 -
$BASE64_IMAGE:图像文件的纯 Base64 编码字符串(不含任何前缀)。
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ.get("MIMO_API_KEY"),
base_url="https://api.xiaomimimo.com/v1"
)
completion = client.chat.completions.create(
model="mimo-v2.6-pro",
messages=[
{
"role": "system",
"content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024."
},
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {
"url": "data:{MIME_TYPE};base64,$BASE64_IMAGE"
}
},
{
"type": "text",
"text": "please describe the content of the image"
}
]
}
],
max_completion_tokens=1024
)
print(completion.model_dump_json())
Anthropic Messages API
import os
from anthropic import Anthropic
client = Anthropic(
api_key=os.environ.get("MIMO_API_KEY"),
base_url="https://api.xiaomimimo.com/anthropic"
)
message = client.messages.create(
model="mimo-v2.6-pro",
max_tokens=1024,
system="You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024.",
messages=[
{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "base64",
"media_type": "{MIME_TYPE}",
"data": "$BASE64_IMAGE"
}
},
{
"type": "text",
"text": "please describe the content of the image"
}
]
}
]
)
print(message.content)
多图输入
支持同时传入多张图像的公网 URL 或 Base64 编码字符串,模型能够解析图像内容并返回贴合图像语义的回复。
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ.get("MIMO_API_KEY"),
base_url="https://api.xiaomimimo.com/v1"
)
completion = client.chat.completions.create(
model="mimo-v2.6-pro",
messages=[
{
"role": "system",
"content": "You are MiMo, an AI assistant developed by Xiaomi. Today is date: Tuesday, December 16, 2025. Your knowledge cutoff date is December 2024."
},
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {
"url": "https://example-files.cnbj1.mi-fds.com/example-files/image/image_example.png"
}
},
{
"type": "image_url",
"image_url": {
"url": "data:{MIME_TYPE};base64,$BASE64_IMAGE"
}
},
{
"type": "text",
"text": "please describe the connections and differences between these two pictures"
}
]
}
],
max_completion_tokens=1024
)
print(completion.model_dump_json())
图片限制
-
图片格式:JPEG,PNG,GIF,WebP,BMP。
-
图像大小:
-
以 URL 方式传入时:单张图片文件大小不超过 50 MB。
-
以 Base64 编码传入时:单张图片 Base64 编码字符串大小不超过 50 MB。
-
-
图片数量:传入多张图片时,图片数量受模型上下文长度限制,所有图片和文本的总 Token 数必须小于模型的上下文长度。
注:计算图像的 Token 请参考 图片 Token 用量说明。模型上下文长度请参考 定价与限速。
图片 Token 用量及缩放规则说明
图片的计算规则较为复杂,Token 转化及缩放规则请参考以下代码。估算结果仅供参考,实际用量以 API 响应为准。
import math
from PIL import Image
PATCH_SIZE = 16
SPATIAL_MERGE_SIZE = 2
TEMPORAL_PATCH_SIZE = 2
IMAGE_MIN_PIXELS = 8192
IMAGE_MAX_PIXELS = 8388608
def calc_image_tokens(image_path: str) -> dict:
image = Image.open(image_path)
height = image.height
width = image.width
factor = PATCH_SIZE * SPATIAL_MERGE_SIZE # 32
h_bar = round(height / factor) * factor
w_bar = round(width / factor) * factor
if h_bar * w_bar > IMAGE_MAX_PIXELS:
beta = math.sqrt((height * width) / IMAGE_MAX_PIXELS)
h_bar = math.floor(height / beta / factor) * factor
w_bar = math.floor(width / beta / factor) * factor
elif h_bar * w_bar < IMAGE_MIN_PIXELS:
beta = math.sqrt(IMAGE_MIN_PIXELS / (height * width))
h_bar = math.ceil(height / beta / factor) * factor
w_bar = math.ceil(width / beta / factor) * factor
grid_t = 1
grid_h = h_bar // PATCH_SIZE
grid_w = w_bar // PATCH_SIZE
num_tokens = (grid_t * grid_h * grid_w) // (SPATIAL_MERGE_SIZE ** 2)
return num_tokens
if __name__ == "__main__":
token = calc_image_tokens(image_path="xxx/test.jpg")
print(token)
计费说明
-
计费:总费用根据输入、输入(命中缓存)和输出 Token 数计算;价格请参考 定价与限速。
- 可通过 图片 Token 用量说明 计算图片的 Token 消耗。估算结果仅供参考,实际用量以 API 响应为准。
-
查看账单:您可以在控制台的 账单明细 页面查看账单及用量。
常见问题
是否支持本地文件上传?
mimo-v2.6-flash,mimo-v2.6-pro,mimo-v2.6-pro-ultraspeed 和 mimo-v2.5 模型暂不支持图片本地文件上传。支持的上传方式请参考 图片传入方式。