MiniMax-H3 is a general-purpose omni-modal video model that understands text, image, video, and audio context together, and generates video with native stereo audio up to 15s at 2K. It supports text-to-video, first/last-frame I2V, and multimodal reference-to-video.
Key strengths
- Omni multimodal context
- Native stereo audio video
- Up to 15s / 2K output
- Reference image / video / audio control
- Strong price-performance on 2K
Use cases
- Cinematic short clips
- Product & character videos
- Multimodal reference editing
- Storyboard-to-video drafts
- Localized creative video
MiniMax's minimax/MiniMax-H3 is a high-fidelity video generation model. It supports text-to-video and image-to-video workflows with configurable duration, aspect ratio, and resolution, plus first-frame and last-frame control for guided scene composition.
Generates cinematic clips with consistent motion, camera control, and optional native audio. Billed per second of generated content.
minimax/MiniMax-H3 is fully OpenAI-compatible — drop in your existing OpenAI Python or Node SDK and switch `baseURL` to `https://api.tokenlx.ai`. TokenLX transparently routes your requests to the optimal provider endpoint while preserving streaming, function-calling, and structured-output semantics.
Performance
Compare different providers across TokenLX · All locations.
Effective Pricing
Pricing is shown by the model billing method, using per-call or per-second prices and resolution tiers.
Recent activity
Total usage per day on TokenLX (last 30 days).
Sample code & API
TokenLX normalizes requests and responses across providers. Use any OpenAI SDK or our native SDK.
# ─── 文生视频示例 ───
import requests, time
headers = {"Authorization": "Bearer sk-aihub-...", "Content-Type": "application/json"}
resp = requests.post(
"https://api.tokenlx.ai/v1/aigc/video/tasks",
headers=headers,
json={
"model": "MiniMax-H3",
"prompt": "Epic cinematic teaser: a captain stands before a starship observation window as the fleet jumps to lightspeed",
"duration": 5,
"resolution": "2K",
"aspectRatio": "16:9",
},
)
task = resp.json()
task_id = task.get("taskId") or task.get("task_id") or task.get("id")
# 轮询结果
while True:
result = requests.get(
f"https://api.tokenlx.ai/v1/aigc/video/tasks/{'{'}task_id{'}'}?model=MiniMax-H3",
headers=headers,
).json()
print("status:", result.get("status"))
if result.get("status") in ("completed", "succeeded", "done"):
print(result)
break
time.sleep(10)
# ─── 图生视频(首帧驱动)───
i2v_resp = requests.post(
"https://api.tokenlx.ai/v1/aigc/video/tasks",
headers=headers,
json={
"model": "MiniMax-H3",
"prompt": "画面中的人物缓缓转头微笑",
"referenceImageUrls": ["https://example.com/first-frame.jpg"],
"videoInputMode": "first_frame",
"duration": 5,
"resolution": "2K",
"aspectRatio": "16:9",
},
)
print("taskId:", i2v_resp.json())
# 轮询方式同上
# ─── 图生视频(首尾帧驱动)───
fl_resp = requests.post(
"https://api.tokenlx.ai/v1/aigc/video/tasks",
headers=headers,
json={
"model": "MiniMax-H3",
"prompt": "平滑过渡",
"referenceImageUrls": [
"https://example.com/start.jpg",
"https://example.com/end.jpg",
],
"videoInputMode": "first_last_frame",
"duration": 5,
"resolution": "2K",
"aspectRatio": "16:9",
},
)
print("taskId:", fl_resp.json())
# ─── 参考视频生成示例 ───
ref_resp = requests.post(
"https://api.tokenlx.ai/v1/aigc/video/tasks",
headers=headers,
json={
"model": "MiniMax-H3",
"prompt": "保持相同的运动风格,换成海边场景",
"videoInputMode": "reference",
"referenceVideoUrls": ["https://example.com/reference.mp4"],
"duration": 5,
"resolution": "2K",
"aspectRatio": "adaptive",
},
)
print("taskId:", ref_resp.json())Replace sk-aihubrouter-… with your key from the dashboard.
Parameter Reference
POST /v1/aigc/video/tasksGET /v1/aigc/video/tasks/{taskId}?model=MiniMax-H3Response is raw upstream JSON passthrough — fields vary by channel.
MiniMax-H3 uses the omni Video V2 API. Billed per output second by resolution (768P / 2K). Prompt ≤7000 characters. Mixed reference files ≤12 total.
| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | Yes | Model name |
| prompt | string | Yes | Video description prompt |
| duration | int | Yes | Duration in seconds (integer 4–15) 456789101112131415 |
| resolution | string | No | Output resolution (default 768P; 2K available) 768P2K |
| aspectRatio | string | No | Aspect ratio. Text-to-video requires an explicit ratio (not adaptive). 16:99:161:14:33:421:9adaptive |
| referenceImageUrls | array | Image-to-video | Reference images (≤9). First/last frame or multimodal reference. Formats: JPG/JPEG/PNG/WebP/HEIC; ≤30MB each. |
| referenceVideoUrls | array | Ref-video | Reference videos (≤3 clips, each 2–15s, total ≤15s, ≤50MB each) |
| referenceAudioUrls | array | No | Reference audio (≤3 clips, each 2–15s; must pair with image or video) |
| videoInputMode | string | No | Input mode. Auto-inferred if omitted: 1 image → first_frame, 2 → first_last_frame, ≥3 or with video/audio → reference first_framefirst_last_framereference |
modelstringYespromptstringYesdurationintYes456789101112131415resolutionstringNo768P2KaspectRatiostringNo16:99:161:14:33:421:9adaptivereferenceImageUrlsarrayImage-to-videoreferenceVideoUrlsarrayRef-videoreferenceAudioUrlsarrayNovideoInputModestringNofirst_framefirst_last_framereference