245 models · 13 providers
Claude Opus — AWS's most capable (and expensive) tier. Reserved for the hardest problems.
DeepSeek V4 Flash Vision (Experimental) is a multimodal understanding model. Images are converted to tokens by size and billed with text at the same peak rates as V4 Flash (Beijing Monday–Friday 09:00–12:00 and 14:00–18:00); weekends and other hours are off-peak at half, and thinking mode uses the same token rates.
Alibaba Model Studio route for DeepSeek V4 Flash. List prices on this page are busy-hour rates (Beijing 08:00–22:00); off-peak (22:00–08:00) is half. Cache-hit is 3× the official DeepSeek cache rate; thinking mode uses the same token rates.
Alibaba Model Studio route for DeepSeek V4 Pro. List prices on this page are busy-hour rates (Beijing 08:00–22:00); off-peak (22:00–08:00) is half. Cache-hit is 3× the official DeepSeek cache rate; thinking mode uses the same token rates.
DeepSeek V4 Flash is the lower-cost V4 tier for high-volume chat, light coding, and budget agents. List prices on this page are peak rates (Beijing Monday–Friday 09:00–12:00 and 14:00–18:00); weekends and other hours are off-peak at half, and thinking mode uses the same token rates.
DeepSeek V4 Pro is the flagship V4 model for coding, agents, and long-context work. List prices on this page are peak rates (Beijing Monday–Friday 09:00–12:00 and 14:00–18:00); weekends and other hours are off-peak at half, and thinking mode uses the same token rates.
Tencent TokenHub route for DeepSeek V4 Flash, billed at the official peak/off-peak rates. List prices on this page are peak rates (Beijing 09:00–12:00 and 14:00–18:00); off-peak is half, and thinking mode uses the same token rates.
Tencent TokenHub route for DeepSeek V4 Pro, billed at the official peak/off-peak rates. List prices on this page are peak rates (Beijing 09:00–12:00 and 14:00–18:00); off-peak is half, and thinking mode uses the same token rates.
Tencent-hosted Kling 3.0-Omni video model. It generates 3–15s clips (3–10s with a reference video) at 720P–4K, with optional native audio, first/last-frame I2V, and multi-image or reference-video control.
Tencent-hosted Kling 3.0 video model. It generates 3–15s clips at 720P–4K with optional audio (and optional voice timbre). It supports text-to-video and first/last-frame I2V, but not reference-video mode.
MiniMax-H3 is a general-purpose omni-modal video model that understands text, image, video, and audio context together, and generates video with native stereo audio up to 15s at 2K. It supports text-to-video, first/last-frame I2V, and multimodal reference-to-video.
Qwen-Image-3.0 Standard supports up to 4.5k token input, stable 10px text rendering, 12-language native typography, and batch-friendly generation for posters, webpages, and UI — optimised for cost-effective continuous creation at 1K/2K (same price).
Qwen-Image-3.0-Pro supports up to 4.5k token input, dense information layouts (newspapers, storyboards, menus, exam papers), 10px small-text rendering, 12-language native typography, and lifelike micro-expression / hair / pore detail at up to 2K resolution.
Qwen3.8-Max is Alibaba’s flagship Qwen model (2.4T MoE, ~95B activated) with a 1M-token context window. It targets coding, office copilots, research, and long-horizon agent tasks, with dual thinking / non-thinking modes and multimodal input.
AIHub routing name for the GLM 5.2 family served through supported cloud channels. It is positioned for bilingual reasoning, coding, and tool-assisted workflows without asserting channel-specific specifications.
AIHub routing name for a Kimi K2.7 Code capability exposed by supported cloud channels. The metadata describes its code-focused placement while avoiding unverified context or benchmark claims.
Claude Fable 5 is an AIHub routing name used across supported channels, not a separately verified Anthropic model name. Its placement emphasizes writing, narrative, and high-care language workflows.
Claude Fable 5 is an AIHub routing name used across supported channels, not a separately verified Anthropic model name. Its placement emphasizes writing, narrative, and high-care language workflows.
Claude Sonnet 5 is an AIHub routing name shared by supported provider channels, pending a separately verified Anthropic model specification. It is positioned as a balanced route for coding, tools, and production language tasks.
Claude Sonnet 5 is an AIHub routing name shared by supported provider channels, pending a separately verified Anthropic model specification. It is positioned as a balanced route for coding, tools, and production language tasks.
Doubao Seed 2.1 Pro is the quality-oriented route in the Seed 2.1 family on Volcano Engine. It is positioned for more demanding reasoning, writing, and enterprise language workflows.
Doubao Seed 2.1 Turbo is the speed-oriented route in the Seed 2.1 family on Volcano Engine. It targets responsive general language and agent interactions without relying on unpublished benchmarks.
Doubao Seed Evolving is an AIHub route name for an evolving hosted Seed capability on Volcano Engine. Because an independent public specification is not assumed, metadata stays limited to its general adaptive language-workflow placement.
Doubao Seed3D 2.0 is a Volcano Engine hosted route for generating three-dimensional content. Its listing describes the 3D workflow role while avoiding unsupported geometry or texture specifications.
Doubao Seedance 2.0 Mini is a compact Volcano Engine route for generated-video workflows. It is positioned for accessible iteration and short-form creative production.
Doubao Seedream 5.0 Lite is the iteration-oriented AIHub route for Volcano Engine image generation. It is positioned for frequent drafts and lightweight visual workflows.
Doubao Seedream 5.0 Pro is the quality-oriented AIHub route for Volcano Engine image generation. It targets demanding creation and editing workflows without adding unpublished resolution claims.
Eleven v3 is an ElevenLabs text-to-speech route focused on expressive, natural voice generation. It is suited to spoken-content workflows that benefit from nuanced delivery.
Gemini Embedding 001 is Google’s text embedding model for mapping language into vectors. It supports semantic retrieval, similarity, clustering, and classification workflows.
Gemini Embedding 2 is an AIHub route for Google’s newer embedding capability, including multimodal inputs where supported. Metadata avoids assigning dimensions or limits not stated by the active provider documentation.
Zhipu GLM — Chinese LLM from Tsinghua. Solid bilingual support with academic training roots.
Zhipu GLM — Chinese LLM from Tsinghua. Solid bilingual support with academic training roots.
GPT-5.6 Luna is an AIHub-managed Azure routing profile, not a separately documented OpenAI model name. Luna denotes an interaction-oriented profile for responsive assistant and iterative tasks.
GPT-5.6 Sol is an AIHub-managed Azure routing profile, not a separately documented OpenAI model name. Sol denotes the balanced default profile for broad production language workloads.
GPT-5.6 Terra is an AIHub-managed Azure routing profile, not a separately documented OpenAI model name. Terra denotes a deliberate profile for sustained analysis and complex workflow orchestration.
AIHub routing name for a hosted Alibaba Model Studio video capability. Publicly independent specifications for this route are not assumed; use it for prompt- and reference-guided video creation.
HappyOyster is an AIHub routing name for an Alibaba-hosted video workflow, not a separately verified public model specification. It is presented as a managed creative-video capability.
HiTem3D 2.0 is an AIHub route name for a 3D asset capability hosted through Volcano Engine. No independent public specification is inferred, so the description remains focused on managed asset generation.
Hyper3D Gen2 is an AIHub route name for a Volcano Engine hosted 3D generation capability. The metadata does not assign unverified mesh, material, or resolution limits.
Moonshot Kimi K3 — flagship model with a 1M-token context window, multimodal understanding, and strong fit for long-horizon coding and knowledge work.
MiniMax-M3 is MiniMax’s long-context language model with up to 1M tokens. Pricing is tiered by input length (≤512K vs 512K–1M). It fits document analysis, bilingual assistants, and general production chat / tool workflows.
OpenAI Whisper is an open-source automatic speech recognition model for multilingual transcription and speech translation. This route provides hosted access for audio-to-text workflows.
Qwen-Image-Layered generates images as semantically separated layers for subsequent composition and editing. It is designed for workflows that need more control than a single flattened image.
Tencent-hosted route for Gemini 2.5 Flash image generation, retaining the fast multimodal family placement. It supports prompt-led visual creation through AIHub without channel-specific capability claims.
AIHub routing name for a Tencent-hosted Gemini 3 Pro image capability. The Pro label indicates a quality-oriented placement; unpublished Google or channel specifications are not inferred.
AIHub routing name for a Tencent-hosted Gemini 3.1 Flash image capability. The Flash label indicates an iteration-oriented placement, without asserting unpublished model limits.
AIHub routing name for the GLM 5.2 family served through supported cloud channels. It is positioned for bilingual reasoning, coding, and tool-assisted workflows without asserting channel-specific specifications.
AIHub routing name for a Kimi K2.7 Code capability exposed by supported cloud channels. The metadata describes its code-focused placement while avoiding unverified context or benchmark claims.
Tripo H3.1 is a hosted 3D generation model for turning text or reference images into three-dimensional assets. It supports rapid asset ideation through Tripo-compatible workflows.
Tripo P1.0 is a hosted 3D asset generation route accepting text or image guidance. It is suited to producing editable starting points for downstream 3D workflows.
Wan 2.2 Animate Move is an Alibaba Model Studio route for transferring or directing motion in generated video. It focuses on animation movement rather than generic text-to-video output.
Wan 2.2 S2V is an Alibaba Model Studio speech-to-video route that drives a visual subject from audio guidance. It is intended for synchronized character and presenter-style clips.
AIHub routing name for a Wan 2.7 image capability optimized for interactive generation sessions. The route emphasizes rapid visual iteration without claiming unpublished model limits.
Moonshot Kimi — long-context Chinese model known for strong document reading and comprehension.
Claude Opus — AWS's most capable (and expensive) tier. Reserved for the hardest problems.
Claude Opus — AWS's most capable (and expensive) tier. Reserved for the hardest problems.
Claude Opus — AWS's most capable (and expensive) tier. Reserved for the hardest problems.
Claude Opus — AWS's most capable (and expensive) tier. Reserved for the hardest problems.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Moonshot Kimi — long-context Chinese model known for strong document reading and comprehension.
Text generation model. Compatible with the OpenAI Chat Completions API.
Text generation model. Compatible with the OpenAI Chat Completions API.
Claude Opus — AWS's most capable (and expensive) tier. Reserved for the hardest problems.
Claude Opus — AWS's most capable (and expensive) tier. Reserved for the hardest problems.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Text-to-image model. Generates original images from natural-language prompts.
Text-to-image model. Generates original images from natural-language prompts.
Video generation model. Produces video clips from text or images.
Text generation model. Compatible with the OpenAI Chat Completions API.
Text generation model. Compatible with the OpenAI Chat Completions API.
Zhipu GLM — Chinese LLM from Tsinghua. Solid bilingual support with academic training roots.
Text generation model. Compatible with the OpenAI Chat Completions API.
MiniMax — Chinese LLM family with hybrid attention for extreme-length contexts.
Text generation model. Compatible with the OpenAI Chat Completions API.
Claude Sonnet — AWS's balanced model. Strong coding, writing, and tool use with 200K context.
Claude Sonnet — AWS's balanced model. Strong coding, writing, and tool use with 200K context.
Latest-generation frontier model with expanded reasoning and faster tool execution. Top choice when quality trumps cost.
Gemini Flash — fast Google multimodal model with long context. Best value for volume tasks.
Gemini 3.1 Flash TTS converts text into natural speech through AIHub’s `/generate/speech` route.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Gemini Flash — fast Google multimodal model with long context. Best value for volume tasks.
Gemini Pro — Google's higher-quality Gemini tier. Strong reasoning with large context windows.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Qwen 3.5 Omni is a multimodal chat model routed through AIHub’s OpenAI-compatible `/chat/completions` endpoint.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Text-to-image model. Generates original images from natural-language prompts.
Moonshot's Kimi K2.5 — Chinese-first model with exceptional long-context ability. Known for strong reading comprehension.
Text-to-image model. Generates original images from natural-language prompts.
Text generation model. Compatible with the OpenAI Chat Completions API.
Text generation model. Compatible with the OpenAI Chat Completions API.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
Text generation model. Compatible with the OpenAI Chat Completions API.
Text generation model. Compatible with the OpenAI Chat Completions API.
Zhipu GLM — Chinese LLM from Tsinghua. Solid bilingual support with academic training roots.
Text generation model. Compatible with the OpenAI Chat Completions API.
Text generation model. Compatible with the OpenAI Chat Completions API.
Zhipu GLM — Chinese LLM from Tsinghua. Solid bilingual support with academic training roots.
Claude Opus — AWS's most capable (and expensive) tier. Reserved for the hardest problems.
Claude Opus — AWS's most capable (and expensive) tier. Reserved for the hardest problems.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Text generation model. Compatible with the OpenAI Chat Completions API.
MiniMax — Chinese LLM family with hybrid attention for extreme-length contexts.
Text generation model. Compatible with the OpenAI Chat Completions API.
Text-to-image model. Generates original images from natural-language prompts.
Video generation model. Produces video clips from text or images.
Video generation model. Produces video clips from text or images.
Upgraded GPT-5 with longer context and improved latency. Production default for demanding agentic workloads.
Codex variant of GPT-5.2 tuned for software engineering. Specialized for repo-aware coding agents.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
Claude Opus — AWS's most capable (and expensive) tier. Reserved for the hardest problems.
Claude Opus — AWS's most capable (and expensive) tier. Reserved for the hardest problems.
Gemini Pro — Google's higher-quality Gemini tier. Strong reasoning with large context windows.
Gemini Pro — Google's higher-quality Gemini tier. Strong reasoning with large context windows.
MiniMax — Chinese LLM family with hybrid attention for extreme-length contexts.
Zhipu GLM — Chinese LLM from Tsinghua. Solid bilingual support with academic training roots.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
Claude Haiku — fast, affordable AWS model. Best for high-volume real-time tasks.
Claude Haiku — fast, affordable AWS model. Best for high-volume real-time tasks.
Video generation model. Produces video clips from text or images.
Video generation model. Produces video clips from text or images.
MiniMax — Chinese LLM family with hybrid attention for extreme-length contexts.
Text-to-image model. Generates original images from natural-language prompts.
Image generation model. Creates or edits images from text prompts.
Video generation model. Produces video clips from text or images.
HappyHouse video edit restyles source videos asynchronously through `/edit/video/tasks`.
Zhipu GLM — Chinese LLM from Tsinghua. Solid bilingual support with academic training roots.
Zhipu GLM — Chinese LLM from Tsinghua. Solid bilingual support with academic training roots.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Claude Sonnet — AWS's balanced model. Strong coding, writing, and tool use with 200K context.
Text generation model. Compatible with the OpenAI Chat Completions API.
DeepSeek — open-weight Chinese LLM family. Strong cost-to-quality ratio and good code generation.
Claude Sonnet — AWS's balanced model. Strong coding, writing, and tool use with 200K context.
Text-to-image model. Generates original images from natural-language prompts.
Video generation model. Produces video clips from text or images.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
Tencent video edit creates asynchronous media-processing tasks through `/edit/video/tasks`.
Tencent video enhance runs asynchronous enhancement jobs through `/edit/video/tasks`.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
MiniMax Music 2.0 generates songs from a style prompt and optional lyrics. AIHub exposes a synchronous music generation route that returns hosted audio URLs.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Upgraded DeepSeek V3.1 with improved reasoning and better tool calling. Pareto-optimal on cost vs quality.
Text generation model. Compatible with the OpenAI Chat Completions API.
Text generation model. Compatible with the OpenAI Chat Completions API.
Text generation model. Compatible with the OpenAI Chat Completions API.
Text generation model. Compatible with the OpenAI Chat Completions API.
Text generation model. Compatible with the OpenAI Chat Completions API.
Doubao witty-remark is an AIHub realtime audio transcription route using the same gRPC streaming ASR contract.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Tencent image translate localizes text inside images through `/image/translate`.
Gemini Flash — fast Google multimodal model with long context. Best value for volume tasks.
Video generation model. Produces video clips from text or images.
Video generation model. Produces video clips from text or images.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
Doubao realtime audio transcription streams speech-to-text over AIHub’s gRPC `realTimeAudioTranscription` API.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Vidu Voice Clone synthesizes speech from text using a cloned voice profile via `/generate/speech`.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
DeepSeek — open-weight Chinese LLM family. Strong cost-to-quality ratio and good code generation.
Claude 4 Sonnet — balance of speed, quality, and cost for agentic workflows and production coding.
Claude Sonnet — AWS's balanced model. Strong coding, writing, and tool use with 200K context.
Video generation model. Produces video clips from text or images.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Alibaba Qwen series — Chinese-first LLMs with strong bilingual support. Wide range from turbo to max tiers.
Latest Azure image model with improved realism and editing. Supports inpainting, outpainting, and mask-guided edits.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Next-generation reasoning model succeeding o1. Solves problems that stumped previous models, at a reasonable cost.
Azure o-series — reasoning-first models that think before answering. Best for hard math, science, and code.
Code-focused GPT-4 successor with stronger instruction following and 1M+ context. Great for long-document analysis and agentic coding.
Tiny model optimized for classification and structured output. Cheapest in the GPT-4 family.
Smaller GPT-4.1 with the same 1M context at a fraction of the cost. The new default for long-context RAG and bulk processing.
MiniMax — Chinese LLM family with hybrid attention for extreme-length contexts.
ByteDance Doubao — Chinese LLM family tuned for the Volcano Engine cloud and ByteDance ecosystem.
Gemini 2.5 Pro — Google's top reasoning model with thinking mode. Frontier performance on coding and math.
Gemini 2.5 Pro TTS focuses on higher-quality text-to-speech generation through AIHub’s speech endpoint.
Gemini Flash — fast Google multimodal model with long context. Best value for volume tasks.
Gemini Flash — fast Google multimodal model with long context. Best value for volume tasks.
Gemini 2.5 Flash TTS is a text-to-speech route for fast spoken-content generation via `/generate/speech`.
Claude Sonnet — AWS's balanced model. Strong coding, writing, and tool use with 200K context.
Claude Sonnet — AWS's balanced model. Strong coding, writing, and tool use with 200K context.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
DeepSeek R1 — a reasoning-first model trained with reinforcement learning. Competes with o1-class models at much lower cost.
Text-to-image model. Generates original images from natural-language prompts.
Text-to-image model. Generates original images from natural-language prompts.
Open-weight DeepSeek V3 — MoE architecture delivering frontier-adjacent quality at a fraction of the cost.
Next-gen Gemini Flash with improved reasoning and native tool use. Drop-in upgrade to 1.5 Flash.
Gemini Flash — fast Google multimodal model with long context. Best value for volume tasks.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Text generation model. Compatible with the OpenAI Chat Completions API.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Balanced Qwen tier with strong Chinese + reasonable cost. The pragmatic default for production Chinese apps.
Alibaba's flagship Qwen model. Strong bilingual (Chinese / English) performance, especially tuned for enterprise scenarios.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Reasoning-first model that thinks before answering. Best for math, science, and multi-step problem solving.
Text generation model. Compatible with the OpenAI Chat Completions API.
Text generation model. Compatible with the OpenAI Chat Completions API.
Text generation model. Compatible with the OpenAI Chat Completions API.
Claude Sonnet — AWS's balanced model. Strong coding, writing, and tool use with 200K context.
Text-to-image model. Generates original images from natural-language prompts.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Text-to-video or image-to-video model. Generates short video clips with configurable duration and resolution.
Video generation model. Produces video clips from text or images.
Video generation model. Produces video clips from text or images.
Video generation model. Produces video clips from text or images.
Text generation model. Compatible with the OpenAI Chat Completions API.
Flagship multimodal model from Azure with native text, vision, and voice understanding. Strong at general-purpose reasoning and instruction following.
Cheap and fast sibling of GPT-4o. Best value for high-volume classification, extraction, and routing tasks.
Qwen variant with very long context (10M+ tokens). Purpose-built for long-document analysis and codebase-level tasks.
Smallest, cheapest Qwen. Good for classification, routing, and high-volume light tasks in Chinese.
Claude Haiku — fast, affordable AWS model. Best for high-volume real-time tasks.
Claude Sonnet — AWS's balanced model. Strong coding, writing, and tool use with 200K context.
Azure's production image generator. Known for strong prompt adherence and coherent in-image text rendering.