DeepSeek V4 Pro is the flagship V4 model for coding, agents, and long-context work. List prices on this page are peak rates (Beijing Monday–Friday 09:00–12:00 and 14:00–18:00); weekends and other hours are off-peak at half, and thinking mode uses the same token rates.
Key strengths
- Flagship V4 quality
- Peak / off-peak billing since 2026-08-17
- 1M-token context
- Coding and agent workflows
Use cases
- Production assistants
- Code and tool agents
- Long-context analysis
- Cost-aware peak scheduling
When Thinking mode is enabled, reasoning tokens generated by the model are counted as billable output. This can increase total usage beyond the visible answer tokens.
DEEPSEEK's deepseek/deepseek-v4-pro is a frontier text generation model in the DEEPSEEK family. It excels at complex reasoning, agentic workflows, code generation, and long-form writing tasks, with native support for streaming, tool calling, JSON mode, and multi-turn conversations.
The model handles long-context inputs gracefully and is particularly effective for software engineering, multi-step research, and end-to-end project execution. Its tokenizer and pricing are optimized for high-throughput production workloads, with a competitive cost profile relative to other models in its tier.
deepseek/deepseek-v4-pro is fully OpenAI-compatible — drop in your existing OpenAI Python or Node SDK and switch `baseURL` to `https://api.tokenlx.ai`. TokenLX transparently routes your requests to the optimal provider endpoint while preserving streaming, function-calling, and structured-output semantics.
Performance
Compare different providers across TokenLX · All locations.
Effective Pricing
Actual cost per million tokens across providers over the past 7 days. Prices above are peak rates (Beijing Mon–Fri 09:00–12:00, 14:00–18:00). Weekends and other hours are off-peak at half.
Recent activity
Total usage per day on TokenLX (last 30 days).
Sample code & API
TokenLX normalizes requests and responses across providers. Use any OpenAI SDK or our native SDK.
Disable Thinking when you do not need explicit reasoning, or set a lower budget_tokens value to cap the reasoning length. Only enable return_thoughts when you need to inspect the thinking process.
from openai import OpenAI
client = OpenAI(
base_url="https://api.tokenlx.ai/v1",
api_key="sk-aihub-...",
)
# Non-streaming
response = client.chat.completions.create(
model="deepseek-v4-pro",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"},
],
# Optional: enable Thinking / reasoning.
extra_body={
"thinking": {
"enabled": True,
"budget_tokens": 2048,
"return_thoughts": True,
}
},
)
print(response.choices[0].message.content)
# Streaming
stream = client.chat.completions.create(
model="deepseek-v4-pro",
messages=[{"role": "user", "content": "Tell me a story"}],
stream=True,
extra_body={
"thinking": {
"enabled": True,
"budget_tokens": 2048,
"return_thoughts": True,
}
},
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)Replace sk-aihubrouter-… with your key from the dashboard.
API Endpoints
Sends a request for a model response for the given chat conversation. Supports both streaming and non-streaming modes.
Creates a streaming or non-streaming response using the OpenAI Responses API format.
Creates a message using the Anthropic Messages API format. Supports text, images, tools, and extended thinking.