Doubao realtime audio transcription streams speech-to-text over AIHub’s gRPC `realTimeAudioTranscription` API.
Key strengths
- Realtime streaming ASR
- Chunked audio frames
- Low-latency transcripts
Use cases
- Live captions
- Meeting assistants
- Voice input pipelines
ByteDance's bytedance/doubao-realtime-audio-transcription is a high-quality speech model. It generates natural-sounding speech across multiple voices and languages, with low-latency streaming output suitable for real-time voice applications.
Supports SSML-style controls, configurable voices, speaking rate, and pitch. Compatible with the OpenAI `/audio/speech` and `/audio/transcriptions` endpoint shapes.
bytedance/doubao-realtime-audio-transcription is fully OpenAI-compatible — drop in your existing OpenAI Python or Node SDK and switch `baseURL` to `https://api.tokenlx.ai`. TokenLX transparently routes your requests to the optimal provider endpoint while preserving streaming, function-calling, and structured-output semantics.
Performance
Compare different providers across TokenLX · All locations.
Effective Pricing
Pricing is shown by the model billing method, using per-call or per-second prices and resolution tiers.
Recent activity
Total usage per day on TokenLX (last 30 days).
Sample code & API
TokenLX normalizes requests and responses across providers. Use any OpenAI SDK or our native SDK.
# doubao-realtime-audio-transcription uses AIHub gRPC streaming realTimeAudioTranscription (not REST /audio/transcriptions).
# Pseudo-flow with the AIHub SDK:
# stream = client.realTimeAudioTranscription(...)
# stream.send(model="doubao-realtime-audio-transcription", userId="user-1", seq=1, isLast=False,
# sampleRate=16000, channels=1, sampleFormat=16, audioData=pcm_chunk)
print("Use AIHub SDK realTimeAudioTranscription for model=doubao-realtime-audio-transcription")Replace sk-aihubrouter-… with your key from the dashboard.