Flagship multimodal model from Azure with native text, vision, and voice understanding. Strong at general-purpose reasoning and instruction following.
Key strengths
- Multimodal (text + vision + audio)
- Function calling and JSON mode
- Strong instruction following
- 128K context window