Cohere
Enterprise-first provider focused on retrieval, grounding, multilingual control, and private deployment. North Mini Code 1.0 (9 Jun, Apache 2.0) extends the stack into agentic coding. Command A+ is the newest open-source workhorse.
Latest model families
- North Mini Code 1.0 (9 Jun) - 30B-A3B Apache coding MoE
- Command A+ - newest open-source enterprise workhorse
- Command A / Command A Reasoning - flagship enterprise models
- Command A Vision / Command A Translate - multimodal and translation lines
- Command R7B / Command R+ / Command R - retrieval and grounding family
- Tiny Aya - live multilingual line
- Transcribe Arabic - speech recognition
- Embed / Rerank / Transcribe - retrieval and audio stack
Best for
- Agentic coding via North Mini Code
- RAG and enterprise search
- Grounded assistants
- Private and sovereign deployments
Amazon Nova / Bedrock
A production platform rather than a single model lab. Bedrock is the control plane; Nova is Amazon's own model line. Current official IDs are Nova 2 Lite, Nova 2 Sonic, and Nova Multimodal Embeddings.
Latest model families
- Amazon Nova 2 Lite - current text/multimodal tier
- Nova 2 Sonic - speech and conversational voice
- Nova Multimodal Embeddings - semantic retrieval
- Third-party model access via Bedrock
- Legacy: Premier / Canvas / Reel / Sonic v1 → EOL Sep 2026
- Command R+ on Bedrock → EOL 19 Aug 2026
Best for
- AWS-native production apps
- Enterprise governance and routing
- Multi-provider deployments
Microsoft Phi
Microsoft's small-model family is optimized for strong performance per parameter and edge use cases. The current work centers on compact reasoning, multimodal SLMs, and vision reasoning.
Latest model families
- Phi-4-reasoning-vision-15B - compact multimodal reasoning
- Phi-4-reasoning / Phi-4-reasoning-plus - compact reasoning
- Phi-4 - small-model flagship
- Phi-4-mini - compact variant
- Phi-4-multimodal - vision, audio, text
- Phi-3.5 - still widely used in light-footprint setups
Best for
- Small-footprint deployments
- STEM-heavy tasks
- Edge and local inference
Microsoft MAI
Microsoft's Foundry model family for enterprise reasoning, coding, image, and speech. MAI is a separate line from Phi; Phi stays the SLM/edge family.
Latest model families
- MAI-Thinking-1 (public preview 12 Aug) - enterprise reasoning
- MAI-Code-1-Flash - coding (shipped)
- MAI-Image-2.5 - image generation (shipped)
- MAI-Transcribe-1.5 - speech-to-text (shipped)
- MAI-Voice-2 - text-to-speech (shipped)
- MAI-Voice-2-Flash - coming soon (not yet shipped)
Best for
- Enterprise reasoning on Foundry
- Coding, image, and speech workloads on Foundry
- Phi remains the choice for SLM and edge use cases
Baidu ERNIE (China)
Baidu's China enterprise stack emphasizes search, multimodal understanding, and document-heavy workflows. ERNIE 5.1 (hosted; launched 8–9 May) joins 5.0 as the native multimodal flagship.
Latest model families
- ERNIE 5.1 (hosted) - text and agent workloads
- ERNIE 5.0 / ERNIE 5.0 Thinking / ERNIE 5.0 Preview - native multimodal flagship
- ERNIE X1.1 / X1.1 Preview
- ERNIE 4.5 Turbo / ERNIE 4.5 Turbo VL
- Qianfan OCR: paddleocr-vl-0.9b - document stack
Best for
- ERNIE 5.1 for text and agent workloads
- ERNIE 5.0 for full-modal enterprise deployments
- Chinese enterprise search-augmented assistants
- Document and multimodal knowledge work
Zhipu AI / GLM (China)
One of China's strongest general-purpose model lines. GLM-5.3 full weights are ungated as of 28 Aug (HF 753.3B; GitHub table 744B-A40B). GLM-5.2 (mid-Jun, 1M, MIT weights) remains the prior open tier. GLM-5.3-Flash (26 Aug) is the first natively multimodal GLM-5 tier with MIT Flash weights. Full GLM-5.3 uses custom glm-5.3 license, not MIT. Do not list GLM-5.4.
Latest model families
- GLM-5.3-Flash (26 Aug) - first natively multimodal GLM-5; 320B-A18B; hosted
glm-5.3-flash; MIT weights zai-org/GLM-5.3-Flash
- GLM-5.3 (full weights 28 Aug) - text-only agentic coding; ungated open weights; custom
glm-5.3 license; HF 753.3B / GitHub 744B-A40B
- GLM-5.3 (hosted 14 Aug) - same base as 5.2 with post-training only
- GLM-5.2 (mid-Jun) - 1M context, MIT open weights
- GLM-5-Turbo - current high-throughput SKU
- GLM-4.7 / GLM-4.7-FlashX - high-intelligence general line
- GLM-5V-Turbo / GLM-4.6V-Flash - multimodal coding and vision
- GLM-Image - optional image generation
Best for
- Cheap multimodal and visual coding via 5.3-Flash
- Text agentic coding via 5.3 hosted
- Agentic and long-horizon coding
- Chinese-language workflows
- Multimodal app stacks
Moonshot / Kimi (China)
A fast-moving China-side assistant platform with strong long-context and coding angles. Kimi K3 (hosted 16 Jul; open weights 27 Jul; 2.8T MoE, native vision, 1M) is the current flagship. k2.5 and moonshot-v1* platform sunset completed 31 Aug 2026. No K4.
Latest model families
- Kimi K3 (hosted 16 Jul; open weights 27 Jul) - 2.8T MoE, native vision, 1M context
- Kimi K2.7 Code (~12 Jun) - coding-focused tier
- Kimi K2.6 - prior flagship still in circulation
- k2.5 and moonshot-v1* - sunset completed 31 Aug 2026
Best for
- 1M long-context assistants
- Native multimodal agents
- First-party Chinese coding and knowledge work
MiniMax (China)
MiniMax spans text, speech, video, and music product lines. M3 (1 Jun; 1M; native image/video) handles agentic coding; MiniMax H3 (31 Jul; omni video, up to 15s / 2K) is the current video line. Official name is MiniMax H3, not Hailuo 3.0.
Latest model families
- MiniMax M3 (1 Jun) - text and agent model, 1M context, native image/video
- MiniMax H3 (31 Jul) - omni video, up to 15s / 2K; partial OSS
- MiniMax Speech 2.8 - current T2A
- MiniMax Music 3.0 (16 Jul) - music generation; new-user Music APIs closed 20 Aug
Best for
- Omni video and audio via H3
- Agentic coding via M3
- Speech products via Speech 2.8
Tencent Hy (Hunyuan) (China)
Tencent Hy (formerly Hunyuan) spans language, role-play, translation, 3D, and video. Hy3 preview became GA 6 Jul (295B-A21B, 256K, Apache 2.0) as the language flagship. Hy4 preview (Aug 2026) is Apache open weights, not GA. UI-Mate (9B / 27B / democua-27B) is the GUI-agent sub-line on this card.
Latest model families
- Hy3 (GA 6 Jul) - 295B-A21B, 256K, Apache 2.0 flagship
- Hy4 preview (Aug 2026) - Apache open weights; not GA
- UI-Mate (9B / 27B / democua-27B) - GUI and computer-use agents
- Hunyuan-role-latest - role-play and conversational character model
- HY-MT2-Pro - translation-focused model
- HY-3D-3.1 / HY-3D-3.0 - 3D generation
- HunyuanVideo / HunyuanVideo-Avatar - video generation
Best for
- GUI and computer-use agents via UI-Mate
- Hy3 agent and productivity workloads
- 3D, simulation, and interactive content
- Media-heavy AI experiences
NVIDIA Nemotron
NVIDIA's open model family for agent workhorses, available via Hugging Face and NIM. Nemotron 3.5 Lightning (11 Aug, 30B-A3B) joins Nano, Super, and Ultra tiers already public.
Latest model families
- Nemotron 3.5 Lightning (11 Aug) - 30B-A3B open MoE
- Nemotron Nano / Super / Ultra - prior public tiers
Best for
- Open agent workhorses
- Local and DGX deployment
- NIM-served inference
IBM Granite
IBM's open enterprise model family for on-prem and Apache-licensed agent workloads. Granite 4.2 (25 Aug) adds dense reasoning models with native thinking on/off; Speech 5.0 Turbo CTC covers high-throughput ASR. A separate NC (non-commercial) variant exists — do not present NC as Apache.
Latest model families
- Granite 4.2 (25 Aug) - dense 3B / 8B / 30B Apache 2.0; native thinking on/off
- Granite 4.2 30B - SWE and terminal workloads
- Granite Speech 5.0 Turbo CTC (470M) - high-throughput ASR
- NC variants - non-commercial license; not Apache
Best for
- On-prem and Apache enterprise agents
- Thinking on/off control
- 30B SWE and terminal tasks
- Speech 5.0 high-throughput ASR
Perplexity
First-party Sonar and Agent API for web-grounded Q&A and research agents. Sonar Chat Completions support ended 27 Sep 2026. Sync and streaming Sonar requests may still work but are being reformulated as Agent API requests (rolling out by model); async Sonar is not reformulated and is no longer supported — use background mode. Build path: Agent API presets (fast / low / medium / high; xhigh for top deep research). Not to be confused with OpenRouter.
Latest model families
- Agent API presets - current build path (Sonar tiers map here)
- Sonar / Sonar Chat Completions - support ended 27 Sep; migrate or use reformulation shim to Agent API
Best for
- Post-Sonar: Agent API presets for web-grounded Q&A and research agents
- Web-grounded Q&A with citations
- Research agents
- Products needing live web retrieval
Black Forest Labs (FLUX) (EU)
EU-based image and video generation API. FLUX.2 and FLUX 3 Video are live on the BFL API (preview). FLUX 3 Image and Dev weights are not GA.
Latest model families
- FLUX.2 - image generation on BFL API
- FLUX 3 Video - video generation on BFL API (preview)
Best for
- EU image generation API
- EU video generation API (preview)