AI Providers Landscape

A current research brief on model families and platform stacks, organized by frontier providers, Tier 1–3, and emerging model classes. It uses provider names and product families, keeps the global market explicit, and prioritizes official docs and announcements first.

Last updated: 2026-10-01
Quick source check on 2026-10-01: OpenAI flagship still GPT-6 Astra; add GPT-6.1 Sol (gpt-6.1-sol) GA 29 Sep (near-Astra coding/computer-use/professional work at $2/$10; Multi-agent beta); keep GPT-6 Sol + Luna (22 Sep) as served prior mid/efficient tiers; Agents API still public beta — 29 Sep added computer use (OpenAI-hosted browser) + Astra Ultrafast mode (service_tier: "ultrafast"). Anthropic: Claude Sonnet 5.5 (claude-sonnet-5-5) GA 28 Sep — best speed/intelligence combo; Opus 5.5 remains recommended default for most workloads; Fable 5.1 / Haiku 4.5 unchanged; no Haiku 5.5 (roadmap “coming weeks”); Sonnet 4.5 deprecation announced 30 Sep (retire 30 Nov). Google: no Gemini 4; model families unchanged vs 24 Sep card; Omni Flash preview endpoint deprecate date 30 Sep is past — use GA gemini-omni-1.1-flash; Antigravity …-05-2026 still shuts 5 Oct (~4 days). SpaceXAI: Grok 4.7 still flagship (no-change); Sep notes add safety_identifier only. DeepSeek: no-change — V4.1-Flash default; V4-Pro continues; no V4.1-Pro. Perplexity: Sonar support ended 27 Sep — rewrite urgent countdown to post-EOL / migrate completed or reformulating; sync+stream may still work as Agent API reformulations; async Sonar unsupported. No new provider cards (TypeSafe = Emerging / Signal-only).

Frontier Providers

OpenAI

OpenAI remains the broadest stack for agentic coding and multimodal app building. The current public flagship is GPT-6 Astra (gpt-6-astra, shipped 3–4 Sep; ~1.05M context; reasoning through max; $10/$50 per MTok). GPT-6.1 Sol (29 Sep) is the current mid-tier for complex coding, computer use, and professional work; GPT-6 Sol and GPT-6 Luna (22 Sep) remain served as prior mid and efficient GPT-6 tiers — distinct from GPT-5.6 Sol / Terra / Luna still served. GPT-5.5 / 5.4 still available; Codex, realtime/audio, GPT-Image-2.5, GPT-Live-1, and open-weight gpt-oss round out the stack. Astra is the first Critical cybersecurity model under the Preparedness Framework; advanced cyber remains via Daybreak. Pro is a mode (reasoning.mode=pro), not a separate model ID.

Latest model families
  • GPT-6 Astra (gpt-6-astra, 3–4 Sep) - current flagship for complex reasoning and coding; Responses service_tier: "ultrafast" (29 Sep) for lower-latency Astra (global + US residency; not EU/other regional)
  • GPT-6.1 Sol (gpt-6.1-sol, GA 29 Sep) - near-Astra coding, computer use, and professional work; $2/$10 per MTok (≤272k prompts); 1.05M context; Multi-agent beta on Responses
  • GPT-6 Sol (gpt-6-sol, GA 22 Sep) - prior mid-tier coding and agentic workflows; $2/$10 per MTok; 1.05M context; still documented
  • GPT-6 Luna (gpt-6-luna, GA 22 Sep) - cost-sensitive high-volume; $0.10/$0.50 per MTok; 1.05M context
  • GPT-5.6 Sol / Terra / Luna - prior GPT-5.6 flagship-class family still served (GA 9 Jul; not GPT-6 Sol/Luna)
  • GPT-5.5 / GPT-5.4 / GPT-5.4 pro - earlier frontier tiers still served
  • GPT-5.4 mini / GPT-5.4 nano - low-cost, high-volume tiers
  • GPT-5.3-Codex - agentic coding and repo operations
  • gpt-realtime-2.1 / gpt-realtime-2.1 mini (6 Jul) - realtime voice
  • GPT-Live-1 (10 Sep) - full-duplex voice API; pairs with backend models including Astra
  • Agents API (public beta 10 Sep; still beta) - managed Codex harness; durable sessions, compaction, recovery; OpenAI-hosted or self-hosted sandboxes; MCP/tools; computer use in OpenAI-hosted browser (29 Sep); pairs with gpt-6-astra et al.; US data residency only; no ZDR in beta; not a separate model ID
  • GPT-Transcribe / GPT-Live-Transcribe (28 Jul) - transcription stack
  • GPT-Image-2.5 Flare / Sunburst (8 Sep; gpt-image-2.5-flare / gpt-image-2.5-sunburst) - current image line
  • GPT-Image-2 (21 Apr) - prior image generation
  • gpt-audio-1.5 - audio-in, audio-out workflows
  • gpt-oss-120b / gpt-oss-20b - open-weight reasoning line
  • gpt-realtime-mini, gpt-audio, gpt-audio-mini - deprecated
  • Daybreak / GPT-5.6 Cyber - approval-gated, not self-serve
  • GPT-5.6 Sol Ultrafast - limited preview, not GA
Best for
  • Astra: complex reasoning, coding, and agents
  • GPT-6.1 Sol: mid-tier coding, computer use, and professional work
  • GPT-6 Sol: prior mid-tier coding and agentic workflows
  • GPT-6 Luna: cost-sensitive high-volume throughput
  • GPT-5.6 Sol: prior flagship-class reasoning and agents
  • GPT-5.6 Terra: production balance
  • GPT-5.6 Luna: high-volume throughput (prior family)
  • Image-2.5 Flare: fast image default; Sunburst: precision/editing
  • GPT-Live-1: full-duplex voice
  • Agents API: durable agent sessions, tools, and MCP orchestration
  • Realtime 2.1 voice products
  • Open weights via gpt-oss
  • Cyber workloads only via Daybreak approval
Anthropic

Anthropic's current Claude line spans Claude Opus 5.5 GA (22 Sep), Claude Sonnet 5.5 GA (28 Sep), Fable 5.1 GA (Sep 2026), Sonnet 5 (30 Jun), and Haiku 4.5. Docs still recommend Opus 5.5 as the default for most workloads; Sonnet 5.5 for the best speed/intelligence combination; Fable 5.1 for demanding long-horizon when Opus 5.5 evals fall short. Claude Mythos 5.1 is invitation-only (Project Glasswing); not GA. Opus 5 remains available as legacy. No Haiku 5.5 shipped (roadmap “coming weeks”). Sonnet 4.5 deprecation announced 30 Sep — Claude API retirement 30 Nov 2026; migrate to Sonnet 5.5.

Latest model families
  • Claude Opus 5.5 (claude-opus-5-5, GA 22 Sep) - recommended default for most workloads; long-running agentic coding; $4/$20 MTok
  • Claude Sonnet 5.5 (claude-sonnet-5-5, GA 28 Sep) - best speed/intelligence combination; $2/$10 MTok; 1M context; adaptive thinking
  • Claude Fable 5.1 (claude-fable-5-1, GA Sep 2026) - demanding and long-horizon agents
  • Claude Sonnet 5 - prior production workhorse; still $2/$10
  • Claude Haiku 4.5 - fast, low-latency tier
  • Claude Mythos 5.1 - invitation-only; not GA
  • Claude Opus 5 - legacy tier still served
  • Claude Fable 5 - legacy tier still served
  • Still available: Opus 4.8 / 4.7 / 4.6; Sonnet 4.6 / 4.5
Best for
  • Opus 5.5: default for most workloads
  • Sonnet 5.5: speed/intelligence mid-tier and production agents
  • Fable 5.1: demanding and long-horizon agents when Opus 5.5 evals fall short
  • Sonnet 5: prior production workhorse
  • Haiku 4.5: latency and cost control
Google / Gemini

Google's flagship model family is Gemini. The current public stack centers on Gemini 3.8 Flash GA (Sep 2026), with Gemini 3.7 Flash (GA 13 Aug) still served, Gemini 3.6 Flash and 3.5 Flash-Lite (GA 21 Jul), and 3.1 Pro Preview as the highest text reasoning tier. Gemma 4 is the current open-weight line. Gemini 2.5 Pro / Flash / Flash-Lite remain documented with official shutdown 16 Oct 2026. No Gemini 4 product.

Latest model families
  • Gemini 3.8 Flash (gemini-3.8-flash, GA Sep 2026) - current public workhorse for coding and agents
  • Gemini 3.8 Live (gemini-3.8-live, GA 15 Sep) - low-latency Live API audio-to-audio / voice agents
  • Gemini 3.8 Live Extended Thinking (gemini-3.8-live-extended-thinking, GA 15 Sep) - high-reasoning live audio with background thinking during conversation
  • Gemini 3.8 Flash Cyber - Fairwind-gated; not public GA
  • Gemini 3.7 Flash (GA 13 Aug) - prior flash workhorse, still served
  • Gemini 3.6 Flash (GA 21 Jul) - flash-tier general model
  • Gemini 3.5 Flash-Lite (GA 21 Jul) - cheap subagent tier
  • Gemini 3.5 Flash (GA 19 May) - balanced flash workhorse
  • Gemini 3.5 Transcribe (GA 26 Aug) - speech-to-text; IDs gemini-3.5-transcribe / gemini-3.5-transcribe-live; Chirp 3 successor
  • Gemini 3.1 Pro Preview - highest text reasoning (still Preview)
  • Gemini Omni Flash (GA 27 Aug; gemini-omni-1.1-flash) - video-capable omni model; gemini-omni-flash-preview deprecate date 30 Sep past — use GA ID
  • Antigravity preview (antigravity-preview-09-2026, GA 17 Sep) - agent preview stack; migrate from antigravity-preview-05-2026 (earliest shutdown 5 Oct — ~4 days as of 1 Oct)
  • Nano Banana 2 Lite (30 Jun) - image generation
  • Gemini 3.8 Flash TTS (gemini-3.8-flash-tts, GA 22 Sep) - flagship creative TTS (studio fidelity, dialects, long-form multi-turn)
  • Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts, GA 22 Sep) - fast/cost TTS; intended replacement for gemini-3.1-flash-tts-preview
  • Voices endpoint (/v1beta/voices, GA 22 Sep) - voice design, voice replication, Extended Voice Library (150+)
  • Lyria 3.5 (lyria-3.5, GA 3 Sep) - music generation (text+image → full-length song)
  • Gemma 4 - current open-weight Google family
  • Gemini 2.5 Pro / Flash / Flash-Lite - still documented; shutdown 16 Oct 2026
Best for
  • 3.8 Flash: public workhorse for coding and agents
  • 3.8 Live: low-latency voice agents
  • 3.8 Live Extended Thinking: high-reasoning live audio
  • 3.7 Flash: prior flash tier still served
  • 3.5 Flash-Lite: cheap subagents
  • 3.5 Transcribe: speech-to-text GA (streaming + file)
  • 3.1 Pro Preview: highest text reasoning
  • Omni Flash: video workflows (GA)
  • 3.8 Flash TTS: creative studio-fidelity speech
  • 3.8 Flash-Lite TTS: fast/cost speech synthesis
  • Voices endpoint: voice design and replication
  • Lyria 3.5: music generation (GA)
  • Nano Banana: image generation
  • Gemma 4: open and on-device deployment
SpaceXAI / Grok

Official pages now say SpaceXAI; domains still use x.ai. Same catalog, not a new lab. The Grok-centric stack spans chat, coding, voice, and media APIs. Grok 4.7 (grok-4.7, ~21 Sep, 500k context) is the current flagship for code, agents, and knowledge work; Grok 4.6 (12 Aug) and Grok 4.5 (16 Jul) are still served. Grok 4.7 Fast is the same model at 2× rates — Cursor + Grok Build only, not public xAI API.

Latest model families
  • Grok 4.7 (grok-4.7, ~21 Sep) - current flagship text/code/agents model, 500k context; reasoning low/medium/high(default)/xhigh
  • Grok 4.6 (12 Aug) - prior flagship, still served
  • Grok 4.5 (16 Jul) - prior flagship, still served
  • Grok 4.7 Fast - same model at 2× rates; Cursor + Grok Build only, not public API
  • Grok 4.3 - earlier tier still available
  • Grok Build 0.1 - coding and agentic workflows
  • Voice Think Fast 2.0 (29 Jul) - realtime voice
  • Imagine Image 2.0 (7 Aug) - image generation
  • Voice / Imagine 1.x - prior media tiers
Best for
  • Default text, code, and agents: Grok 4.7
  • Prior-tier workloads: Grok 4.6 / 4.5
  • Voice: Voice Think Fast 2.0
  • Image: Imagine Image 2.0

Tier 1 - Global Scale Leaders

Mistral AI (EU)

Europe's strongest independent model lab. Mistral adds in-region sovereign inference (regional endpoints GA 11 Aug), document intelligence via OCR 4.0 / OCR 4.1, and hosted third-party open models starting with Z.ai GLM 5.2 public preview.

Latest model families
  • Mistral Large 3 - flagship open-weight multimodal model
  • Mistral Medium 3.5 - frontier-class multimodal workhorse
  • Mistral Small 4 - compact general-purpose model
  • Ministral 3 - edge-friendly 3B / 8B / 14B family
  • Shieldstral 1.0 - safety and moderation
  • Leanstral 1.5 - efficient reasoning
  • Codestral - low-latency code generation
  • OCR 4.0 / OCR 4.1 (mistral-ocr-latest) - document intelligence
  • Voxtral Mini Transcribe 2 / Voxtral Mini Transcribe Realtime / Voxtral TTS - audio stack
  • Hosted Z.ai GLM 5.2 - third-party open model (public preview)
  • Magistral Medium/Small 1.2 and Devstral 2 - retired 31 Jul; use Medium 3.5 / Small 4
  • Mistral Medium 3 / 3.1 - retired 31 Aug
Best for
  • In-region and sovereign inference
  • Document intelligence and OCR
  • Multilingual European workloads
  • Coding via Medium 3.5 / Small 4
DeepSeek (China)

DeepSeek remains the cost/performance pressure valve in the market. Current default is DeepSeek-V4.1-Flash (GA 10 Sep; call as deepseek-flash) — native multimodal Causal Encoder–Decoder MoE. V4-Flash and V4-Flash-Vision-Exp are retired as current (compat aliases temporarily route to V4.1-Flash). deepseek-v4-pro (V4-Pro-0813) continues as a separate Pro SKU after 14 Sep with unchanged Pro billing per ops docs/changelog — do not claim Pro→Flash routing as current. No V4.1-Pro. Aliases deepseek-chat / deepseek-reasoner retired 24 Jul 2026.

Latest model families
  • DeepSeek-V4.1-Flash (GA 10 Sep; deepseek-flash) - current default / fast multimodal tier
  • DeepSeek-V4-Pro-0813 (deepseek-v4-pro) - separate Pro tier; continues after 14 Sep with unchanged Pro billing (not routed to Flash)
  • DeepSeek-V4-Flash-0731 / deepseek-v4-flash-vision-exp - retired as current; compat aliases route to V4.1-Flash
  • No V4.1-Pro (referenced only as future)
  • Peak/off-peak pricing continues (new rates effective 10 Sep)
Best for
  • Current default: V4.1-Flash (deepseek-flash)
  • Native multimodal / vision on V4.1-Flash
  • Higher-concurrency Pro workloads via V4-Pro-0813 (separate SKU)
  • Cost control via peak/off-peak pricing
Qwen / Alibaba (China)

One of the broadest stacks anywhere. Qwen spans proprietary flagships, open-weight multimodal models, coding, translation, audio, and image families. Qwen3.8-Max is the current hosted flagship; hosted qwen3.8-flash is now catalog-live on Model Studio. Qwen3.8-Flash-Next (26 Aug) is an open-weights architecture preview toward Qwen4 — not Qwen4 itself.

Latest model families
  • Qwen3.8-Max (hosted GA 3 Aug; qwen3.8-max, 2.4T, 1M) - current flagship
  • Qwen3.8-Flash (qwen3.8-flash, hosted on Model Studio) - catalog-live hosted flash tier
  • Qwen3.8-Flash-Next (open weights 26 Aug) - architecture preview; HF Qwen/Qwen3.8-Flash-Next; not Qwen4
  • Qwen3.8-2.4T-A95B - open weights
  • Qwen3.8-27B Apache VL - open weights; hosted API live
  • Qwen3.7-Plus / Qwen3.7-Flash - current balanced and cheap tiers
  • Qwen-Image-3.0 / Qwen-Audio-3.0 - image and audio stack
  • Qwen3-Coder - agentic coding
  • Qwen3.5-Plus/Flash and Qwen3-Max - legacy
Best for
  • Max as hosted flagship
  • Hosted flash via qwen3.8-flash on Model Studio
  • Self-host efficiency preview via Flash-Next weights
  • Max-class open weights via Qwen3.8
  • Long-horizon agents on 1M context
  • Chinese-English and multilingual workflows
  • Vision, translation, speech, and coding tools
Meta / Llama

The dominant open-weight ecosystem outside China. Meta's current headline models are Llama 4 Scout and Maverick, with Llama Guard 4 as the safety line.

Latest model families
  • Llama 4 Scout - natively multimodal, long-context
  • Llama 4 Maverick - multimodal general-purpose model
  • Llama Guard 4 - safety and policy filtering
  • Llama 3.2 / Llama 3.1 - legacy broad-deployment baseline
Best for
  • Self-hosting and fine-tuning
  • Open-weight multimodal apps
  • Safety filtering and local control
Meta Muse (MSL)

Meta ships two model families: Llama for open-weight foundation models and Muse for agentic, coding, and media workloads. Muse is a separate product line, not a Llama replacement. Spark 1.3 is the recommended agentic/coding tier; max / xhigh reasoning is live in docs. Muse Code is called out of beta on the Meta AI developer hub.

Latest model families
  • Muse Spark 1.3 (muse-spark-1.3, 2 Sep) - recommended for new agentic/coding work via Muse Code + Meta Model API; all reasoning_effort levels including xhigh / max are live; audio not fully supported on 1.3
  • Muse Spark 1.2 - prior hosted API default; docs/pricing lag
  • Muse Voice Transcribe 1.0 (muse-voice-transcribe-1.0, 1 Sep) - STT family; not a Spark replacement
  • Muse Glimmer 30B - Apache open-weight local agent model
  • Muse Code - coding agent (out of beta)
  • Muse Image - image generation (shipped)
  • Muse Video - video generation (preview)
Best for
  • Hosted agentic and coding API via Spark 1.3 (incl. max / xhigh reasoning)
  • Speech-to-text via Muse Voice Transcribe
  • Local open-weight agents via Glimmer
  • Meta-ecosystem image generation

Tier 2 - Enterprise, Regional, and Platform Providers

Cohere

Enterprise-first provider focused on retrieval, grounding, multilingual control, and private deployment. North Mini Code 1.0 (9 Jun, Apache 2.0) extends the stack into agentic coding. Command A+ is the newest open-source workhorse.

Latest model families
  • North Mini Code 1.0 (9 Jun) - 30B-A3B Apache coding MoE
  • Command A+ - newest open-source enterprise workhorse
  • Command A / Command A Reasoning - flagship enterprise models
  • Command A Vision / Command A Translate - multimodal and translation lines
  • Command R7B / Command R+ / Command R - retrieval and grounding family
  • Tiny Aya - live multilingual line
  • Transcribe Arabic - speech recognition
  • Embed / Rerank / Transcribe - retrieval and audio stack
Best for
  • Agentic coding via North Mini Code
  • RAG and enterprise search
  • Grounded assistants
  • Private and sovereign deployments
Amazon Nova / Bedrock

A production platform rather than a single model lab. Bedrock is the control plane; Nova is Amazon's own model line. Current official IDs are Nova 2 Lite, Nova 2 Sonic, and Nova Multimodal Embeddings.

Latest model families
  • Amazon Nova 2 Lite - current text/multimodal tier
  • Nova 2 Sonic - speech and conversational voice
  • Nova Multimodal Embeddings - semantic retrieval
  • Third-party model access via Bedrock
  • Legacy: Premier / Canvas / Reel / Sonic v1 → EOL Sep 2026
  • Command R+ on Bedrock → EOL 19 Aug 2026
Best for
  • AWS-native production apps
  • Enterprise governance and routing
  • Multi-provider deployments
Microsoft Phi

Microsoft's small-model family is optimized for strong performance per parameter and edge use cases. The current work centers on compact reasoning, multimodal SLMs, and vision reasoning.

Latest model families
  • Phi-4-reasoning-vision-15B - compact multimodal reasoning
  • Phi-4-reasoning / Phi-4-reasoning-plus - compact reasoning
  • Phi-4 - small-model flagship
  • Phi-4-mini - compact variant
  • Phi-4-multimodal - vision, audio, text
  • Phi-3.5 - still widely used in light-footprint setups
Best for
  • Small-footprint deployments
  • STEM-heavy tasks
  • Edge and local inference
Microsoft MAI

Microsoft's Foundry model family for enterprise reasoning, coding, image, and speech. MAI is a separate line from Phi; Phi stays the SLM/edge family.

Latest model families
  • MAI-Thinking-1 (public preview 12 Aug) - enterprise reasoning
  • MAI-Code-1-Flash - coding (shipped)
  • MAI-Image-2.5 - image generation (shipped)
  • MAI-Transcribe-1.5 - speech-to-text (shipped)
  • MAI-Voice-2 - text-to-speech (shipped)
  • MAI-Voice-2-Flash - coming soon (not yet shipped)
Best for
  • Enterprise reasoning on Foundry
  • Coding, image, and speech workloads on Foundry
  • Phi remains the choice for SLM and edge use cases
Baidu ERNIE (China)

Baidu's China enterprise stack emphasizes search, multimodal understanding, and document-heavy workflows. ERNIE 5.1 (hosted; launched 8–9 May) joins 5.0 as the native multimodal flagship.

Latest model families
  • ERNIE 5.1 (hosted) - text and agent workloads
  • ERNIE 5.0 / ERNIE 5.0 Thinking / ERNIE 5.0 Preview - native multimodal flagship
  • ERNIE X1.1 / X1.1 Preview
  • ERNIE 4.5 Turbo / ERNIE 4.5 Turbo VL
  • Qianfan OCR: paddleocr-vl-0.9b - document stack
Best for
  • ERNIE 5.1 for text and agent workloads
  • ERNIE 5.0 for full-modal enterprise deployments
  • Chinese enterprise search-augmented assistants
  • Document and multimodal knowledge work
Zhipu AI / GLM (China)

One of China's strongest general-purpose model lines. GLM-5.3 full weights are ungated as of 28 Aug (HF 753.3B; GitHub table 744B-A40B). GLM-5.2 (mid-Jun, 1M, MIT weights) remains the prior open tier. GLM-5.3-Flash (26 Aug) is the first natively multimodal GLM-5 tier with MIT Flash weights. Full GLM-5.3 uses custom glm-5.3 license, not MIT. Do not list GLM-5.4.

Latest model families
  • GLM-5.3-Flash (26 Aug) - first natively multimodal GLM-5; 320B-A18B; hosted glm-5.3-flash; MIT weights zai-org/GLM-5.3-Flash
  • GLM-5.3 (full weights 28 Aug) - text-only agentic coding; ungated open weights; custom glm-5.3 license; HF 753.3B / GitHub 744B-A40B
  • GLM-5.3 (hosted 14 Aug) - same base as 5.2 with post-training only
  • GLM-5.2 (mid-Jun) - 1M context, MIT open weights
  • GLM-5-Turbo - current high-throughput SKU
  • GLM-4.7 / GLM-4.7-FlashX - high-intelligence general line
  • GLM-5V-Turbo / GLM-4.6V-Flash - multimodal coding and vision
  • GLM-Image - optional image generation
Best for
  • Cheap multimodal and visual coding via 5.3-Flash
  • Text agentic coding via 5.3 hosted
  • Agentic and long-horizon coding
  • Chinese-language workflows
  • Multimodal app stacks
Moonshot / Kimi (China)

A fast-moving China-side assistant platform with strong long-context and coding angles. Kimi K3 (hosted 16 Jul; open weights 27 Jul; 2.8T MoE, native vision, 1M) is the current flagship. k2.5 and moonshot-v1* platform sunset completed 31 Aug 2026. No K4.

Latest model families
  • Kimi K3 (hosted 16 Jul; open weights 27 Jul) - 2.8T MoE, native vision, 1M context
  • Kimi K2.7 Code (~12 Jun) - coding-focused tier
  • Kimi K2.6 - prior flagship still in circulation
  • k2.5 and moonshot-v1* - sunset completed 31 Aug 2026
Best for
  • 1M long-context assistants
  • Native multimodal agents
  • First-party Chinese coding and knowledge work
MiniMax (China)

MiniMax spans text, speech, video, and music product lines. M3 (1 Jun; 1M; native image/video) handles agentic coding; MiniMax H3 (31 Jul; omni video, up to 15s / 2K) is the current video line. Official name is MiniMax H3, not Hailuo 3.0.

Latest model families
  • MiniMax M3 (1 Jun) - text and agent model, 1M context, native image/video
  • MiniMax H3 (31 Jul) - omni video, up to 15s / 2K; partial OSS
  • MiniMax Speech 2.8 - current T2A
  • MiniMax Music 3.0 (16 Jul) - music generation; new-user Music APIs closed 20 Aug
Best for
  • Omni video and audio via H3
  • Agentic coding via M3
  • Speech products via Speech 2.8
Tencent Hy (Hunyuan) (China)

Tencent Hy (formerly Hunyuan) spans language, role-play, translation, 3D, and video. Hy3 preview became GA 6 Jul (295B-A21B, 256K, Apache 2.0) as the language flagship. Hy4 preview (Aug 2026) is Apache open weights, not GA. UI-Mate (9B / 27B / democua-27B) is the GUI-agent sub-line on this card.

Latest model families
  • Hy3 (GA 6 Jul) - 295B-A21B, 256K, Apache 2.0 flagship
  • Hy4 preview (Aug 2026) - Apache open weights; not GA
  • UI-Mate (9B / 27B / democua-27B) - GUI and computer-use agents
  • Hunyuan-role-latest - role-play and conversational character model
  • HY-MT2-Pro - translation-focused model
  • HY-3D-3.1 / HY-3D-3.0 - 3D generation
  • HunyuanVideo / HunyuanVideo-Avatar - video generation
Best for
  • GUI and computer-use agents via UI-Mate
  • Hy3 agent and productivity workloads
  • 3D, simulation, and interactive content
  • Media-heavy AI experiences
NVIDIA Nemotron

NVIDIA's open model family for agent workhorses, available via Hugging Face and NIM. Nemotron 3.5 Lightning (11 Aug, 30B-A3B) joins Nano, Super, and Ultra tiers already public.

Latest model families
  • Nemotron 3.5 Lightning (11 Aug) - 30B-A3B open MoE
  • Nemotron Nano / Super / Ultra - prior public tiers
Best for
  • Open agent workhorses
  • Local and DGX deployment
  • NIM-served inference
IBM Granite

IBM's open enterprise model family for on-prem and Apache-licensed agent workloads. Granite 4.2 (25 Aug) adds dense reasoning models with native thinking on/off; Speech 5.0 Turbo CTC covers high-throughput ASR. A separate NC (non-commercial) variant exists — do not present NC as Apache.

Latest model families
  • Granite 4.2 (25 Aug) - dense 3B / 8B / 30B Apache 2.0; native thinking on/off
  • Granite 4.2 30B - SWE and terminal workloads
  • Granite Speech 5.0 Turbo CTC (470M) - high-throughput ASR
  • NC variants - non-commercial license; not Apache
Best for
  • On-prem and Apache enterprise agents
  • Thinking on/off control
  • 30B SWE and terminal tasks
  • Speech 5.0 high-throughput ASR
Perplexity

First-party Sonar and Agent API for web-grounded Q&A and research agents. Sonar Chat Completions support ended 27 Sep 2026. Sync and streaming Sonar requests may still work but are being reformulated as Agent API requests (rolling out by model); async Sonar is not reformulated and is no longer supported — use background mode. Build path: Agent API presets (fast / low / medium / high; xhigh for top deep research). Not to be confused with OpenRouter.

Latest model families
  • Agent API presets - current build path (Sonar tiers map here)
  • Sonar / Sonar Chat Completions - support ended 27 Sep; migrate or use reformulation shim to Agent API
Best for
  • Post-Sonar: Agent API presets for web-grounded Q&A and research agents
  • Web-grounded Q&A with citations
  • Research agents
  • Products needing live web retrieval
Black Forest Labs (FLUX) (EU)

EU-based image and video generation API. FLUX.2 and FLUX 3 Video are live on the BFL API (preview). FLUX 3 Image and Dev weights are not GA.

Latest model families
  • FLUX.2 - image generation on BFL API
  • FLUX 3 Video - video generation on BFL API (preview)
Best for
  • EU image generation API
  • EU video generation API (preview)
Sources: docs.bfl.ai · bfl.ai

Tier 3 - Open-Weight, Local, and Ecosystem

Ollama

Runtime and ecosystem layer, not a model lab. It is the simplest way to run open models locally, switch families, and keep work private.

What it provides
  • Local model runtime and serving
  • Cloud offload for larger models
  • Easy switching between open families
  • Private/offline experimentation
Common families people run
  • gpt-oss
  • Qwen3 / 3.5 / 3.6 / 3.8 / Coder
  • DeepSeek V4.1-Flash / V4-Pro (compat)
  • Gemma 4 / 3 / 3n
  • Llama 4 / 3.x
  • Phi-4
  • Kimi K2.6 / K3
  • Muse Glimmer
  • Nemotron 3.5 Lightning
  • North Mini Code
  • GLM-5.2
  • GLM-5.3-Flash
  • glm-5.3:cloud (cloud-only; not local full-5.3 GGUF)
  • Qwen3.8-Flash-Next
  • Granite 4.2
Best for
  • Local prototyping
  • Privacy-first workflows
  • Developer experimentation
Hugging Face

The distribution and discovery layer for open models. Baseten was added as an Inference Provider (6 Aug). Hugging Face is the hub, not an HF-owned model family.

What it provides
  • Model hub and hosting
  • Open-weight distribution
  • Community evaluation and tooling
  • Baseten Inference Provider (added 6 Aug)
Best for
  • Model discovery
  • Open-model release tracking
  • Community experimentation

Emerging model class — System One / decision models

TypeSafe AI (Emerging · System One)

TypeSafe AI (San Francisco) exited stealth 15 Sep 2026 with System One Models — machine-native decision models, not chat LLMs. First public model Jev (jev-1.13.0) returns typed Choice/Score/Noul answers with calibrated probabilities via POST /v1/systemone. Trained with company method RLCD (Reinforcement Learning for Calibrated Decisions). Early access / waitlist; public docs at docs.typesafe.ai. Seed ~$40M led by DCVC (company IR / FinSMEs). The company calls this “frontier”; Twiniti treats it as a new model class, not a Frontier LLM peer.

Current model / product families (early access)
  • System One Models (product class) — typed structured decisions for software, not string generation
  • Jev 1.13 / jev-1.13.0 (flagship; docs as of Sep 2026)
  • Primitives: Choice, Score, Noul
Best for
  • In-software semantic judgment, routing, classification, and confidence-gated automation
  • Low-latency typed decisions inside agents and workflows (code stays in control)
  • Parallel fan-out over structured questions (company claim)
  • Not for: chat, long-form generation, coding assistants, or multimodal (text/JSON only today per docs)