Model tracking45 releases
Frontier and notable model releases
A running record of the models that defined the field: who released them, when, under what access terms, and why they mattered.
Data compiled as of July 2026. Corrections by pull request are welcome.
45/
2026
| Model | Developer | Released | Access | Context | Significance |
|---|---|---|---|---|---|
| Claude Opus 5 | Anthropic | Jul 2026 | Proprietary | 1M | Opus-tier Claude 5 release; close to Fable 5 on agentic coding at half the price, with an optional faster “fast mode”. |
| Kimi K3 | Moonshot AI | Jul 2026 | Open weights | 1M | 2.8T-parameter MoE, the largest open-weight release to date; weights published July 27 under a bespoke (non-OSI) license. |
| GPT-5.6 (Sol, Terra, Luna) | OpenAI | Jul 2026 | Proprietary | 1M | Three-tier family distilled from a single base run; replaced GPT-5.5 across ChatGPT and the API. |
| Grok 4.5 | xAI (SpaceX) | Jul 2026 | Proprietary | 500K | First Grok built specifically for agentic coding; first major release since xAI folded into SpaceX in February 2026. |
| Claude Sonnet 5 | Anthropic | Jun 2026 | Proprietary | 1M | Default mid-tier Claude 5 model; near-Opus-4.8 performance at lower prices, and the new free-tier model on Claude.ai. |
| GLM-5.2 | Zhipu AI (Z.ai) | Jun 2026 | Open weights | 1M | 753B MoE under an MIT license; at release the strongest open-weight coding model, at a fraction of frontier API prices. |
| Claude Fable 5 / Mythos 5 | Anthropic | Jun 2026 | Proprietary | 1M | First public Mythos-class model, a tier above Opus; access was suspended for most of June under a U.S. export-control order before being restored; the less-restricted Mythos 5 variant is limited to approved organizations. |
| Qwen3.7-Max | Alibaba | May 2026 | Proprietary | 1M | API-only flagship built for long-horizon agent runs; marked Alibaba’s shift away from open weights for its top tier. |
| Gemini 3.5 Flash | Google DeepMind | May 2026 | Proprietary | — | I/O 2026 workhorse that beat Gemini 3.1 Pro on coding and agentic benchmarks; shipped while the 3.5 Pro flagship was repeatedly delayed. |
| DeepSeek V4 (Pro, Flash) | DeepSeek | Apr 2026 | Open weights | 1M | MIT-licensed MoE pair (V4-Pro 1.6T/49B active, V4-Flash 284B); previewed in April, stable in July; its thinking mode absorbed the R1 reasoning line. |
| GPT-5.5 | OpenAI | Apr 2026 | Proprietary | — | Agent-focused flagship pitched as a step toward an AI “super app”; GPT-5.5 Instant became the ChatGPT default in May. |
| Muse Spark | Meta | Apr 2026 | Proprietary | — | First model from Meta Superintelligence Labs and Meta’s first closed-weight flagship, ending its open-release strategy. |
| GPT-5.4 | OpenAI | Mar 2026 | Proprietary | — | Mainline GPT-5 refresh on OpenAI’s quickened release cadence; its xHigh reasoning tier led agentic tool-use benchmarks. |
| Gemini 3.1 Pro | Google DeepMind | Feb 2026 | Proprietary | 1M | Incremental Pro flagship update; native multimodal input and output across text, image, audio, video, and code. |
| Claude Opus 4.6–4.8 | Anthropic | Feb 2026 | Proprietary | — | Rapid-cadence Opus point releases (4.6 in February, 4.7 in April, 4.8 in May) bridging the fourth generation to Claude 5. |
2025
| Model | Developer | Released | Access | Context | Significance |
|---|---|---|---|---|---|
| GPT-5.2 | OpenAI | Dec 2025 | Proprietary | 400K | Instant, Thinking, and Pro variants; OpenAI’s rapid answer to Gemini 3, with stronger long-context and enterprise coding. |
| Mistral Large 3 | Mistral AI | Dec 2025 | Open weights | 256K | 675B MoE (41B active) multimodal flagship under Apache 2.0; Europe’s strongest open-weight release. |
| Claude Opus 4.5 | Anthropic | Nov 2025 | Proprietary | 200K | Flagship coding and agentic model; substantial price cut from prior Opus tiers. |
| Gemini 3 Pro | Google DeepMind | Nov 2025 | Proprietary | 1M | Led most reasoning and multimodal benchmarks at launch; deep integration into Search. |
| GPT-5.1 | OpenAI | Nov 2025 | Proprietary | 400K | Refinement of GPT-5 with adaptive reasoning effort and improved instruction following. |
| Claude Sonnet 4.5 | Anthropic | Sep 2025 | Proprietary | 200K–1M | Mid-tier flagship focused on long-horizon agentic coding; 1M-token context in beta. |
| GPT-5 | OpenAI | Aug 2025 | Proprietary | 400K | Unified router across fast and reasoning modes; replaced the GPT-4 line in ChatGPT. |
| gpt-oss-120b / 20b | OpenAI | Aug 2025 | Open weights | — | OpenAI’s first open-weight release since GPT-2; Apache-2.0 licensed MoE reasoning models. |
| Grok 4 | xAI | Jul 2025 | Proprietary | 256K | Reasoning-first flagship trained on the Colossus cluster; strong benchmark results. |
| Kimi K2 | Moonshot AI | Jul 2025 | Open weights | — | 1T-parameter MoE (32B active); at release the strongest open agentic/coding model. |
| Claude Opus 4 / Sonnet 4 | Anthropic | May 2025 | Proprietary | 200K | Fourth-generation flagships; extended thinking with tool use during reasoning. |
| Qwen 3 | Alibaba | Apr 2025 | Open weights | — | Dense and MoE family (0.6B–235B) with switchable thinking mode; broadly adopted base for fine-tunes. |
| Llama 4 (Scout, Maverick) | Meta | Apr 2025 | Open weights | up to 10M (Scout) | Natively multimodal MoE family; mixed reception relative to open-weight competitors. |
| o3 / o4-mini | OpenAI | Apr 2025 | Proprietary | 200K | Reasoning models with full tool use during chain of thought; o3 topped many benchmarks at launch. |
| Gemini 2.5 Pro | Google DeepMind | Mar 2025 | Proprietary | 1M | Thinking-by-default flagship; long-context multimodal reasoning at competitive prices. |
| Claude 3.7 Sonnet | Anthropic | Feb 2025 | Proprietary | 200K | First hybrid reasoning model: a single model with a controllable extended-thinking budget. |
| DeepSeek-R1 | DeepSeek | Jan 2025 | Open weights | 128K | Open reasoning model rivaling o1 at a much lower reported training cost; triggered a market-wide repricing of AI capex assumptions. |
2024
| Model | Developer | Released | Access | Context | Significance |
|---|---|---|---|---|---|
| DeepSeek-V3 | DeepSeek | Dec 2024 | Open weights | 128K | 671B MoE (37B active); the reported ~$5.6M compute cost covered only the final training run and is widely considered to understate total cost; basis for R1. |
| o1 | OpenAI | Dec 2024 | Proprietary | 200K | First production reasoning model line (previewed September 2024); established test-time compute scaling. |
| Gemini 2.0 Flash | Google DeepMind | Dec 2024 | Proprietary | 1M | Fast agentic workhorse model; native tool use and multimodal output. |
| Claude 3.5 Sonnet | Anthropic | Jun 2024 | Proprietary | 200K | Outperformed the larger Opus tier; the October update added computer use — the first frontier agent able to operate a GUI. |
| GPT-4o | OpenAI | May 2024 | Proprietary | 128K | Natively multimodal (“omni”) flagship with real-time voice; free-tier default in ChatGPT. |
| Llama 3 / 3.1 | Meta | Apr 2024 | Open weights | 128K (3.1) | Llama 3.1 405B (July 2024) was the first open-weight model at rough parity with frontier proprietary models. |
| Claude 3 (Opus, Sonnet, Haiku) | Anthropic | Mar 2024 | Proprietary | 200K | Three-tier family; Opus was the first model to clearly match GPT-4. |
| Gemini 1.5 Pro | Google DeepMind | Feb 2024 | Proprietary | 1M–2M | Broke the long-context barrier with near-perfect million-token recall. |
2023
| Model | Developer | Released | Access | Context | Significance |
|---|---|---|---|---|---|
| Mixtral 8x7B | Mistral AI | Dec 2023 | Open weights | 32K | Sparse MoE that beat much larger dense models; mainstreamed MoE in open models. |
| GPT-4 Turbo | OpenAI | Nov 2023 | Proprietary | 128K | Cheaper, faster GPT-4 with 128K context; launched alongside the GPT Store and Assistants API. |
| Llama 2 | Meta | Jul 2023 | Open weights | 4K | First openly licensed Llama for commercial use; seeded the open-model ecosystem. |
| GPT-4 | OpenAI | Mar 2023 | Proprietary | 8K–32K | Defined the frontier for over a year; passed professional exams that stumped GPT-3.5. |
2022
| Model | Developer | Released | Access | Context | Significance |
|---|---|---|---|---|---|
| ChatGPT (GPT-3.5) | OpenAI | Nov 2022 | Proprietary | 4K | The consumer breakout: 100M users in two months; started the current investment cycle. |