AI hardware23 accelerators
Accelerator generations and prices
The chips the era is built on, by vendor and generation, with indicative unit prices and cloud rates. Vendors rarely publish list prices; every figure here is approximate.
Data compiled as of July 2026. Corrections by pull request are welcome.
23/
NVIDIA8
| Chip | Year | Memory | Bandwidth | Compute | Unit price (approx.) | Cloud rate (approx.) |
|---|---|---|---|---|---|---|
V100 (Volta) Introduced tensor cores; trained the GPT-2/GPT-3-era models. | 2017 | 16–32 GB HBM2 | 900 GB/s | 125 TFLOPS FP16 | ~$10,000 | ~$0.30–2/hr |
A100 (Ampere) The workhorse of the ChatGPT moment; GPT-4 reportedly trained on ~25k A100s. | 2020 | 40–80 GB HBM2e | 2.0 TB/s | 312 TFLOPS BF16 | ~$10,000–15,000 | ~$0.75–3/hr |
H100 (Hopper) The scarcest commodity of 2023–24; secondary-market prices have since fallen sharply as Blackwell ramped. | 2022 | 80 GB HBM3 | 3.35 TB/s | ~1.0 PFLOPS FP8 | ~$25,000–40,000 new; used units now ~$10,000–20,000 | ~$2–8/hr |
H200 Memory-expanded Hopper; nearly doubled inference throughput on large models. | 2024 | 141 GB HBM3e | 4.8 TB/s | ~1.0 PFLOPS FP8 | ~$30,000–40,000 | ~$3.50–9/hr |
B200 (Blackwell) Dual-die design with FP4 inference; sold primarily as GB200 NVL72 racks (~$3–3.4M each). | 2024 | 180–192 GB HBM3e | 8 TB/s | ~4.5 PFLOPS FP8 / 9 PFLOPS FP4 | ~$30,000–40,000 | ~$4–12/hr |
GB300 / B300 (Blackwell Ultra) Mid-cycle refresh for reasoning-model inference; now the volume Blackwell part as Rubin ramps. | 2025 | 288 GB HBM3e | 8 TB/s | ~1.5× B200 (FP4) | rack-scale (~$3.7–4M per NVL72; ~$40–55k per GPU) | ~$5–18/hr |
VR200 (Vera Rubin) Sampling now, volume H2 2026; HBM4 costs (~$500+ per stack) pushed rack prices to roughly double GB300. | 2026 | 288 GB HBM4 | ~13 TB/s | ~50 PFLOPS FP4 (per dual-die package) | rack-scale (~$7.8–8.8M per NVL72; ~$55k per GPU est.) | — |
Groq 3 LPX (LPU) Product of NVIDIA’s ~$20B Groq licensing acquihire (Dec 2025); deterministic low-latency decode engine paired with Rubin racks. | 2026 | 500 MB SRAM per LPU | 150 TB/s (on-chip SRAM) | ~315 PFLOPS per 256-LPU rack | rack-scale; pricing not yet public | — |
Google4
| Chip | Year | Memory | Bandwidth | Compute | Unit price (approx.) | Cloud rate (approx.) |
|---|---|---|---|---|---|---|
TPU v4 Optical circuit switching across 4,096-chip pods; trained PaLM. | 2021 | 32 GB HBM2 | 1.2 TB/s | 275 TFLOPS BF16 | not sold; cloud only | ~$2–3/hr |
TPU v5e / v5p Split the line into efficiency (v5e) and performance (v5p) tiers; trained Gemini 1.0. | 2023 | 16 / 95 GB HBM | 0.8 / 2.8 TB/s | 197 / 459 TFLOPS BF16 | not sold; cloud only | ~$1.20 / $4.20/hr |
TPU v6 (Trillium) 4.7× v5e per-chip performance; trained Gemini 2.x. | 2024 | 32 GB HBM | 1.6 TB/s | ~926 TFLOPS BF16 | not sold; cloud only | ~$2.70/hr |
TPU v7 (Ironwood) GA on Google Cloud April 2026; Anthropic contracted for up to ~1M chips. TPU v8 splits into Sunfish (training) and Zebrafish (inference) for ~2027. | 2025 | 192 GB HBM3e | 7.4 TB/s | ~4.6 PFLOPS FP8 | not sold; cloud only | ~$5.40–12/hr (commit vs. on-demand; large contracts far lower) |
AMD4
| Chip | Year | Memory | Bandwidth | Compute | Unit price (approx.) | Cloud rate (approx.) |
|---|---|---|---|---|---|---|
MI250X (CDNA 2) Powered Frontier, the first exascale supercomputer; little LLM traction. | 2021 | 128 GB HBM2e | 3.2 TB/s | 383 TFLOPS FP16 | ~$12,000–15,000 | — |
MI300X (CDNA 3) AMD’s first credible H100 rival; adopted by Microsoft, Meta, and OpenAI for inference. | 2023 | 192 GB HBM3 | 5.3 TB/s | ~1.3 PFLOPS FP8 | ~$15,000–20,000 | ~$2–4/hr |
MI355X (CDNA 4) Anchor of OpenAI’s 6-gigawatt AMD partnership announced October 2025. | 2025 | 288 GB HBM3e | 8 TB/s | ~5 PFLOPS FP8 / 10 PFLOPS FP4 | ~$25,000–30,000 | ~$2.60–8.60/hr |
MI455X (CDNA 5) / Helios AMD’s NVL72 answer: 72 GPUs, 31 TB HBM4, open Ethernet fabric; claims ~30% more tokens per dollar than Rubin. Ships late 2026. | 2026 | 432 GB HBM4 | ~19.6 TB/s | ~40 PFLOPS FP4 / 20 PFLOPS FP8 | rack-scale (~$5–5.5M per 72-GPU Helios rack) | — |
Amazon2
| Chip | Year | Memory | Bandwidth | Compute | Unit price (approx.) | Cloud rate (approx.) |
|---|---|---|---|---|---|---|
Trainium 2 Anthropic’s Project Rainier cluster runs on hundreds of thousands of these chips. | 2024 | 96 GB HBM3 | 2.9 TB/s | ~650 TFLOPS FP8 | not sold; cloud only | ~$1–2/hr (per chip, Trn2 instances) |
Trainium 3 AWS’s first 3nm chip, GA December 2025; UltraServers scale to 144 chips, with Anthropic and Bedrock running production workloads. | 2025 | 144 GB HBM3e | 4.9 TB/s | ~2.5 PFLOPS FP8 | not sold; cloud only | ~$1.80/hr (per chip, est.) |
Microsoft1
| Chip | Year | Memory | Bandwidth | Compute | Unit price (approx.) | Cloud rate (approx.) |
|---|---|---|---|---|---|---|
Maia 200 Microsoft’s first competitive inference silicon (Jan 2026); claims ~3× the FP4 of Trainium 3 or Ironwood per chip. | 2026 | 216 GB HBM3e | 7 TB/s | ~5 PFLOPS FP8 / 10 PFLOPS FP4 | not sold; internal Azure only | — |
Meta1
| Chip | Year | Memory | Bandwidth | Compute | Unit price (approx.) | Cloud rate (approx.) |
|---|---|---|---|---|---|---|
MTIA v2 / "Iris" Runs Meta’s ads and recommendation inference; the GenAI-focused "Iris" chip ramps late 2026 on a roughly six-month silicon cadence. | 2026 | 128 GB LPDDR5 (v2) | ~0.2 TB/s (off-chip, v2) | ~354 TOPS INT8 (v2) | not sold; internal only | — |
Intel1
| Chip | Year | Memory | Bandwidth | Compute | Unit price (approx.) | Cloud rate (approx.) |
|---|---|---|---|---|---|---|
Gaudi 3 Priced aggressively but found little adoption; Intel shelved Falcon Shores and pivoted to the Crescent Island inference GPU, sampling H2 2026. | 2024 | 128 GB HBM2e | 3.7 TB/s | ~1.8 PFLOPS FP8 | ~$16,000 (8-chip board ~$125k) | — |
Cerebras1
| Chip | Year | Memory | Bandwidth | Compute | Unit price (approx.) | Cloud rate (approx.) |
|---|---|---|---|---|---|---|
WSE-3 A single wafer-scale chip; IPO’d May 2026 (~$23B target) on the back of a ~$20B, 750 MW OpenAI inference agreement. | 2024 | 44 GB on-wafer SRAM | 21 PB/s (on-chip) | ~125 PFLOPS FP16 (sparse) | ~$2–3M per CS-3 system | — |
SambaNova1
| Chip | Year | Memory | Bandwidth | Compute | Unit price (approx.) | Cloud rate (approx.) |
|---|---|---|---|---|---|---|
SN40L (RDU) Dataflow architecture for trillion-parameter inference; JPMorgan deployment and a $350M Series E (Feb 2026) after Intel acquisition talks collapsed. | 2023 | 64 GB HBM + 1.5 TB DDR (three-tier) | ~1.6 TB/s (HBM) | 688 TFLOPS BF16 | sold as full systems/cloud; pricing not public | — |
Prices are indicative street or list prices per accelerator at launch in USD. Real transactions vary widely with volume, form factor, and packaging; rack-scale systems are priced as systems, not chips. Cloud rates are typical on-demand per-accelerator-hour figures across major and specialist providers.