AI Model Pricing

What every major AI model actually costs, per million tokens.

Pulled from each provider’s official pricing page and reviewed by hand before publishing — not an unattended scraper. Last verified 2026-07-27.

Workload cost calculator

Estimate monthly cost across every tracked model for a workload you describe -- start from a preset, then adjust the numbers to match your actual usage.

Short conversational replies with a system prompt and recent chat history.

Processing

Batch rates aren’t modeled yet for these providers — real-time pricing is used either way.

Requires tool calling
Requires reasoning
Required modalities
ModelProviderMonthly costFitInputOutputContextSpeedStatus
Qwen3.7 Flashtiered
Qwen (Alibaba Model Studio)$2.45Compatible$0.03$0.131Mfastproduction
DeepSeek V4 Flash
DeepSeek$6.96Compatible$0.14$0.281Mfastproduction
GPT-5.4 Nano
OpenAI$21.15Compatible$0.2$1.25400Kfastproduction
Qwen3.7 Plustiered
Qwen (Alibaba Model Studio)$21.38Compatible$0.276$1.101Mstandardproduction
DeepSeek V4 Pro
DeepSeek$21.45Compatible$0.435$0.871Mstandardproduction
Qwen3.6 Flashtiered
Qwen (Alibaba Model Studio)$25.65Compatible$0.25$1.501Mfastproduction
Moonshot V1 8K
Kimi (Moonshot)$37.00Compatible$0.2$2.008Kstandardproduction
Moonshot V1 8K Visionpreview
Kimi (Moonshot)$37.00Compatible$0.2$2.008Kstandardpreview
Gemini 2.5 Flash
Google$39.53Compatible$0.3$2.501Mfastproduction
Kimi K2.5
Kimi (Moonshot)$55.50Compatible$0.6$3.00262Kstandardproduction
Grok Build 0.1tieredpreview
xAI$56.20Compatible$1.00$2.00256Kfastpreview
Grok 4.20 0309 Non-Reasoningtiered
xAI$68.45Compatible$1.25$2.501Mfastproduction
Grok 4.20 0309 Reasoningtiered
xAI$68.45Compatible$1.25$2.501Mstandardproduction
Grok 4.20 Multi-Agent 0309tieredpreview
xAI$68.45Compatible$1.25$2.501Mslowpreview
Grok 4.3tiered
xAI$68.45Compatible$1.25$2.501Mstandardproduction
GPT-5.4 Mini
OpenAI$76.95Compatible$0.75$4.50400Kfastproduction
Kimi K2.6
Kimi (Moonshot)$78.56Compatible$0.95$4.00262Kstandardproduction
Kimi K2.7 Code
Kimi (Moonshot)$79.64Compatible$0.95$4.00262Kstandardproduction
Moonshot V1 32K
Kimi (Moonshot)$97.50Compatible$1.00$3.0033Kstandardproduction
Moonshot V1 32K Visionpreview
Kimi (Moonshot)$97.50Compatible$1.00$3.0033Kstandardpreview
GPT-5.6 Lunatiered
OpenAI$102.60Compatible$1.00$6.001.1Mfastproduction
Qwen3.7 Max
Qwen (Alibaba Model Studio)$107.43Compatible$1.65$4.951Mstandardproduction
Grok 4.5tiered
xAI$133.80Compatible$2.00$6.00500Kstandardproduction
Kimi K2.7 Code HighSpeed
Kimi (Moonshot)$159.28Compatible$1.90$8.00262Kfastproduction
Moonshot V1 128K
Kimi (Moonshot)$182.50Compatible$2.00$5.00131Kstandardproduction
Moonshot V1 128K Visionpreview
Kimi (Moonshot)$182.50Compatible$2.00$5.00131Kstandardpreview
GPT-5.3 Codex
OpenAI$223.30Compatible$1.75$14.00400Kstandardproduction
GPT-5.4tiered
OpenAI$256.50Compatible$2.50$15.001.1Mstandardproduction
GPT-5.6 Terratiered
OpenAI$256.50Compatible$2.50$15.001.1Mfastproduction
Kimi K3
Kimi (Moonshot)$270.30Compatible$3.00$15.001.0Mslowproduction
Chat Latest
OpenAI$513.00Compatible$5.00$30.00400Kfastproduction
GPT-5.5tiered
OpenAI$513.00Compatible$5.00$30.001.1Mstandardproduction
GPT-5.6 Soltiered
OpenAI$513.00Compatible$5.00$30.001.1Mstandardproduction
GPT-5.4 Protiered
OpenAI$4,050.00Compatible$30.00$180.001.1Mslowproduction
GPT-5.5 Protiered
OpenAI$4,050.00Compatible$30.00$180.001.1Mslowproduction
Ministral 3 — 3B
Mistral$7.25Partial / unknown$0.1$0.1fastproduction
Gemini 2.5 Flash-Lite
Google$7.76Partial / unknown$0.1$0.4fastproduction
Devstral Small 2preview
Mistral$9.75Partial / unknown$0.1$0.3fastpreview
Ministral 3 — 8B
Mistral$10.88Partial / unknown$0.15$0.15fastproduction
Mistral NeMo
Mistral$10.88Partial / unknown$0.15$0.15fastproduction
Ministral 3 — 14B
Mistral$14.50Partial / unknown$0.2$0.2fastproduction
Mistral Small 4
Mistral$16.50Partial / unknown$0.15$0.6fastproduction
Codestral
Mistral$29.25Partial / unknown$0.3$0.9fastproduction
Gemini 3.5 Flash-Lite
Google$39.53Partial / unknown$0.3$2.50fastproduction
Magistral Small
Mistral$48.75Partial / unknown$0.5$1.50standardproduction
Mistral Large 3
Mistral$48.75Partial / unknown$0.5$1.50standardproduction
Devstral 2
Mistral$49.00Partial / unknown$0.4$2.00standardproduction
Mixtral 8x7B
Mistral$50.75Partial / unknown$0.7$0.7standardproduction
Claude Haiku 4.5
Anthropic$90.10Partial / unknown$1.00$5.00fastproduction
Gemini 3.6 Flash
Google$135.15Partial / unknown$1.50$7.50fastproduction
Gemini 3.5 Flash
Google$153.90Partial / unknown$1.50$9.00fastproduction
Gemini 2.5 Protiered
Google$159.50Partial / unknown$1.25$10.00standardproduction
Claude Sonnet 5 (introductory)
Anthropic$180.20Partial / unknown$2.00$10.00fastproduction
Magistral Medium
Mistral$182.50Partial / unknown$2.00$5.00slowproduction
Mistral Medium 3.5
Mistral$183.75Partial / unknown$1.50$7.50standardproduction
Mixtral 8x22B
Mistral$195.00Partial / unknown$2.00$6.00standardproduction
Gemini 3.1 Pro Previewtieredpreview
Google$205.20Partial / unknown$2.00$12.00standardpreview
Claude Sonnet 5 (standard)
Anthropic$270.30Partial / unknown$3.00$15.00fastproduction
Claude Opus 5
Anthropic$450.50Partial / unknown$5.00$25.00standardproduction
Claude Fable 5
Anthropic$901.00Partial / unknown$10.00$50.00slowproduction
Claude Opus 4.1
Anthropic$1,351.50Partial / unknown$15.00$75.00slowdeprecated
Claude Haiku 3.5
Anthropic$72.08Incompatible$0.8$4.00fastretired

Assumptions

  • Standard published prices are used; negotiated or enterprise pricing is not reflected.
  • Provider-specific tiers (request-length, promotional, regional) are not yet calculated -- tiered models are flagged, but the estimate uses their base rate.
  • Batch discounts are only applied where a structured batch rate exists in the data -- currently none, so batch mode uses real-time rates.
  • Tool-invocation and media-processing charges are excluded from this estimate.
  • Real-world token usage and retries vary; treat this as a planning estimate, not a bill.
  • Estimates exclude taxes and regional pricing uplifts.

Anthropic

USD · per 1M tokens
ModelStatusInputCached inputOutputContextSource
Claude Fable 5

Highest-capability long-running agents and demanding knowledge work

textimagetool callingreasoningslow
production$10.00$1.00$50.00link
Claude Haiku 3.5

Fast, low-cost text and vision workloads

textimagetool callingfast

Retired except on Bedrock/Google Cloud

retired$0.8$0.08$4.00link
Claude Haiku 4.5

High-volume, latency-sensitive agents and assistants

textimagetool callingreasoningfast
production$1.00$0.1$5.00link
Claude Opus 4.1

Complex analysis, coding, and long-running agent tasks

textimagetool callingreasoningslow
deprecated$15.00$1.50$75.00link
Claude Opus 5

Complex agentic coding and enterprise knowledge work

textimagetool callingreasoningstandard
production$5.00$0.5$25.00link
Claude Sonnet 5 (introductory)

Balanced production coding, agents, and enterprise workflows

textimagetool callingreasoningfast

Introductory pricing, ends 2026-08-31

production$2.00$0.2$10.00link
Claude Sonnet 5 (standard)

Balanced production coding, agents, and enterprise workflows

textimagetool callingreasoningfast
production$3.00$0.3$15.00link

DeepSeek

USD · per 1M tokens
ModelStatusInputCached inputOutputContextSource
DeepSeek V4 Flash

Fast, economical reasoning and straightforward agent tasks

texttool callingreasoningfast
production$0.14$0.0028$0.281Mlink
DeepSeek V4 Pro

Advanced reasoning, STEM, coding, and agentic workflows

texttool callingreasoningstandard
production$0.435$0.0036$0.871Mlink

Google

USD · per 1M tokens
ModelStatusInputCached inputOutputContextSource
Gemini 2.5 Flash

Fast multimodal processing and high-volume agent workflows

textimageaudiovideotool callingreasoningfast

Text/image/video input; audio input billed at $1.00

production$0.3$0.03$2.501Mlink
Gemini 2.5 Flash-Lite

Low-cost, high-throughput classification and extraction

textimageaudiovideotool callingreasoningfast

Text/image/video input; audio input billed at $0.30

production$0.1$0.01$0.4link
Gemini 2.5 Pro

Complex reasoning, coding, and long multimodal documents

textimageaudiovideotool callingreasoningstandard

Price rises with prompt length; shown value is the lower tier -- see source for the upper tier ($2.50 / $0.25 / $15.00)

production$1.25$0.125$10.00link
Gemini 3.1 Pro Preview

Frontier multimodal reasoning and complex agentic tasks

textimageaudiovideotool callingreasoningstandard

Price rises with prompt length; shown value is the lower tier -- see source for the upper tier ($4.00 / $0.40 / $18.00)

preview$2.00$0.2$12.00link
Gemini 3.5 Flash

Agentic coding and sustained multimodal workflows

textimageaudiovideotool callingreasoningfast
production$1.50$0.15$9.00link
Gemini 3.5 Flash-Lite

Lowest-cost high-throughput execution in the Gemini 3.5 family

textimageaudiovideotool callingreasoningfast
production$0.3$0.03$2.50link
Gemini 3.6 Flash

Balanced agentic and multimodal tasks with efficient token use

textimageaudiovideotool callingreasoningfast
production$1.50$0.15$7.50link

Kimi (Moonshot)

USD · per 1M tokens
ModelStatusInputCached inputOutputContextSource
Kimi K2.5

Cost-efficient multimodal reasoning and agent tasks

textimagevideotool callingreasoningstandard

Kimi notes that web_search is currently being updated and does not recommend it in the near term.

production$0.6$0.1$3.00262Klink
Kimi K2.6needs review

Coding agents, visual understanding, and flexible thinking workloads

textimagevideotool callingreasoningstandard

Kimi notes that web_search is currently being updated and does not recommend it in the near term.

production$0.95$0.16$4.00262Klink
Kimi K2.7 Code

Long-context programming, deep reasoning, and coding agents

textimagevideotool callingreasoningstandard
production$0.95$0.19$4.00262Klink
Kimi K2.7 Code HighSpeed

High-speed long-context programming and coding agents

textimagevideotool callingreasoningfast

Same model as Kimi K2.7 Code, served at approximately 180 tokens/s and up to 260 tokens/s for short contexts.

production$1.90$0.38$8.00262Klink
Kimi K3needs review

Long-horizon coding, deep reasoning, and end-to-end knowledge work

textimagevideotool callingreasoningslow

Kimi notes that web_search is currently being updated and does not recommend it in the near term.

production$3.00$0.3$15.001.0Mlink
Moonshot V1 128K

Long-document text generation and analysis

textstandard
production$2.00$5.00131Klink
Moonshot V1 128K Vision

Image understanding with long-document context

textimagestandard
preview$2.00$5.00131Klink
Moonshot V1 32K

Medium-length documents and general text generation

textstandard
production$1.00$3.0033Klink
Moonshot V1 32K Vision

Image understanding with medium-length context

textimagestandard
preview$1.00$3.0033Klink
Moonshot V1 8K

Short-context general text generation

textstandard
production$0.2$2.008Klink
Moonshot V1 8K Vision

Short-context image understanding and text generation

textimagestandard
preview$0.2$2.008Klink

Mistral

USD · per 1M tokens
ModelStatusInputCached inputOutputContextSource
Codestral

Low-latency code completion, fill-in-the-middle, and generation

textfast

Premier low-latency coding model.

production$0.3$0.9link
Devstral 2

Autonomous software engineering and agentic coding

texttool callingstandard

Open-weights agentic coding model.

production$0.4$2.00link
Devstral Small 2

Lightweight multimodal coding agents

textimagetool callingfast

Labs model for lightweight coding agents.

preview$0.1$0.3link
Magistral Medium

Domain-specific, transparent, and multilingual reasoning

textimagetool callingreasoningslow

Premier thinking model.

production$2.00$5.00link
Magistral Small

Cost-efficient multilingual reasoning and multimodal analysis

textimagetool callingreasoningstandard

Premier lightweight thinking model.

production$0.5$1.50link
Ministral 3 — 14B

Higher-quality edge inference and compact agent workloads

texttool callingfast

Open lightweight edge model.

production$0.2$0.2link
Ministral 3 — 3B

Small edge deployments and lightweight agents

texttool callingfast

Open lightweight edge model.

production$0.1$0.1link
Ministral 3 — 8B

Balanced edge inference and lightweight agents

texttool callingfast

Open lightweight edge model.

production$0.15$0.15link
Mistral Large 3

General-purpose multilingual and multimodal enterprise workloads

textimagetool callingstandard

Open-weight flagship multimodal and multilingual model.

production$0.5$1.50link
Mistral Medium 3.5

Enterprise reasoning, coding, agents, and multimodal work

textimagetool callingreasoningstandard

Open model. State-of-the-art performance with simplified enterprise deployment.

production$1.50$7.50link
Mistral NeMo

Low-cost lightweight coding workloads

textfast

Open lightweight model trained specifically for code tasks.

production$0.15$0.15link
Mistral Small 4

Cost-efficient multilingual, multimodal, and agentic workloads

textimagetool callingfast

Open Apache 2.0 model.

production$0.15$0.6link
Mixtral 8x22B

Higher-quality open-weight general text generation

textstandard

Open sparse Mixture-of-Experts model.

production$2.00$6.00link
Mixtral 8x7B

Lightweight general text generation with open weights

textstandard

Open sparse Mixture-of-Experts model.

production$0.7$0.7link

OpenAI

USD · per 1M tokens
ModelStatusInputCached inputOutputContextSource
Chat Latest

Testing the latest Instant model experience used in ChatGPT

textimagetool callingfast

Latest Instant model used in ChatGPT; its underlying snapshot is regularly updated. OpenAI recommends GPT-5.6 for production API usage.

production$5.00$0.5$30.00400Klink
GPT-5.3 Codex

Long-running software engineering and agentic coding

textimagetool callingreasoningstandard

Specialized Codex model for agentic software engineering.

production$1.75$0.175$14.00400Klink
GPT-5.4

General complex reasoning, coding, and professional workflows

textimagetool callingreasoningstandard

Affordable model for coding and professional work. Prompts over 272K input tokens use long-context pricing for the full request: $5 input, $0.5 cached input, $22.5 output per 1M. Eligible regional-processing endpoints add 10%.

production$2.50$0.25$15.001.1Mlink
GPT-5.4 Mini

Cost-efficient production agents and routine reasoning tasks

textimagetool callingreasoningfast

Strong mini model for coding, computer use, and subagents. Eligible regional-processing endpoints add 10%.

production$0.75$0.075$4.50400Klink
GPT-5.4 Nano

High-volume classification, extraction, and simple automation

textimagetool callingreasoningfast

Lowest-cost GPT-5.4-class model for simple high-volume tasks. Eligible regional-processing endpoints add 10%.

production$0.2$0.02$1.25400Klink
GPT-5.4 Pro

Highest-quality GPT-5.4 reasoning for difficult professional work

textimagetool callingreasoningslow

Higher-compute GPT-5.4 model for maximum response quality. Prompts over 272K input tokens use long-context pricing for the full request: $60 input, $270 output per 1M. Eligible regional-processing endpoints add 10%.

production$30.00$180.001.1Mlink
GPT-5.5

Complex production reasoning, coding, and tool-driven workflows

textimagetool callingreasoningstandard

Frontier coding and professional-work model. Prompts over 272K input tokens use long-context pricing for the full request: $10 input, $1 cached input, $45 output per 1M. Eligible regional-processing endpoints add 10%.

production$5.00$0.5$30.001.1Mlink
GPT-5.5 Pro

Maximum-quality GPT-5.5 reasoning and complex professional tasks

textimagetool callingreasoningslow

Higher-compute GPT-5.5 model for maximum response quality. Prompts over 272K input tokens use long-context pricing for the full request: $60 input, $270 output per 1M. Eligible regional-processing endpoints add 10%.

production$30.00$180.001.1Mlink
GPT-5.6 Luna

Cost-sensitive, high-volume production workloads

textimagetool callingreasoningfast

GPT-5.6 model for cost-sensitive, high-volume workloads. Prompts over 272K input tokens use long-context pricing for the full request: $2 input, $0.2 cached input, $2.5 cache write, $9 output per 1M. Eligible regional-processing endpoints add 10%.

production$1.00$0.1$6.001.1Mlink
GPT-5.6 Sol

Frontier complex reasoning, coding, and professional work

textimagetool callingreasoningstandard

Frontier GPT-5.6 model for complex professional work. Prompts over 272K input tokens use long-context pricing for the full request: $10 input, $1 cached input, $12.5 cache write, $45 output per 1M. Eligible regional-processing endpoints add 10%.

production$5.00$0.5$30.001.1Mlink
GPT-5.6 Terra

Strong intelligence with balanced cost and latency

textimagetool callingreasoningfast

GPT-5.6 model balancing intelligence and cost. Prompts over 272K input tokens use long-context pricing for the full request: $5 input, $0.5 cached input, $6.25 cache write, $22.5 output per 1M. Eligible regional-processing endpoints add 10%.

production$2.50$0.25$15.001.1Mlink

Qwen (Alibaba Model Studio)

USD · per 1M tokens
ModelStatusInputCached inputOutputContextSource
Qwen3.6 Flash

Low-cost multimodal understanding and fast agent workflows

textimagevideotool callingreasoningfast

International list price shown for requests up to 256K tokens. Above 256K through 1M: $1 input and $4 output. Explicit cache hits cost 10%; implicit hits cost 20%.

production$0.25$0.025$1.501Mlink
Qwen3.7 Flash

Ultra-low-cost, high-throughput reasoning and text generation

texttool callingreasoningfast

International list price shown for requests up to 32K tokens. 32K-256K: $0.10 input/$0.40 output; 256K-1M: $0.20 input/$0.80 output. Explicit cache hits cost 10%; implicit hits cost 20%.

production$0.03$0.003$0.131Mlink
Qwen3.7 Max

Highest-capability Qwen reasoning and text generation

texttool callingreasoningstandard

Global list price through 1M tokens. Explicit cache hits cost 10% and implicit cache hits 20% of input list price. Time-limited night/day discounts may apply.

production$1.65$0.165$4.951Mlink
Qwen3.7 Plus

Balanced multimodal reasoning, long video, and tool use

textimagevideotool callingreasoningstandard

Global list price shown for requests up to 256K tokens. Above 256K through 1M: $0.826 input and $3.301 output. Explicit cache hits cost 10%; implicit hits cost 20%. Time-limited night/day discounts may apply.

production$0.276$0.0276$1.101Mlink

xAI

USD · per 1M tokens
ModelStatusInputCached inputOutputContextSource
Grok 4.20 0309 Non-Reasoning

Fast general-purpose responses and tool-driven workflows

textimagetool callingfast

Short-context prices shown. At prompts of 200K tokens or more, all tokens cost $2.5 input, $0.4 cached input, and $5 output per 1M. Batch API discount: 20%. Priority processing costs 2x standard rates.

production$1.25$0.2$2.501Mlink
Grok 4.20 0309 Reasoning

Deep reasoning, analysis, and tool-driven workflows

textimagetool callingreasoningstandard

Short-context prices shown. At prompts of 200K tokens or more, all tokens cost $2.5 input, $0.4 cached input, and $5 output per 1M. Batch API discount: 20%. Priority processing costs 2x standard rates.

production$1.25$0.2$2.501Mlink
Grok 4.20 Multi-Agent 0309

Parallel multi-agent deep research and complex investigations

textimagetool callingreasoningslow

Short-context prices shown. At prompts of 200K tokens or more, all tokens cost $2.5 input, $0.4 cached input, and $5 output per 1M. Batch API discount: 20%. Priority processing costs 2x standard rates.

preview$1.25$0.2$2.501Mlink
Grok 4.3

General reasoning and tool-driven production workflows

textimagetool callingreasoningstandard

Short-context prices shown. At prompts of 200K tokens or more, all tokens cost $2.5 input, $0.4 cached input, and $5 output per 1M. Batch API discount: 20%. Priority processing costs 2x standard rates.

production$1.25$0.2$2.501Mlink
Grok 4.5

Agentic software engineering and complex coding workflows

textimagetool callingreasoningstandard

Short-context prices shown. At prompts of 200K tokens or more, all tokens cost $4 input, $0.6 cached input, and $12 output per 1M. Priority processing costs 2x standard rates.

production$2.00$0.3$6.00500Klink
Grok Build 0.1

Fast agentic coding, debugging, web development, and MCP

texttool callingfast

Short-context prices shown. At prompts of 200K tokens or more, all tokens cost $2 input, $0.4 cached input, and $4 output per 1M. Priority processing costs 2x standard rates.

preview$1.00$0.2$2.00256Klink

Sources: OpenAI, Anthropic, Google, DeepSeek, Qwen (Alibaba Model Studio), Kimi (Moonshot), Mistral, xAI. Prices are informational and may lag behind provider changes between refreshes — always confirm against the linked source before billing decisions.