Models

41 models under one key. A single price per 1M tokens — input and output cost the same, billed by usage.

Modalities
TextImagesAudioVideoPDFText
1M context
NewFlagship

OpenAI's new flagship: GPT-6 Astra. The strongest model in the lineup for complex reasoning, coding and agentic workloads. 1M-token context.

by openai|Sep 5, 2026|1M context|$0.489/M input tokens|$0.489/M output tokens|gpt-6-astra
1M context
FastCodinglowmediumhigh

Google's newest fast model with separate low, medium and high reasoning modes, multimodal input and a 1M-token context.

by google|Sep 2, 2026|1M context|$0.075/M input tokens|$0.075/M output tokens|gemini-3.8-flash

Z.ai: GLM 5.3

Reasoning
1.05M context
Coding

A large Z.ai reasoning model for complex engineering and long agentic runs: 1M-token context, up to 131K output. Reasoning is always on (low / high / max, max by default), with better token economy than GLM-5.2.

by z-ai|Aug 18, 2026|1.05M context|$0.143/M input tokens|$0.143/M output tokens|glm-5.3
1.05M context
FastCoding

A fast, low-cost version of GLM 5.3: the same 1M-token context and up to 131K output, but significantly cheaper and with lower latency. Built for high-throughput agentic runs, code completion and bulk text processing.

by z-ai|Aug 18, 2026|1.05M context|$0.045/M input tokens|$0.045/M output tokens|glm-5.3-flash
1M context
FastCodinglowmediumhigh

Google's fast Antigravity model with separate low, medium and high reasoning modes, multimodal input and a 1M-token context.

by google|Aug 14, 2026|1M context|$0.075/M input tokens|$0.075/M output tokens|gemini-3.7-flash

xAI: Grok 4.6

Reasoning
500K context
STEM

xAI's strongest model: frontier results in code, knowledge work and STEM. Accepts text, images and files (including PDF), 500K-token context. Widely used in agentic developer tooling.

by x-ai|Aug 12, 2026|500K context|$0.060/M input tokens|$0.060/M output tokens|grok-4.6
1M context

DeepSeek's large MoE model, the general-availability V4 Pro release. 1M-token context, served both by DeepSeek and third-party inference providers; strong results in reasoning, code and agentic benchmarks.

by deepseek|Aug 12, 2026|1M context|$0.045/M input tokens|$0.045/M output tokens|deepseek-v4-pro
Cheaplowmediumhigh

An open lightweight MoE model: 3B active parameters out of 30B total, built for high throughput in agentic workloads and domain fine-tuning. 1M-token context, part of NVIDIA's open model family for agentic AI.

by nvidia|Aug 11, 2026|1M context|$0.0038/M input tokens|$0.0038/M output tokens|nemotron-3.5-lightning
1M context
Multimodal

The flagship of Alibaba's Qwen3.8 series: a multimodal reasoning model for hard tasks, visual understanding, code and agentic work. 1M-token context and top placements in the Artificial Analysis intelligence, coding and agentic indexes.

by qwen|Aug 3, 2026|1M context|$0.023/M input tokens|$0.023/M output tokens|qwen3.8-max
1M context
Open weights

Moonshot AI's open natively multimodal model with 2.8T parameters — the first open model in the 3T class. Built on Kimi Delta Attention and Attention Residuals with a 1M-token context; designed for long-horizon code, knowledge work and reasoning.

by moonshotai|Jul 17, 2026|1M context|$0.414/M input tokens|$0.414/M output tokens|kimi-k3
1M context
CodingAgents

The flagship of the GPT-5.6 series: complex reasoning, coding and agentic workloads. Especially strong on multi-step terminal tasks and long chains of code work. 1M-token context, world knowledge up to February 2026.

by openai|Jul 9, 2026|1M context|$0.188/M input tokens|$0.188/M output tokens|gpt-5.6-sol
1M context
Balanced

The balanced GPT-5.6 tier between the flagship Sol and the budget Luna. Built for everyday coding, reasoning and agentic tasks at roughly half the price of Sol. 1M-token context.

by openai|Jul 9, 2026|1M context|$0.090/M input tokens|$0.090/M output tokens|gpt-5.6-terra
1M context
Fast

The fastest and cheapest tier of the GPT-5.6 family for high-throughput workloads. Supports reasoning-effort control and prompt caching, just like the larger models in the series.

by openai|Jul 9, 2026|1M context|$0.017/M input tokens|$0.017/M output tokens|gpt-5.6-luna

xAI: Grok 4.5

Reasoning
500K context

A frontier xAI model with strong results in code, knowledge work and STEM. 500K-token context with step-by-step reasoning. Widely used in coding assistants and autonomous agent frameworks.

by x-ai|Jul 8, 2026|500K context|$0.060/M input tokens|$0.060/M output tokens|grok-4.5
1M context
FastCodinglowmediumhigh

Google's new fast Antigravity model with separate low, medium and high reasoning modes, multimodal input and a 1M-token context.

by google|Jul 2, 2026|1M context|$0.075/M input tokens|$0.075/M output tokens|gemini-3.6-flash
1M context
Popular

The most agentic Sonnet model: it closes the gap to Opus 4.8 in reasoning, tool use and code at a markedly lower price. Hallucinates and sycophantically agrees less often than Sonnet 4.6.

by anthropic|Jun 30, 2026|1M context|$0.150/M input tokens|$0.150/M output tokens|claude-sonnet-5
262K context
Coding

The coding model of the Kimi K2 family, based on K2.6: a 1T-parameter MoE with 32B active (384 experts, 8 per token) and the MoonViT vision encoder. 262K context, always in reasoning mode with the chain preserved between turns.

by moonshotai|Jun 12, 2026|262K context|$0.113/M input tokens|$0.113/M output tokens|kimi-k2.7-code
1M context
Open weights

An open frontier model for reasoning and orchestration: 550B parameters, 55B active, a hybrid Transformer-Mamba MoE architecture with LatentMoE and native speculative decoding via MTP. Up to 1M context, tuned for long agentic workflows.

by nvidia|Jun 4, 2026|1M context|$0.0038/M input tokens|$0.0038/M output tokens|nemotron-3-ultra
Cheap

The cost-efficient Qwen3.7 model with text and image input. Vision is substantially stronger while agentic intelligence is preserved: it reads scenes and interfaces, generates code from visual references and can navigate mobile apps.

by qwen|Jun 3, 2026|1M context|$0.0075/M input tokens|$0.0075/M output tokens|qwen3.7-plus
1M context
Multimodal

MiniMax's multimodal model: text, image and video input with a 1M-token context. Built on MiniMax Sparse Attention with KV-block selection instead of full attention — roughly 20× less compute per token at long context.

by minimax|May 31, 2026|1M context|$0.034/M input tokens|$0.034/M output tokens|minimax-m3
1M context
Coding

An evolution of Opus 4.7: better code, more reliable agentic runs, roughly 4× less likely to let faulty code pass without comment. Adds user-set effort levels and Claude Code "dynamic workflows" with parallel subagents.

by anthropic|May 28, 2026|1M context|$0.301/M input tokens|$0.301/M output tokens|claude-opus-4-8
1M context

The Qwen3.7 flagship: text in, text out, 1M-token context. Aimed at agentic workloads — coding, office tasks and long autonomous execution — with prompt caching support.

by qwen|May 21, 2026|1M context|$0.015/M input tokens|$0.015/M output tokens|qwen3.7-max
Coding

Cursor's own agentic coding model, trained from a Moonshot Kimi K2.5 checkpoint with extra RL on long-horizon tasks. The Fast variant is the default interactive mode, tuned for tool use, multi-file edits and the terminal.

by cursor|May 18, 2026|256K context|$0.038/M input tokens|$0.038/M output tokens|composer-2.5-fast
1M context
FastMultimodallowmediumhigh

Google's fast, solid-quality model: good accuracy, high speed and a 1M-token context. Simpler than the Pro tier, but it answers faster and costs less.

by google|May 6, 2026|1M context|$0.075/M input tokens|$0.075/M output tokens|gemini-3.5-flash
400K context
Agents

OpenAI's spring 2026 flagship: the most agentic model in the lineup at launch, 82.7% on Terminal-Bench 2.0, strong at computer use and knowledge work. Same speed per token as GPT-5.4, but it spends fewer tokens on equivalent Codex tasks.

by openai|Apr 23, 2026|400K context|$0.188/M input tokens|$0.188/M output tokens|gpt-5.5
1M context
Cheap

DeepSeek's cost-efficient MoE model: 284B parameters, 13B active, hybrid attention for cheap long-context work up to 1M tokens. Supports high and xhigh reasoning effort — a fit for fast assistants and high-volume agents.

by deepseek|Apr 23, 2026|1M context|$0.0075/M input tokens|$0.0075/M output tokens|deepseek-v4-flash
1M context
Agents

Xiaomi's flagship trillion-parameter agentic model: strong results on ClawEval, GDPVal and SWE-bench Pro. Sustains coherent autonomous execution across nearly 1,000 tool calls, with up to 1M context and 128K output.

by xiaomi|Apr 22, 2026|1M context|$0.023/M input tokens|$0.023/M output tokens|mimo-v2.5-pro
Cheap

Xiaomi's natively omni-modal model: Pro-level agentic performance at roughly half the inference cost, and better image and video understanding than MiMo-V2-Omni. The 1M-token context fits whole documents and long conversations.

by xiaomi|Apr 22, 2026|1M context|$0.0038/M input tokens|$0.0038/M output tokens|mimo-v2.5
400K context

A fast, cost-efficient version of GPT-5.4 for high throughput. Handles text and images, works confidently with tools — a fit for chat apps, coding assistants and high-volume agentic workloads.

by openai|Mar 17, 2026|400K context|$0.060/M input tokens|$0.060/M output tokens|gpt-5.4-mini
203K context
Fast

A Z.ai model for fast inference in agentic environments with long execution chains. 203K context, with improved decomposition of complex instructions, tool use and stability across long runs.

by z-ai|Mar 15, 2026|203K context|$0.104/M input tokens|$0.104/M output tokens|glm-5-turbo
1M context
MultimodalReasoninglowhigh

Google's strong high-end general-purpose model: high accuracy, deep reasoning and a 1M-token context. Accepts text, images, audio, video and PDF, and replies with text.

by google|Mar 10, 2026|1M context|$0.150/M input tokens|$0.150/M output tokens|gemini-3.1-pro
1M context

The model that merged the Codex and GPT lines into one system: 922K input context, up to 128K output, text and image input. Improved coding, document understanding, tool use and instruction following.

by openai|Mar 5, 2026|1M context|$0.113/M input tokens|$0.113/M output tokens|gpt-5.4
1M context

The first Opus-class model with a 1M-token context (beta): plans more carefully and sustains agentic tasks longer across large codebases. SOTA on Terminal-Bench 2.0 and Humanity's Last Exam at release.

by anthropic|Feb 5, 2026|1M context|$0.301/M input tokens|$0.301/M output tokens|claude-opus-4-6
Audio

OpenAI's speech recognition model: it takes an audio file and returns text via the /v1/audio/transcriptions route. A fit for transcribing calls, interviews and voice messages.

by openai|Mar 20, 2025|16K context|$0.075/M input tokens|$0.075/M output tokens|gpt-4o-transcribe
200K context
Top tier

Anthropic's top price tier ($10 / $50 per 1M at the vendor) — the same rate as the limited-availability Claude Mythos 5. Shipped with stricter cybersecurity safeguards than the later Sonnet models.

by anthropic|200K context|$0.602/M input tokens|$0.602/M output tokens|claude-fable-5
1M context
Flagship

The Opus-class flagship: $5 / $25 per 1M at the vendor — the same rate as Opus 4.5–4.8. Claude Sonnet 5 is reported to approach Opus 5, which makes the latter the accuracy benchmark of the lineup.

by anthropic|1M context|$0.301/M input tokens|$0.301/M output tokens|claude-opus-5
1M context

An Opus-class model, the predecessor of Opus 4.8, vendor price $5 / $25 per 1M. Used as the comparison baseline in benchmarks for code, computer use and legal agents.

by anthropic|1M context|$0.301/M input tokens|$0.301/M output tokens|claude-opus-4-7
1M context

The predecessor of Claude Sonnet 5, vendor price $3 / $15 per 1M. Served as the baseline against which Sonnet 5 showed gains in agentic reasoning, tool use, code and safety.

by anthropic|1M context|$0.150/M input tokens|$0.150/M output tokens|claude-sonnet-4-6
Fast

Anthropic's fast, low-cost model, vendor price $1 / $5 per 1M. Part of the Claude 4.5 generation that introduced regional and multi-region endpoints on Bedrock and Google Cloud.

by anthropic|200K context|$0.068/M input tokens|$0.068/M output tokens|claude-haiku-4-5

Z.ai: GLM 5.2

Reasoning
1M context

A Z.ai reasoning model with a 1M-token context for long agentic scenarios, project-scale development and multi-step automation. Supports high and xhigh reasoning effort (xhigh being the maximum).

by z-ai|1M context|$0.143/M input tokens|$0.143/M output tokens|glm-5.2
mu256K context
Cheap

A light, very cheap model for high-volume work: classification, data extraction, short answers and throughput-heavy scenarios with a minimal cost per token.

by muse|256K context|$0.0038/M input tokens|$0.0038/M output tokens|muse-spark-1.2

Every model has a single price per 1M tokens — input and output cost the same, billed by actual usage, no subscriptions or minimum charges. Billed proportionally: 100K tokens is a tenth of the listed amount. Ruble amounts use an internal rate of ₽100 per $1.