OpenAI's new flagship: GPT-6 Astra. The strongest model in the lineup for complex reasoning, coding and agentic workloads. 1M-token context.
gpt-6-astra41 models under one key. A single price per 1M tokens — input and output cost the same, billed by usage.
OpenAI's new flagship: GPT-6 Astra. The strongest model in the lineup for complex reasoning, coding and agentic workloads. 1M-token context.
gpt-6-astraGoogle's newest fast model with separate low, medium and high reasoning modes, multimodal input and a 1M-token context.
gemini-3.8-flashA large Z.ai reasoning model for complex engineering and long agentic runs: 1M-token context, up to 131K output. Reasoning is always on (low / high / max, max by default), with better token economy than GLM-5.2.
glm-5.3A fast, low-cost version of GLM 5.3: the same 1M-token context and up to 131K output, but significantly cheaper and with lower latency. Built for high-throughput agentic runs, code completion and bulk text processing.
glm-5.3-flashGoogle's fast Antigravity model with separate low, medium and high reasoning modes, multimodal input and a 1M-token context.
gemini-3.7-flashxAI's strongest model: frontier results in code, knowledge work and STEM. Accepts text, images and files (including PDF), 500K-token context. Widely used in agentic developer tooling.
grok-4.6DeepSeek's large MoE model, the general-availability V4 Pro release. 1M-token context, served both by DeepSeek and third-party inference providers; strong results in reasoning, code and agentic benchmarks.
deepseek-v4-proAn open lightweight MoE model: 3B active parameters out of 30B total, built for high throughput in agentic workloads and domain fine-tuning. 1M-token context, part of NVIDIA's open model family for agentic AI.
nemotron-3.5-lightningThe flagship of Alibaba's Qwen3.8 series: a multimodal reasoning model for hard tasks, visual understanding, code and agentic work. 1M-token context and top placements in the Artificial Analysis intelligence, coding and agentic indexes.
qwen3.8-maxMoonshot AI's open natively multimodal model with 2.8T parameters — the first open model in the 3T class. Built on Kimi Delta Attention and Attention Residuals with a 1M-token context; designed for long-horizon code, knowledge work and reasoning.
kimi-k3The flagship of the GPT-5.6 series: complex reasoning, coding and agentic workloads. Especially strong on multi-step terminal tasks and long chains of code work. 1M-token context, world knowledge up to February 2026.
gpt-5.6-solThe balanced GPT-5.6 tier between the flagship Sol and the budget Luna. Built for everyday coding, reasoning and agentic tasks at roughly half the price of Sol. 1M-token context.
gpt-5.6-terraThe fastest and cheapest tier of the GPT-5.6 family for high-throughput workloads. Supports reasoning-effort control and prompt caching, just like the larger models in the series.
gpt-5.6-lunaA frontier xAI model with strong results in code, knowledge work and STEM. 500K-token context with step-by-step reasoning. Widely used in coding assistants and autonomous agent frameworks.
grok-4.5Google's new fast Antigravity model with separate low, medium and high reasoning modes, multimodal input and a 1M-token context.
gemini-3.6-flashThe most agentic Sonnet model: it closes the gap to Opus 4.8 in reasoning, tool use and code at a markedly lower price. Hallucinates and sycophantically agrees less often than Sonnet 4.6.
claude-sonnet-5The coding model of the Kimi K2 family, based on K2.6: a 1T-parameter MoE with 32B active (384 experts, 8 per token) and the MoonViT vision encoder. 262K context, always in reasoning mode with the chain preserved between turns.
kimi-k2.7-codeAn open frontier model for reasoning and orchestration: 550B parameters, 55B active, a hybrid Transformer-Mamba MoE architecture with LatentMoE and native speculative decoding via MTP. Up to 1M context, tuned for long agentic workflows.
nemotron-3-ultraThe cost-efficient Qwen3.7 model with text and image input. Vision is substantially stronger while agentic intelligence is preserved: it reads scenes and interfaces, generates code from visual references and can navigate mobile apps.
qwen3.7-plusMiniMax's multimodal model: text, image and video input with a 1M-token context. Built on MiniMax Sparse Attention with KV-block selection instead of full attention — roughly 20× less compute per token at long context.
minimax-m3An evolution of Opus 4.7: better code, more reliable agentic runs, roughly 4× less likely to let faulty code pass without comment. Adds user-set effort levels and Claude Code "dynamic workflows" with parallel subagents.
claude-opus-4-8The Qwen3.7 flagship: text in, text out, 1M-token context. Aimed at agentic workloads — coding, office tasks and long autonomous execution — with prompt caching support.
qwen3.7-maxCursor's own agentic coding model, trained from a Moonshot Kimi K2.5 checkpoint with extra RL on long-horizon tasks. The Fast variant is the default interactive mode, tuned for tool use, multi-file edits and the terminal.
composer-2.5-fastGoogle's fast, solid-quality model: good accuracy, high speed and a 1M-token context. Simpler than the Pro tier, but it answers faster and costs less.
gemini-3.5-flashOpenAI's spring 2026 flagship: the most agentic model in the lineup at launch, 82.7% on Terminal-Bench 2.0, strong at computer use and knowledge work. Same speed per token as GPT-5.4, but it spends fewer tokens on equivalent Codex tasks.
gpt-5.5DeepSeek's cost-efficient MoE model: 284B parameters, 13B active, hybrid attention for cheap long-context work up to 1M tokens. Supports high and xhigh reasoning effort — a fit for fast assistants and high-volume agents.
deepseek-v4-flashXiaomi's flagship trillion-parameter agentic model: strong results on ClawEval, GDPVal and SWE-bench Pro. Sustains coherent autonomous execution across nearly 1,000 tool calls, with up to 1M context and 128K output.
mimo-v2.5-proXiaomi's natively omni-modal model: Pro-level agentic performance at roughly half the inference cost, and better image and video understanding than MiMo-V2-Omni. The 1M-token context fits whole documents and long conversations.
mimo-v2.5A fast, cost-efficient version of GPT-5.4 for high throughput. Handles text and images, works confidently with tools — a fit for chat apps, coding assistants and high-volume agentic workloads.
gpt-5.4-miniA Z.ai model for fast inference in agentic environments with long execution chains. 203K context, with improved decomposition of complex instructions, tool use and stability across long runs.
glm-5-turboGoogle's strong high-end general-purpose model: high accuracy, deep reasoning and a 1M-token context. Accepts text, images, audio, video and PDF, and replies with text.
gemini-3.1-proThe model that merged the Codex and GPT lines into one system: 922K input context, up to 128K output, text and image input. Improved coding, document understanding, tool use and instruction following.
gpt-5.4The first Opus-class model with a 1M-token context (beta): plans more carefully and sustains agentic tasks longer across large codebases. SOTA on Terminal-Bench 2.0 and Humanity's Last Exam at release.
claude-opus-4-6OpenAI's speech recognition model: it takes an audio file and returns text via the /v1/audio/transcriptions route. A fit for transcribing calls, interviews and voice messages.
gpt-4o-transcribeAnthropic's top price tier ($10 / $50 per 1M at the vendor) — the same rate as the limited-availability Claude Mythos 5. Shipped with stricter cybersecurity safeguards than the later Sonnet models.
claude-fable-5The Opus-class flagship: $5 / $25 per 1M at the vendor — the same rate as Opus 4.5–4.8. Claude Sonnet 5 is reported to approach Opus 5, which makes the latter the accuracy benchmark of the lineup.
claude-opus-5An Opus-class model, the predecessor of Opus 4.8, vendor price $5 / $25 per 1M. Used as the comparison baseline in benchmarks for code, computer use and legal agents.
claude-opus-4-7The predecessor of Claude Sonnet 5, vendor price $3 / $15 per 1M. Served as the baseline against which Sonnet 5 showed gains in agentic reasoning, tool use, code and safety.
claude-sonnet-4-6Anthropic's fast, low-cost model, vendor price $1 / $5 per 1M. Part of the Claude 4.5 generation that introduced regional and multi-region endpoints on Bedrock and Google Cloud.
claude-haiku-4-5A Z.ai reasoning model with a 1M-token context for long agentic scenarios, project-scale development and multi-step automation. Supports high and xhigh reasoning effort (xhigh being the maximum).
glm-5.2A light, very cheap model for high-volume work: classification, data extraction, short answers and throughput-heavy scenarios with a minimal cost per token.
muse-spark-1.2Every model has a single price per 1M tokens — input and output cost the same, billed by actual usage, no subscriptions or minimum charges. Billed proportionally: 100K tokens is a tenth of the listed amount. Ruble amounts use an internal rate of ₽100 per $1.