All Models
Browse all available AI models — compare pricing, features, and start integrating.

DeepSeek V4.1 Flash is a new 552B MoE with 8B to 16B active parameters, 1M context, native image input, and the top Terminal-Bench and DeepSWE scores in the DeepSeek line. AIReiter bills it below the official off-peak rate with no peak-hour surcharge.

OpenAI frontier model for complex reasoning, coding, and long-context work.

Build with Gemini 3.8 Flash for software engineering, autonomous agents, and complex enterprise tasks.

Claude Fable 5.1 keeps Fable 5 pricing on input and output, cuts cache-read cost to a quarter, and is built for long-running agentic coding and document work.

GLM-5.3 Flash reads both text and images with a 1M-token context and 128K output, priced for workloads you run thousands of times a day. GLM-5.3 itself is text-only.

Gemini 3.6 Flash combines speed and intelligence for advanced reasoning, coding, multimodal understanding, and agentic workflows.

Multimodal intelligence for understanding diverse information.

Kimi K3 is a long-context reasoning model available through an OpenAI-compatible Chat Completions API, suitable for code understanding, writing, analysis, and agent-style workflows.

Deep reasoning model for complex problem solving.

High-performance multimodal AI for diverse applications.

Claude Opus 5 is suited for demanding reasoning, advanced coding, analysis, long-form writing, and agent workflows through the Anthropic Messages API.

Claude Fable 5 is a premium model for deep reasoning, complex analysis, long-form creation, and demanding agent workflows.

Claude Sonnet 5 balances intelligence, speed, and cost for advanced reasoning, coding, writing, and everyday professional work.

Claude Opus 4.8 is designed for demanding reasoning, coding, analysis, and professional knowledge workflows.

Advanced reasoning for deep analysis and complex thinking.

Balanced AI assistant with speed and accuracy.

Reliable AI assistant with strong overall performance.

Use GLM-5.3 for repository-scale coding, bilingual analysis, agent review, and recommendations that need explicit tradeoffs.

Use Gemini 3.7 Flash for responsive chat, focused coding help, document processing, and frequent API workloads.

Use Grok 4.5 for code review, debugging, technical planning, and stable agent workflows already validated on this version.

Evaluate Grok 4.6 for implementation planning, complex debugging, code modernization, and engineering agent workflows.

Use DeepSeek V4 Pro when a prompt needs deeper reasoning and a careful final recommendation.

Use DeepSeek V4 Flash as the first-pass route for cost-sensitive technical traffic.

Use Kimi K2.7 Code when code context is the bottleneck.

Use Doubao Seed 2.1 Turbo for fast Chinese-language production traffic.

Use GLM 5.2 when reasoning trace quality and structured conclusions matter.

Use MiniMax M3 when documents or agent memory need to stay in one request.

Frontier AI for advanced reasoning and complex intelligence tasks.

Frontier AI for advanced reasoning and complex intelligence tasks.

Powerful reasoning model for analysis and problem solving.

Next-generation AI for understanding and creative generation.

Qwen3.8 Max is a 1M-context flagship. Use Chat Completions for dialogue and PDF; use Responses for web_search, code_interpreter, and image search.