Supported models
Every model below is available through one LLMRPM key on a daily request budget — GPT, Claude, Gemini, DeepSeek, Qwen, and more. Compatible with the OpenAI and Anthropic SDKs.
LLMGPT
LLMGPT is a managed multi-account GPT route that automatically selects an available subscribed Codex model and reports the model actually used.
GPT-5.6 Luna
GPT-5.6 Luna from OpenAI, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 400K-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
GPT-5.6 Sol
GPT-5.6 Sol from OpenAI, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 400K-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
GPT-5.6 Terra
GPT-5.6 Terra from OpenAI, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 400K-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
GPT-5.5
GPT-5.5 from OpenAI, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 400K-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
GPT-5.4
GPT-5.4 from OpenAI, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Claude Opus 4.6
Claude Opus 4.6 from Anthropic, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Claude Opus 4.7
Claude Opus 4.7 from Anthropic, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Claude Opus 4.8
Claude Opus 4.8 from Anthropic, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Claude Sonnet 4.6
Claude Sonnet 4.6 from Anthropic, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Claude Sonnet 5
Claude Sonnet 5 from Anthropic, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Claude Haiku 4.5
Claude Haiku 4.5 from Anthropic, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 200K-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Gemini 3 Flash (Preview)
Gemini 3 Flash (Preview) from Google, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1.05M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Gemini 3.1 Flash Lite (Preview)
Gemini 3.1 Flash Lite (Preview) from Google, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1.05M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Gemini 3.1 Pro (Preview)
Gemini 3.1 Pro (Preview) from Google, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Gemini 3.5 Flash
Gemini 3.5 Flash from Google, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Gemini 2.5 Flash
Gemini 2.5 Flash from Google, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1.05M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Gemini 2.5 Flash Lite
Gemini 2.5 Flash Lite from Google, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1.05M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 from DeepSeek, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
DeepSeek V4 Pro
DeepSeek V4 Pro from DeepSeek, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Qwen 3.5 Plus
Qwen 3.5 Plus from Alibaba Cloud, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 128K-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Qwen 3.6 Plus
Qwen 3.6 Plus from Alibaba Cloud, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Qwen 3.7 Max
Qwen 3.7 Max from Alibaba Cloud, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Qwen 3.7 Plus
Qwen 3.7 Plus from Alibaba Cloud, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Kimi K2.5
Kimi K2.5 from Moonshot AI, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 128K-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Kimi K2.6
Kimi K2.6 from Moonshot AI, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 256K-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Kimi K2.7 Code
Kimi K2.7 Code from Moonshot AI, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 262K-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
GLM-5
GLM-5 from Zhipu AI, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 200K-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
GLM-5.1
GLM-5.1 from Zhipu AI, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 200K-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
GLM-5.2
GLM-5.2 from Zhipu AI, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1.05M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
MiniMax M2.5
MiniMax M2.5 from MiniMax, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 200K-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
MiniMax M2.7
MiniMax M2.7 from MiniMax, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 200K-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
MiniMax M3
MiniMax M3 from MiniMax, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1.05M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
MiMo V2.5
MiMo V2.5 from Xiaomi, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
MiMo V2.5 Pro
MiMo V2.5 Pro from Xiaomi, served through the LLMRPM proxy with a unified OpenAI- and Anthropic-compatible API. 1M-token context window. Chat completions, streaming, and tool calling. Supports adjustable reasoning effort for harder multi-step tasks.
Claude Opus 4.5
Claude Opus 4.5 from Anthropic, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Claude Sonnet 4.5
Claude Sonnet 4.5 from Anthropic, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
DeepSeek R1
DeepSeek R1 from DeepSeek, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
DeepSeek V3.2
DeepSeek V3.2 from DeepSeek, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Gemini 2.5 Pro
Gemini 2.5 Pro from Google, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Gemini 3 Flash
Gemini 3 Flash from Google, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Gemini 3 Pro
Gemini 3 Pro from Google, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Gemini 3.1 Pro
Gemini 3.1 Pro from Google, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
GPT-4.1
GPT-4.1 from OpenAI, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
GPT-4.5
GPT-4.5 from OpenAI, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
GPT-4o
GPT-4o from OpenAI, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
GPT-5
GPT-5 from OpenAI, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
GPT-5 Mini
GPT-5 Mini from OpenAI, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
GPT-5.1
GPT-5.1 from OpenAI, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
GPT-5.2
GPT-5.2 from OpenAI, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
GPT-5.3
GPT-5.3 from OpenAI, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
GPT-5.5 Pro
GPT-5.5 Pro from OpenAI, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
o3
o3 from OpenAI, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
o4
o4 from OpenAI, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Grok 3
Grok 3 from xAI, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Grok 3 Mini
Grok 3 Mini from xAI, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Grok 4
Grok 4 from xAI, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Grok 4.1
Grok 4.1 from xAI, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Llama 3.3 70B
Llama 3.3 70B from Meta, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Llama 4 Maverick
Llama 4 Maverick from Meta, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Llama 4 Scout
Llama 4 Scout from Meta, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Codestral
Codestral from Mistral, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Mistral Large 3
Mistral Large 3 from Mistral, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Mistral Medium 3.5
Mistral Medium 3.5 from Mistral, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Qwen 3.6 Plus
Qwen 3.6 Plus from Qwen, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Qwen3 Coder
Qwen3 Coder from Qwen, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
Qwen3 Max
Qwen3 Max from Qwen, verified through LLMRPM for chat completions, Messages, streaming, and automatic tool calling.
GLM
GLM served through the LLMRPM unified API.
Kimi
Kimi served through the LLMRPM unified API.
MiMo
MiMo served through the LLMRPM unified API.
Qwen
Qwen served through the LLMRPM unified API.
MiniMax
MiniMax served through the LLMRPM unified API.
DeepSeek
DeepSeek served through the LLMRPM unified API.
LongCat 2.0
LongCat 2.0 served through the LLMRPM unified API. 1M-token context window with tool calling and reasoning.
Unified format
Send requests in OpenAI or Anthropic shape and let LLMRPM handle the rest.
Streaming ready
Server-sent events are forwarded transparently for low-latency first-token responses.
Constant additions
New frontier models are added as providers release them — no code changes on your side.