OpenAI compatible API · Attested · Public status

DeepInfra

DeepInfra models on TrustedRouter with prices, routes, policy notes, and source links.

Verify gateway
Onebase URL to migrate
100sof models and routes
Noneprompt logs by default

deepinfra

No provider claim

All providers

ProviderDeepInfra
Models46 public models
Prepaid routes46
BYOK routes46
Zero data retentionno
Confidential computenot claimed
Provider E2EEnot claimed
Policy noteTracked as no-store, not strict ZDR. DeepInfra documents memory-only handling and no training for ordinary inference, but reserves the right to log a small portion of requests for debugging or security. Google- and Anthropic-backed routes also inherit those vendors' terms.
Policy source

Measured performance

191 samples

Continuously sampled across DeepInfra's routed models — p50 TTFT, throughput, and success rate. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.

p50 TTFT4415 ms
Throughput0 tok/s
Uptime96.34%
Modelp50 TTFTp50 TTFBThroughputUptimeConfig excludedSamples
mistralai/mistral-small-24b-instruct-2501 678 ms 678 ms 100.00% 2
qwen/qwen3-next-80b-a3b-instruct 960 ms 960 ms 100.00% 1
openai/gpt-oss-20b 1219 ms 1219 ms 100.00% 2
meta-llama/llama-3.1-70b-instruct 1728 ms 1728 ms 100.00% 5
z-ai/glm-5 1762 ms 1762 ms 100.00% 3
qwen/qwen3.5-122b-a10b 1988 ms 1988 ms 50.00% 2
qwen/qwen3.5-27b 2076 ms 2076 ms 100.00% 3
qwen/qwen3-30b-a3b 2142 ms 2142 ms 100.00% 7
qwen/qwen3-vl-30b-a3b-instruct 2343 ms 2342 ms 100.00% 6
qwen/qwen3-235b-a22b-thinking-2507 2575 ms 2575 ms 83.33% 6
deepseek/deepseek-v3.1-terminus 2842 ms 2842 ms 80.00% 5
qwen/qwen3.6-27b 2868 ms 2868 ms 100.00% 2
deepseek/deepseek-v4-flash 3114 ms 3113 ms 100.00% 4
z-ai/glm-5.2 3161 ms 3161 ms 100.00% 8
moonshotai/kimi-k2.5 3210 ms 3210 ms 100.00% 4
gryphe/mythomax-l2-13b 3545 ms 3545 ms 100.00% 9
google/gemini-2.5-pro 3846 ms 3846 ms 100.00% 3
google/gemini-2.5-flash 3863 ms 3863 ms 100.00% 9
google/gemma-4-26b-a4b-it 3997 ms 3997 ms 100.00% 6
deepseek/deepseek-r1-0528 4269 ms 4269 ms 100.00% 6
google/gemma-3-4b-it 4415 ms 4414 ms 100.00% 8
qwen/qwen3.6-35b-a3b 4527 ms 4527 ms 50.00% 4
openai/gpt-oss-120b 4892 ms 4892 ms 100.00% 5
google/gemma-4-31b-it 4951 ms 4951 ms 0 tok/s 100.00% 4
deepseek/deepseek-v4-pro 5843 ms 5843 ms 100.00% 1
google/gemma-3-27b-it 5934 ms 5933 ms 100.00% 5
google/gemini-3.1-flash-lite 6005 ms 6005 ms 100.00% 8
qwen/qwen3.5-397b-a17b 6017 ms 6017 ms 100.00% 4
moonshotai/kimi-k2.6 6219 ms 6219 ms 100.00% 4
qwen/qwen3.5-35b-a3b 6441 ms 6440 ms 100.00% 5
qwen/qwen3-vl-235b-a22b-instruct 6842 ms 6842 ms 100.00% 7
microsoft/phi-4 7726 ms 7726 ms 100.00% 2
z-ai/glm-4.6 8301 ms 8301 ms 100.00% 2
deepseek/deepseek-v3.2 8497 ms 8496 ms 66.67% 3
nvidia/nemotron-3-nano-30b-a3b 9239 ms 9239 ms 100.00% 2
qwen/qwen3-14b 9570 ms 9570 ms 100.00% 3
z-ai/glm-4.7 10103 ms 10103 ms 100.00% 3
qwen/qwen3.5-9b 10272 ms 10271 ms 100.00% 5
minimax/minimax-m3 10462 ms 10461 ms 100.00% 7
tencent/hy3 11317 ms 11316 ms 100.00% 5
google/gemma-3-12b-it 11516 ms 11516 ms 100.00% 3
minimax/minimax-m2.7 14380 ms 14380 ms 100.00% 6
meta-llama/llama-guard-4-12b 15583 ms 15583 ms 100.00% 1
Qwen/Qwen3-Embedding-8B 0.00% 1

DeepInfra performance history · Full provider & model leaderboard.

Provider models

Models served by DeepInfra.

Each row links to pricing, provider, benchmark, and API pages for the model.

Model AI IQ Context Endpoints Prompt Completion Routes
Qwen/Qwen3-Embedding-8B
Qwen3 Embedding 8B
32,000 2 $0.0105/1M selected route prepaid BYOK
deepseek/deepseek-r1-0528
DeepSeek: R1 0528
163,840 2 $0.525/1M $2.2575/1M prepaid BYOK
deepseek/deepseek-v3.1-terminus
DeepSeek: DeepSeek V3.1 Terminus
163,840 2 $0.2835/1M $0.9975/1M prepaid BYOK
deepseek/deepseek-v3.2
DeepSeek: DeepSeek V3.2
IQ 104#54 163,840 2 $0.273/1M $0.399/1M prepaid BYOK
deepseek/deepseek-v4-flash
DeepSeek: DeepSeek V4 Flash
IQ 109#42 1,048,576 2 $0.0945/1M $0.189/1M prepaid BYOK
deepseek/deepseek-v4-pro
DeepSeek: DeepSeek V4 Pro
IQ 114#27 1,048,576 2 $1.365/1M $2.73/1M prepaid BYOK
google/gemini-2.5-flash
Google: Gemini 2.5 Flash
1,048,576 2 $0.315/1M $2.625/1M prepaid BYOK
google/gemini-2.5-pro
Google: Gemini 2.5 Pro
IQ 104#55 1,048,576 2 $1.3125/1M $10.5/1M prepaid BYOK
google/gemini-3.1-flash-lite
Google: Gemini 3.1 Flash Lite
IQ 101#64 1,048,576 2 $0.2625/1M $1.575/1M prepaid BYOK
google/gemma-3-12b-it
Google: Gemma 3 12B
131,072 2 $0.0525/1M $0.1575/1M prepaid BYOK
google/gemma-3-27b-it
Google: Gemma 3 27B
262,144 2 $0.084/1M $0.168/1M prepaid BYOK
google/gemma-3-4b-it
Google: Gemma 3 4B
131,072 2 $0.0525/1M $0.105/1M prepaid BYOK
google/gemma-4-26b-a4b-it
Google: Gemma 4 26B A4B
IQ 97#77 262,144 2 $0.0735/1M $0.357/1M prepaid BYOK
google/gemma-4-31b-it
Google: Gemma 4 31B
IQ 101#65 262,144 2 $0.1365/1M $0.399/1M prepaid BYOK
gryphe/mythomax-l2-13b
MythoMax 13B
8,192 2 $0.42/1M $0.42/1M prepaid BYOK
meta-llama/llama-3.1-70b-instruct
Meta: Llama 3.1 70B Instruct
131,072 2 $0.42/1M $0.42/1M prepaid BYOK
meta-llama/llama-guard-4-12b
Meta: Llama Guard 4 12B
1,048,576 2 $0.189/1M $0.189/1M prepaid BYOK
microsoft/phi-4
Microsoft: Phi 4
16,384 2 $0.0735/1M $0.147/1M prepaid BYOK
minimax/minimax-m2.7
MiniMax: MiniMax M2.7
IQ 107#49 204,800 2 $0.2625/1M $1.05/1M prepaid BYOK
minimax/minimax-m3
MiniMax: MiniMax M3
IQ 112#37 1,048,576 2 $0.315/1M $1.26/1M prepaid BYOK
mistralai/mistral-small-24b-instruct-2501
Mistral: Mistral Small 3
32,768 2 $0.0525/1M $0.084/1M prepaid BYOK
moonshotai/kimi-k2.5
MoonshotAI: Kimi K2.5
IQ 111#39 262,144 2 $0.4725/1M $2.3625/1M prepaid BYOK
moonshotai/kimi-k2.6
MoonshotAI: Kimi K2.6
IQ 119#19 262,144 2 $0.7875/1M $3.675/1M prepaid BYOK
nousresearch/hermes-3-llama-3.1-405b
Nous: Hermes 3 405B Instruct
131,072 2 $1.05/1M $1.05/1M prepaid BYOK
nvidia/nemotron-3-nano-30b-a3b
NVIDIA: Nemotron 3 Nano 30B A3B
262,144 2 $0.0525/1M $0.21/1M prepaid BYOK
openai/gpt-oss-120b
OpenAI: gpt-oss-120b
IQ 103#57 131,072 2 $0.03885/1M $0.1785/1M prepaid BYOK
openai/gpt-oss-20b
OpenAI: gpt-oss-20b
IQ 100#68 131,072 2 $0.0315/1M $0.147/1M prepaid BYOK
qwen/qwen3-14b
Qwen: Qwen3 14B
131,072 2 $0.126/1M $0.252/1M prepaid BYOK
qwen/qwen3-235b-a22b-thinking-2507
Qwen: Qwen3 235B A22B Thinking 2507
262,144 2 $0.2415/1M $2.415/1M prepaid BYOK
qwen/qwen3-30b-a3b
Qwen: Qwen3 30B A3B
131,072 2 $0.126/1M $0.525/1M prepaid BYOK
qwen/qwen3-next-80b-a3b-instruct
Qwen: Qwen3 Next 80B A3B Instruct
262,144 2 $0.0945/1M $1.155/1M prepaid BYOK
qwen/qwen3-vl-235b-a22b-instruct
Qwen: Qwen3 VL 235B A22B Instruct
262,144 2 $0.21/1M $0.924/1M prepaid BYOK
qwen/qwen3-vl-30b-a3b-instruct
Qwen: Qwen3 VL 30B A3B Instruct
262,144 2 $0.1575/1M $0.63/1M prepaid BYOK
qwen/qwen3.5-122b-a10b
Qwen: Qwen3.5-122B-A10B
262,144 2 $0.3045/1M $2.52/1M prepaid BYOK
qwen/qwen3.5-27b
Qwen: Qwen3.5-27B
262,144 2 $0.273/1M $2.73/1M prepaid BYOK
qwen/qwen3.5-35b-a3b
Qwen: Qwen3.5-35B-A3B
262,144 2 $0.147/1M $1.05/1M prepaid BYOK
qwen/qwen3.5-397b-a17b
Qwen: Qwen3.5 397B A17B
262,144 2 $0.4725/1M $3.15/1M prepaid BYOK
qwen/qwen3.5-9b
Qwen: Qwen3.5-9B
IQ 93#88 262,144 2 $0.105/1M $0.1575/1M prepaid BYOK
qwen/qwen3.6-27b
Qwen: Qwen3.6 27B
IQ 112#38 262,144 2 $0.336/1M $3.36/1M prepaid BYOK
qwen/qwen3.6-35b-a3b
Qwen: Qwen3.6 35B A3B
IQ 100#70 262,144 2 $0.1575/1M $0.9975/1M prepaid BYOK
tencent/hy3
Tencent: Hy3
IQ 103#58 262,144 2 $0.147/1M $0.609/1M prepaid BYOK
z-ai/glm-4.6
Z.ai: GLM 4.6
204,800 2 $0.525/1M $2.1/1M prepaid BYOK
z-ai/glm-4.7
Z.ai: GLM 4.7
IQ 103#56 204,800 2 $0.42/1M $1.8375/1M prepaid BYOK
z-ai/glm-4.7-flash
Z.ai: GLM 4.7 Flash
202,752 2 $0.063/1M $0.42/1M prepaid BYOK
z-ai/glm-5
Z.ai: GLM 5
IQ 106#50 204,800 2 $0.63/1M $2.184/1M prepaid BYOK
z-ai/glm-5.2
Z.ai: GLM 5.2
IQ 120#16 1,048,576 2 $1.26/1M $4.41/1M prepaid BYOK
Workspace access

Sign in

Choose a sign in method. New email and OAuth accounts include $0.10 in starter credit; wallet-only accounts start at $0.

By signing in you agree to the terms of service and privacy policy.