OpenAI compatible API · Attested · Public status

DeepInfra

Explore DeepInfra models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

DeepInfradeepinfra

No provider claim

All providers

ProviderDeepInfra
Provider websitehttps://deepinfra.com/
Models95 public models
Prepaid routes90
BYOK routes94
Zero data retentionno
Confidential computenot claimed
Provider E2EEnot claimed
Policy noteTracked as no-store, not strict ZDR. DeepInfra documents memory-only handling and no training for ordinary inference, but reserves the right to log a small portion of requests for debugging or security. Google- and Anthropic-backed routes also inherit those vendors' terms.
Policy source

Measured performance

92 samples

Continuously sampled across DeepInfra's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.

p50 TTFT3743 ms
Effective throughput27 tok/s n=21
Uptime98.91%
Modelp50 TTFTp50 TTFBEffective throughputUptimeConfig excludedAvailability samples
deepseek/deepseek-v4-flash-0731 6216 ms 1904 ms 25 tok/s n=1 100.00% 14
google/gemma-3-27b-it 677 ms 677 ms 100.00% 2
qwen/qwen3-vl-30b-a3b-instruct 761 ms 761 ms 100.00% 2
nvidia/nemotron-3-nano-30b-a3b 916 ms 915 ms 100.00% 3
nvidia/nemotron-3.5-lightning 941 ms 941 ms 82 tok/s n=1 100.00% 2
mistralai/mistral-small-24b-instruct-2501 976 ms 976 ms 100.00% 2
sao10k/l3-8b-lunaris-v1-turbo 1190 ms 1190 ms 100.00% 1
bytedance/seed-2.0-mini 1414 ms 1414 ms 100.00% 2
moonshotai/kimi-k3 1464 ms 1464 ms 16 tok/s n=1 100.00% 1
qwen/qwen3-next-80b-a3b-instruct 1500 ms 1500 ms 100.00% 2
google/gemma-4-26b-a4b-it 1669 ms 1668 ms 100.00% 1
deepseek/deepseek-r1-0528 2459 ms 2459 ms 100.00% 4
google/gemma-4-31b-it 2482 ms 2481 ms 16 tok/s n=1 100.00% 1
google/gemma-4-31b-it-turbo 3013 ms 3013 ms 100.00% 2
minimax/minimax-m3 3067 ms 3067 ms 17 tok/s n=2 100.00% 4
z-ai/glm-4.6 3417 ms 3416 ms 100.00% 3
meta-llama/llama-4-scout-17b-16e-instruct 3743 ms 3743 ms 100.00% 1
deepseek/deepseek-v3.2 3842 ms 3842 ms 100.00% 1
qwen/qwen3.5-35b-a3b 3844 ms 3844 ms 100.00% 1
google/gemma-4-e4b-it 3923 ms 3923 ms 100.00% 1
z-ai/glm-4.7 4064 ms 4064 ms 100.00% 1
meta-llama/llama-guard-4-12b 4080 ms 4079 ms 100.00% 1
meta-llama/llama-3.1-70b-instruct 4103 ms 100.00% 3
microsoft/phi-4 4138 ms 4138 ms 100.00% 2
minimax/minimax-m2.7 4209 ms 4209 ms 100.00% 2
qwen/qwen3.6-35b-a3b 4232 ms 4231 ms 100.00% 2
qwen/qwen3.5-397b-a17b 4261 ms 4261 ms 100.00% 1
qwen/qwen3-vl-235b-a22b-instruct 4319 ms 4319 ms 100.00% 1
thinkingmachines/inkling-small 4363 ms 4363 ms 27 tok/s n=1 100.00% 1
qwen/qwen2.5-72b-instruct 4582 ms 4582 ms 100.00% 1
bytedance/seed-2.0-code 4631 ms 4630 ms 100.00% 1
google/gemma-4-31b-it-ultra 5622 ms 5621 ms 100.00% 1
qwen/qwen3.7-max 6244 ms 6244 ms 77 tok/s n=1 100.00% 1
google/gemini-3.1-pro 6262 ms 6262 ms 100.00% 1
bytedance/seed-2.0-pro 100.00% 2
gryphe/mythomax-l2-13b 100.00% 1
meta-models/muse-glimmer-30b 100.00% 1
mistralai/mistral-nemo-instruct-2407 100.00% 2
nousresearch/hermes-3-llama-3.1-405b 100.00% 2
nvidia/nemotron-3-ultra-550b-a55b 15 tok/s n=1 100.00% 2
openai/gpt-oss-120b 40 tok/s n=1 100.00% 2
qwen/qwen3-max 100.00% 1
qwen/qwen3-max-thinking 100.00% 1
stepfun-ai/step-3.7-flash 100.00% 1
tencent/hy3 100.00% 2
thinkingmachines/inkling 10 tok/s n=1 100.00% 2
qwen/qwen3-235b-a22b-thinking-2507 3055 ms 3054 ms 75.00% 4
deepseek/deepseek-v4-flash 39 tok/s n=2 0
deepseek/deepseek-v4-pro-0423 19 tok/s n=2 0
google/gemini-3.7-flash 51 tok/s n=1 0
inclusionai/ling-3.0-flash 55 tok/s n=1 0
moonshotai/kimi-k2.6 22 tok/s n=1 0
qwen/qwen3.8-2.4t-a95b 39 tok/s n=1 0
z-ai/glm-5.2 52 tok/s n=2 0

DeepInfra performance history · Full provider & model leaderboard.

Provider models

Models served by DeepInfra.

Each row links to pricing, provider, benchmark, and API pages for the model.

Model AI IQ Context Endpoints Prompt Completion Routes
Qwen/Qwen3-Embedding-8B
Qwen3 Embedding 8B
32,000 2 $0.01055/1M selected route prepaid BYOK
anthropic/claude-haiku-4-5
anthropic/claude-haiku-4-5
200,000 1 $1.055/1M $5.275/1M BYOK
anthropic/claude-opus-4-7
anthropic/claude-opus-4-7
1,000,000 1 $5.275/1M $26.375/1M BYOK
anthropic/claude-opus-4-8
anthropic/claude-opus-4-8
1,000,000 1 $5.275/1M $26.375/1M BYOK
anthropic/claude-opus-5
Claude Opus 5
IQ 134#3 1,000,000 1 $5.275/1M $26.375/1M BYOK
anthropic/claude-sonnet-4-6
anthropic/claude-sonnet-4-6
1,000,000 1 $3.165/1M $15.825/1M BYOK
bytedance/seed-1.8
ByteDance/Seed-1.8
256,000 2 $0.26375/1M $2.11/1M prepaid BYOK
bytedance/seed-2.0-code
ByteDance/Seed-2.0-code
256,000 2 $0.5275/1M $3.165/1M prepaid BYOK
bytedance/seed-2.0-mini
ByteDance/Seed-2.0-mini
256,000 2 $0.1055/1M $0.422/1M prepaid BYOK
bytedance/seed-2.0-pro
ByteDance/Seed-2.0-pro
256,000 2 $0.5275/1M $3.165/1M prepaid BYOK
deepseek/deepseek-r1-0528
DeepSeek: R1 0528
163,840 2 $0.5275/1M $2.26825/1M prepaid BYOK
deepseek/deepseek-v3
deepseek-ai/DeepSeek-V3
163,840 2 $0.3376/1M $0.93895/1M prepaid BYOK
deepseek/deepseek-v3-0324
DeepSeek V3 0324
163,840 2 $0.2532/1M $0.9495/1M prepaid BYOK
deepseek/deepseek-v3.1
DeepSeek V3.1
IQ 96#88 131,072 2 $0.26375/1M $1.00225/1M prepaid BYOK
deepseek/deepseek-v3.2
DeepSeek: DeepSeek V3.2
IQ 103#64 163,840 2 $0.2743/1M $0.4009/1M prepaid BYOK
deepseek/deepseek-v4-flash
DeepSeek: DeepSeek V4 Flash 0423
IQ 113#41 1,048,576 2 $0.09495/1M $0.1899/1M prepaid BYOK
deepseek/deepseek-v4-flash-0731
DeepSeek: DeepSeek V4 Flash 0731
1,048,576 2 $0.0844/1M $0.1899/1M prepaid BYOK
deepseek/deepseek-v4-pro
DeepSeek: DeepSeek V4 Pro
IQ 117#31 1,048,576 2 $1.3715/1M $2.743/1M prepaid BYOK
deepseek/deepseek-v4-pro-0423
DeepSeek V4 Pro 0423
1,048,576 1 $1.3715/1M $2.743/1M prepaid
google/gemini-2.5-flash
Google: Gemini 2.5 Flash
1,048,576 2 $0.3165/1M $2.6375/1M prepaid BYOK
google/gemini-2.5-pro
Google: Gemini 2.5 Pro
IQ 103#65 1,048,576 2 $1.31875/1M $10.55/1M prepaid BYOK
google/gemini-3.1-flash-lite
Google: Gemini 3.1 Flash Lite
IQ 101#72 1,048,576 2 $0.26375/1M $1.5825/1M prepaid BYOK
google/gemini-3.1-pro
google/gemini-3.1-pro
IQ 127#10 1,000,000 2 $2.11/1M $12.66/1M prepaid BYOK
google/gemini-3.7-flash
Google: Gemini 3.7 Flash
IQ 122#17 1,048,576 2 $0.79125/1M $3.95625/1M prepaid BYOK
google/gemma-3-12b-it
Google: Gemma 3 12B
131,072 2 $0.05275/1M $0.15825/1M prepaid BYOK
google/gemma-3-27b-it
Google: Gemma 3 27B
262,144 2 $0.0844/1M $0.1688/1M prepaid BYOK
google/gemma-3-4b-it
Google: Gemma 3 4B
131,072 2 $0.05275/1M $0.1055/1M prepaid BYOK
google/gemma-4-26b-a4b-it
Google: Gemma 4 26B A4B
IQ 96#89 262,144 2 $0.07385/1M $0.3587/1M prepaid BYOK
google/gemma-4-31b-it
Google: Gemma 4 31B
IQ 101#73 262,144 2 $0.13715/1M $0.4009/1M prepaid BYOK
google/gemma-4-31b-it-turbo
google/gemma-4-31B-it-turbo
262,144 2 $0.09495/1M $0.3587/1M prepaid BYOK
google/gemma-4-31b-it-ultra
google/gemma-4-31B-it-Ultra
131,072 2 $0.28485/1M $0.8018/1M prepaid BYOK
google/gemma-4-e4b-it
google/gemma-4-E4B-it
IQ 93#98 131,072 2 $0.0211/1M $0.1055/1M prepaid BYOK
gryphe/mythomax-l2-13b
MythoMax 13B
8,192 2 $0.422/1M $0.422/1M prepaid BYOK
inclusionai/ling-3.0-flash
Ling-3.0-flash
IQ 92#100 262,144 2 $0.0633/1M $0.1899/1M prepaid BYOK
meta-llama/llama-3.1-70b-instruct
Meta: Llama 3.1 70B Instruct
131,072 2 $0.422/1M $0.422/1M prepaid BYOK
meta-llama/llama-3.3-70b-instruct-turbo
meta-llama/Llama-3.3-70B-Instruct-Turbo
131,072 2 $0.1055/1M $0.3376/1M prepaid BYOK
meta-llama/llama-4-maverick-17b-128e-instruct-fp8
Llama 4 Maverick Instruct
1,048,576 2 $0.211/1M $0.844/1M prepaid BYOK
meta-llama/llama-4-scout-17b-16e-instruct
Llama 4 Scout Instruct
131,072 2 $0.1055/1M $0.3165/1M prepaid BYOK
meta-llama/llama-guard-4-12b
Meta: Llama Guard 4 12B
1,048,576 2 $0.1899/1M $0.1899/1M prepaid BYOK
meta-llama/meta-llama-3.1-8b-instruct-turbo
meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo
131,072 2 $0.0211/1M $0.0422/1M prepaid BYOK
meta-models/muse-glimmer-30b
meta-models/Muse-Glimmer-30B
131,072 2 $0.3165/1M $1.266/1M prepaid BYOK
microsoft/phi-4
Microsoft: Phi 4
16,384 2 $0.07385/1M $0.1477/1M prepaid BYOK
minimax/minimax-m2.7
MiniMax: MiniMax M2.7
IQ 109#50 204,800 2 $0.26375/1M $1.055/1M prepaid BYOK
minimax/minimax-m2.7-turbo
MiniMaxAI/MiniMax-M2.7-Turbo
196,608 2 $0.4009/1M $1.7935/1M prepaid BYOK
minimax/minimax-m3
MiniMax: MiniMax M3
IQ 114#39 1,048,576 2 $0.2954/1M $1.1605/1M prepaid BYOK
mistralai/mistral-nemo-instruct-2407
mistralai/Mistral-Nemo-Instruct-2407
131,072 2 $0.020045/1M $0.03165/1M prepaid BYOK
mistralai/mistral-small-24b-instruct-2501
Mistral: Mistral Small 3
32,768 2 $0.05275/1M $0.0844/1M prepaid BYOK
mistralai/mistral-small-3.2-24b-instruct-2506
mistralai/Mistral-Small-3.2-24B-Instruct-2506
128,000 2 $0.079125/1M $0.211/1M prepaid BYOK
moonshotai/kimi-k2.5
MoonshotAI: Kimi K2.5
IQ 111#46 262,144 2 $0.47475/1M $2.37375/1M prepaid BYOK
moonshotai/kimi-k2.6
MoonshotAI: Kimi K2.6
IQ 119#23 262,144 2 $0.79125/1M $3.6925/1M prepaid BYOK
moonshotai/kimi-k3
MoonshotAI: Kimi K3
IQ 123#16 1,048,576 2 $3.00675/1M $15.03375/1M prepaid BYOK
nousresearch/hermes-3-llama-3.1-405b
Nous: Hermes 3 405B Instruct
131,072 2 $1.055/1M $1.055/1M prepaid BYOK
nvidia/nemotron-3-nano-30b-a3b
NVIDIA: Nemotron 3 Nano 30B A3B
262,144 2 $0.05275/1M $0.211/1M prepaid BYOK
nvidia/nemotron-3-super-120b-a12b
NVIDIA: Nemotron 3 Super
1,000,000 2 $0.089675/1M $0.422/1M prepaid BYOK
nvidia/nemotron-3-ultra-550b-a55b
NVIDIA: Nemotron 3 Ultra
512,288 2 $0.5275/1M $2.321/1M prepaid BYOK
nvidia/nemotron-3.5-lightning
NVIDIA: Nemotron 3.5 Lightning
1,000,000 2 $0.0844/1M $0.211/1M prepaid BYOK
nvidia/nemotron-content-safety-3.5
nvidia/Nemotron-Content-Safety-3.5
131,072 2 $0.211/1M $0.211/1M prepaid BYOK
openai/gpt-oss-120b
OpenAI: gpt-oss-120b
IQ 105#61 131,072 2 $0.039035/1M $0.17935/1M prepaid BYOK
openai/gpt-oss-120b-turbo
openai/gpt-oss-120b-Turbo
131,072 2 $0.15825/1M $0.633/1M prepaid BYOK
openai/gpt-oss-120b-ultra
openai/gpt-oss-120b-Ultra
131,072 2 $0.211/1M $1.00225/1M prepaid BYOK
openai/gpt-oss-20b
OpenAI: gpt-oss-20b
IQ 100#77 131,072 2 $0.03165/1M $0.1477/1M prepaid BYOK
qwen/qwen2.5-72b-instruct
Qwen/Qwen2.5-72B-Instruct
32,768 2 $0.3798/1M $0.422/1M prepaid BYOK
qwen/qwen3-14b
Qwen: Qwen3 14B
131,072 2 $0.1266/1M $0.2532/1M prepaid BYOK
qwen/qwen3-235b-a22b-instruct-2507
Qwen3 235B A22B Instruct 2507
131,072 2 $0.09495/1M $0.58025/1M prepaid BYOK
qwen/qwen3-235b-a22b-thinking-2507
Qwen: Qwen3 235B A22B Thinking 2507
262,144 2 $0.24265/1M $2.4265/1M prepaid BYOK
qwen/qwen3-30b-a3b
Qwen: Qwen3 30B A3B
131,072 2 $0.1266/1M $0.5275/1M prepaid BYOK
qwen/qwen3-coder-480b-a35b-instruct-turbo
Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo
262,144 2 $0.3165/1M $1.055/1M prepaid BYOK
qwen/qwen3-max
Qwen3 Max
262,144 2 $1.266/1M $6.33/1M prepaid BYOK
qwen/qwen3-max-thinking
Qwen/Qwen3-Max-Thinking
256,000 2 $1.266/1M $6.33/1M prepaid BYOK
qwen/qwen3-next-80b-a3b-instruct
Qwen: Qwen3 Next 80B A3B Instruct
262,144 2 $0.09495/1M $1.1605/1M prepaid BYOK
qwen/qwen3-vl-235b-a22b-instruct
Qwen: Qwen3 VL 235B A22B Instruct
262,144 2 $0.211/1M $0.9284/1M prepaid BYOK
qwen/qwen3-vl-30b-a3b-instruct
Qwen: Qwen3 VL 30B A3B Instruct
262,144 2 $0.15825/1M $0.633/1M prepaid BYOK
qwen/qwen3.5-122b-a10b
Qwen: Qwen3.5-122B-A10B
262,144 2 $0.30595/1M $2.532/1M prepaid BYOK
qwen/qwen3.5-27b
Qwen: Qwen3.5-27B
262,144 2 $0.2743/1M $2.743/1M prepaid BYOK
qwen/qwen3.5-35b-a3b
Qwen: Qwen3.5-35B-A3B
262,144 2 $0.1477/1M $1.055/1M prepaid BYOK
qwen/qwen3.5-397b-a17b
Qwen: Qwen3.5 397B A17B
262,144 2 $0.47475/1M $3.165/1M prepaid BYOK
qwen/qwen3.5-9b
Qwen: Qwen3.5-9B
IQ 93#99 262,144 2 $0.1055/1M $0.15825/1M prepaid BYOK
qwen/qwen3.6-27b
Qwen: Qwen3.6 27B
IQ 111#47 262,144 2 $0.3376/1M $3.376/1M prepaid BYOK
qwen/qwen3.6-35b-a3b
Qwen: Qwen3.6 35B A3B
IQ 100#78 262,144 2 $0.1055/1M $1.00225/1M prepaid BYOK
qwen/qwen3.7-max
Qwen3.7-Max
IQ 118#29 1,000,000 2 $2.6375/1M $7.9125/1M prepaid BYOK
qwen/qwen3.8-2.4t-a95b
Qwen: Qwen3.8 2.4T A95B
IQ 119#25 1,048,576 2 $2.11/1M $6.33/1M prepaid BYOK
qwen/qwen3.8-max
Qwen3.8 Max
IQ 119#26 1,000,000 2 $1.74075/1M $5.223305/1M prepaid BYOK
sao10k/l3-8b-lunaris-v1-turbo
Sao10K/L3-8B-Lunaris-v1-Turbo
8,192 2 $0.0422/1M $0.05275/1M prepaid BYOK
sao10k/l3.1-70b-euryale-v2.2
Sao10K/L3.1-70B-Euryale-v2.2
131,072 2 $0.89675/1M $0.89675/1M prepaid BYOK
stepfun-ai/step-3.7-flash
stepfun-ai/Step-3.7-Flash
IQ 101#76 262,144 2 $0.211/1M $1.21325/1M prepaid BYOK
tencent/hy3
Tencent: Hy3
IQ 103#67 262,144 2 $0.1477/1M $0.6119/1M prepaid BYOK
thinkingmachines/inkling
Thinking Machines: Inkling
IQ 108#53 1,048,576 2 $1.00225/1M $4.27275/1M prepaid BYOK
thinkingmachines/inkling-small
Thinking Machines: Inkling Small
IQ 106#59 524,288 2 $0.47475/1M $1.266/1M prepaid BYOK
xiaomimimo/mimo-v2.5
XiaomiMiMo/MiMo-V2.5
IQ 109#49 262,144 2 $0.422/1M $2.11/1M prepaid BYOK
xiaomimimo/mimo-v2.5-pro
XiaomiMiMo/MiMo-V2.5-Pro
IQ 115#36 1,048,576 2 $1.055/1M $3.165/1M prepaid BYOK
z-ai/glm-4.6
Z.ai: GLM 4.6
204,800 2 $0.5275/1M $2.11/1M prepaid BYOK
z-ai/glm-4.7
Z.ai: GLM 4.7
IQ 103#66 204,800 2 $0.422/1M $1.84625/1M prepaid BYOK
z-ai/glm-4.7-flash
Z.ai: GLM 4.7 Flash
202,752 2 $0.0633/1M $0.422/1M prepaid BYOK
z-ai/glm-5
Z.ai: GLM 5
IQ 105#60 204,800 2 $0.633/1M $2.1944/1M prepaid BYOK
z-ai/glm-5.2
Z.ai: GLM 5.2
IQ 120#19 1,048,576 2 $0.79125/1M $2.532/1M prepaid BYOK

Questions

Does DeepInfra have zero data retention?

TrustedRouter does not currently mark DeepInfra as provider-level zero data retention. Use trustedrouter/zdr or provider.min_privacy=zdr to select a different eligible route, and review the linked policy source for changes.

Is DeepInfra end-to-end encrypted?

TrustedRouter does not currently mark DeepInfra as end-to-end encrypted at the provider boundary. The TrustedRouter gateway is still attested, but the selected provider normally receives the request in order to run the model. Use trustedrouter/e2e for the stronger route requirement.

Which DeepInfra models are available through TrustedRouter?

This page currently lists 95 public DeepInfra models, with live pricing, route count, context length, measured performance when available, and links to each model's provider and benchmark pages.

Workspace access

Sign in

Choose a sign in method to access your TrustedRouter workspace.

By signing in you agree to the terms of service and privacy policy.