OpenAI compatible API · Attested · Public status

DeepInfra

Explore DeepInfra models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

DeepInfradeepinfra

No provider claim

All providers

ProviderDeepInfra
Routing statusActive
Provider websitehttps://deepinfra.com/
Models93 public models
Credits routes93
Zero data retentionno
Confidential computenot claimed
Provider E2EEnot claimed
Policy noteTracked as no-store, not strict ZDR. DeepInfra documents memory-only handling and no training for ordinary inference, but reserves the right to log a small portion of requests for debugging or security. Google- and Anthropic-backed routes also inherit those vendors' terms.
Policy source

Measured performance

55 samples

Continuously sampled across DeepInfra's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.

p50 TTFT2142 ms
Effective throughput29 tok/s n=15
Uptime98.18%
Modelp50 TTFTEffective throughputUptimeConfig excludedAvailability samples
qwen/qwen3.5-27b 387 ms 100.00% 2
qwen/qwen3-vl-30b-a3b-instruct 455 ms 100.00% 1
thinkingmachines/inkling 707 ms 100.00% 2
stepfun-ai/step-3.7-flash 935 ms 100.00% 1
inclusionai/ling-3.0-flash-fin 973 ms 100.00% 2
gryphe/mythomax-l2-13b 1067 ms 100.00% 2
qwen/qwen3-vl-235b-a22b-instruct 1302 ms 100.00% 1
bytedance/seed-2.0-code 1426 ms 100.00% 1
openai/gpt-oss-120b-ultra 1695 ms 100.00% 1
deepseek/deepseek-v3.1 1747 ms 100.00% 1
meta-llama/llama-4-maverick-17b-128e-instruct-fp8 1843 ms 100.00% 1
google/gemma-4-31b-it-turbo 1955 ms 100.00% 1
qwen/qwen3.5-35b-a3b 2142 ms 100.00% 1
google/gemini-2.5-flash 2170 ms 100.00% 1
microsoft/phi-4 2229 ms 100.00% 1
deepseek/deepseek-v3.2 2398 ms 100.00% 1
openai/gpt-oss-120b-turbo 2645 ms 100.00% 1
qwen/qwen3-max 2875 ms 100.00% 1
deepseek/deepseek-v4-flash-0731 3534 ms 43 tok/s n=1 100.00% 4
deepseek/deepseek-v3 3563 ms 100.00% 1
nvidia/nemotron-3.5-lightning 4665 ms 100.00% 1
meta-llama/llama-4-scout-17b-16e-instruct 5483 ms 100.00% 1
deepseek/deepseek-v4-flash-vision-exp 6116 ms 100.00% 3
qwen/qwen3-235b-a22b-instruct-2507 8534 ms 100.00% 1
minimax/minimax-m3 19 tok/s n=2 100.00% 18
qwen/qwen3.8-27b 100.00% 3
deepseek/deepseek-v4-flash 22 tok/s n=1 0
deepseek/deepseek-v4-pro 36 tok/s n=1 0
deepseek/deepseek-v4-pro-0423 95 tok/s n=2 0
google/gemma-4-31b-it 9 tok/s n=1 0
moonshotai/kimi-k2.6 15 tok/s n=1 0
moonshotai/kimi-k3 8 tok/s n=2 0
xiaomi/mimo-v2.5-pro 29 tok/s n=2 0
z-ai/glm-4.6 0.00% 1
z-ai/glm-5.2 53 tok/s n=2 0

DeepInfra performance history · Full provider & model leaderboard.

Provider models

Models served by DeepInfra.

Each row links to pricing, provider, benchmark, and API pages for the model.

Model AI IQ Context Input Cached input Output
Qwen/Qwen3-Embedding-8B
Qwen3 Embedding 8B
32,000 $0.01055/1M Not published selected route
bytedance/seed-1.8
ByteDance/Seed-1.8
256,000 $0.26375/1M $0.05275/1M $2.11/1M
bytedance/seed-2.0-code
ByteDance/Seed-2.0-code
256,000 $0.5275/1M $0.1055/1M $3.165/1M
bytedance/seed-2.0-mini
ByteDance/Seed-2.0-mini
256,000 $0.1055/1M $0.0211/1M $0.422/1M
bytedance/seed-2.0-pro
ByteDance/Seed-2.0-pro
256,000 $0.5275/1M $0.1055/1M $3.165/1M
deepseek/deepseek-r1-0528
DeepSeek: R1 0528
163,840 $0.5275/1M $0.36925/1M $2.26825/1M
deepseek/deepseek-v3
deepseek-ai/DeepSeek-V3
163,840 $0.3376/1M Not published $0.93895/1M
deepseek/deepseek-v3-0324
DeepSeek V3 0324
163,840 $0.2532/1M $0.142425/1M $0.9495/1M
deepseek/deepseek-v3.1
DeepSeek V3.1
IQ 94#102 131,072 $0.26375/1M $0.13715/1M $1.00225/1M
deepseek/deepseek-v3.2
DeepSeek: DeepSeek V3.2
IQ 103#73 163,840 $0.2743/1M $0.13715/1M $0.4009/1M
deepseek/deepseek-v4-flash
DeepSeek: DeepSeek V4 Flash 0423
IQ 113#46 1,024,000 $0.09495/1M $0.01899/1M $0.1899/1M
deepseek/deepseek-v4-flash-0731
DeepSeek: DeepSeek V4 Flash 0731
1,048,576 $0.0633/1M $0.015825/1M $0.1899/1M
deepseek/deepseek-v4-flash-vision-exp
DeepSeek: DeepSeek V4 Flash Vision Exp
IQ 113#47 1,048,576 $0.4642/1M $0.1477/1M $1.3926/1M
deepseek/deepseek-v4-pro
DeepSeek: DeepSeek V4 Pro 0423
IQ 117#36 1,024,000 $1.3715/1M $0.1055/1M $2.743/1M
deepseek/deepseek-v4-pro-0423
DeepSeek V4 Pro 0423
1,024,000 $1.3715/1M $0.1055/1M $2.743/1M
google/gemini-2.5-flash
Google: Gemini 2.5 Flash
1,048,576 $0.3165/1M Not published $2.6375/1M
google/gemini-2.5-pro
Google: Gemini 2.5 Pro
IQ 100#85 1,048,576 $1.31875/1M Not published $10.55/1M
google/gemini-3.1-flash-lite
Google: Gemini 3.1 Flash Lite
IQ 101#79 1,048,576 $0.26375/1M Not published $1.5825/1M
google/gemini-3.1-pro
google/gemini-3.1-pro
IQ 127#11 1,000,000 $2.11/1M Not published $12.66/1M
google/gemini-3.7-flash
Google: Gemini 3.7 Flash
IQ 123#18 1,048,576 $0.79125/1M Not published $3.95625/1M
google/gemma-3-12b-it
Google: Gemma 3 12B
131,072 $0.05275/1M Not published $0.15825/1M
google/gemma-3-27b-it
Google: Gemma 3 27B
131,072 $0.0844/1M Not published $0.1688/1M
google/gemma-3-4b-it
Google: Gemma 3 4B
131,072 $0.05275/1M Not published $0.1055/1M
google/gemma-4-26b-a4b-it
Google: Gemma 4 26B A4B
IQ 96#98 262,144 $0.07385/1M Not published $0.3587/1M
google/gemma-4-31b-it
Google: Gemma 4 31B
IQ 101#80 262,144 $0.13715/1M Not published $0.4009/1M
google/gemma-4-31b-it-turbo
google/gemma-4-31B-it-turbo
262,144 $0.09495/1M $0.05275/1M $0.3587/1M
google/gemma-4-31b-it-ultra
google/gemma-4-31B-it-Ultra
131,072 $0.28485/1M Not published $0.8018/1M
google/gemma-4-e4b-it
google/gemma-4-E4B-it
IQ 92#107 131,072 $0.0211/1M Not published $0.1055/1M
gryphe/mythomax-l2-13b
MythoMax 13B
4,096 $0.422/1M Not published $0.422/1M
ibm-granite/granite-4.2-30b
ibm-granite/granite-4.2-30b
131,072 $0.1688/1M $0.0422/1M $0.68575/1M
ibm-granite/granite-4.2-3b
ibm-granite/granite-4.2-3b
131,072 $0.03165/1M $0.01/1M $0.1266/1M
ibm-granite/granite-4.2-8b
IBM: Granite 4.2 8B
131,072 $0.0633/1M $0.015825/1M $0.26375/1M
inclusionai/ling-3.0-flash
inclusionAI: Ling 3.0 Flash
IQ 92#109 262,144 $0.0633/1M $0.01266/1M $0.1899/1M
inclusionai/ling-3.0-flash-fin
inclusionAI: Ling 3.0 Flash Fin
262,144 $0.0633/1M $0.01266/1M $0.1899/1M
meta-llama/llama-3.1-70b-instruct
Meta: Llama 3.1 70B Instruct
131,072 $0.422/1M Not published $0.422/1M
meta-llama/llama-3.3-70b-instruct-turbo
meta-llama/Llama-3.3-70B-Instruct-Turbo
131,072 $0.1055/1M Not published $0.3376/1M
meta-llama/llama-4-maverick-17b-128e-instruct-fp8
Llama 4 Maverick Instruct
1,048,576 $0.211/1M Not published $0.844/1M
meta-llama/llama-4-scout-17b-16e-instruct
Llama 4 Scout Instruct
131,072 $0.1055/1M Not published $0.3165/1M
meta-llama/llama-guard-4-12b
Meta: Llama Guard 4 12B
163,840 $0.1899/1M Not published $0.1899/1M
meta-llama/meta-llama-3.1-8b-instruct-turbo
meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo
131,072 $0.0211/1M Not published $0.0422/1M
meta-models/muse-glimmer-30b
Muse Glimmer 30B on Fireworks
131,072 $0.3165/1M $0.0422/1M $1.266/1M
microsoft/phi-4
Microsoft: Phi 4
16,384 $0.07385/1M Not published $0.1477/1M
minimax/minimax-m2.7-turbo
MiniMaxAI/MiniMax-M2.7-Turbo
196,608 $0.4009/1M $0.07385/1M $1.7935/1M
minimax/minimax-m3
MiniMax: MiniMax M3
IQ 114#44 524,288 $0.2954/1M $0.05908/1M $1.1605/1M
mistralai/mistral-nemo-instruct-2407
mistralai/Mistral-Nemo-Instruct-2407
131,072 $0.020045/1M Not published $0.03165/1M
mistralai/mistral-small-24b-instruct-2501
Mistral: Mistral Small 3
32,768 $0.05275/1M Not published $0.0844/1M
mistralai/mistral-small-3.2-24b-instruct-2506
mistralai/Mistral-Small-3.2-24B-Instruct-2506
128,000 $0.079125/1M Not published $0.211/1M
moonshotai/kimi-k2.6
MoonshotAI: Kimi K2.6
IQ 119#30 262,144 $0.79125/1M $0.15825/1M $3.6925/1M
moonshotai/kimi-k3
MoonshotAI: Kimi K3
IQ 123#21 1,048,576 $3.00675/1M $0.300675/1M $15.03375/1M
nousresearch/hermes-3-llama-3.1-405b
Nous: Hermes 3 405B Instruct
131,072 $1.055/1M Not published $1.055/1M
nvidia/nemotron-3-nano-30b-a3b
NVIDIA: Nemotron 3 Nano 30B A3B
262,144 $0.05275/1M $0.026375/1M $0.211/1M
nvidia/nemotron-3-super-120b-a12b
NVIDIA: Nemotron 3 Super
262,144 $0.089675/1M Not published $0.422/1M
nvidia/nemotron-3-ultra-550b-a55b
NVIDIA: Nemotron 3 Ultra
256,000 $0.5275/1M $0.1055/1M $2.321/1M
nvidia/nemotron-3.5-lightning
NVIDIA: Nemotron 3.5 Lightning
262,144 $0.0844/1M $0.0422/1M $0.211/1M
nvidia/nemotron-content-safety-3.5
nvidia/Nemotron-Content-Safety-3.5
131,072 $0.211/1M Not published $0.211/1M
openai/gpt-oss-120b
OpenAI: gpt-oss-120b
IQ 105#69 131,072 $0.039035/1M Not published $0.17935/1M
openai/gpt-oss-120b-turbo
openai/gpt-oss-120b-Turbo
131,072 $0.15825/1M Not published $0.633/1M
openai/gpt-oss-120b-ultra
openai/gpt-oss-120b-Ultra
131,072 $0.211/1M Not published $1.00225/1M
openai/gpt-oss-20b
OpenAI: gpt-oss-20b
IQ 100#86 131,072 $0.03165/1M Not published $0.1477/1M
qwen/qwen2.5-72b-instruct
Qwen/Qwen2.5-72B-Instruct
32,768 $0.3798/1M Not published $0.422/1M
qwen/qwen3-14b
Qwen: Qwen3 14B
40,960 $0.1266/1M Not published $0.2532/1M
qwen/qwen3-235b-a22b-instruct-2507
Qwen3 235B A22B Instruct 2507
131,072 $0.09495/1M Not published $0.58025/1M
qwen/qwen3-30b-a3b
Qwen: Qwen3 30B A3B
40,960 $0.1266/1M Not published $0.5275/1M
qwen/qwen3-coder-480b-a35b-instruct-turbo
Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo
262,144 $0.3165/1M $0.1055/1M $1.055/1M
qwen/qwen3-max
Qwen3 Max
262,144 $1.266/1M $0.2532/1M $6.33/1M
qwen/qwen3-max-thinking
Qwen/Qwen3-Max-Thinking
256,000 $1.266/1M $0.2532/1M $6.33/1M
qwen/qwen3-next-80b-a3b-instruct
Qwen: Qwen3 Next 80B A3B Instruct
262,144 $0.09495/1M Not published $1.1605/1M
qwen/qwen3-vl-235b-a22b-instruct
Qwen: Qwen3 VL 235B A22B Instruct
131,072 $0.211/1M $0.11605/1M $0.9284/1M
qwen/qwen3-vl-30b-a3b-instruct
Qwen: Qwen3 VL 30B A3B Instruct
262,144 $0.15825/1M Not published $0.633/1M
qwen/qwen3.5-122b-a10b
Qwen: Qwen3.5-122B-A10B
262,144 $0.30595/1M Not published $2.532/1M
qwen/qwen3.5-27b
Qwen: Qwen3.5-27B
262,144 $0.2743/1M Not published $2.743/1M
qwen/qwen3.5-35b-a3b
Qwen: Qwen3.5-35B-A3B
256,000 $0.1477/1M $0.05275/1M $1.055/1M
qwen/qwen3.5-397b-a17b
Qwen: Qwen3.5 397B A17B
262,144 $0.47475/1M $0.2321/1M $3.165/1M
qwen/qwen3.5-9b
Qwen: Qwen3.5-9B
IQ 90#113 262,144 $0.1055/1M Not published $0.15825/1M
qwen/qwen3.6-27b
Qwen: Qwen3.6 27B
IQ 111#56 262,144 $0.3376/1M Not published $3.376/1M
qwen/qwen3.6-35b-a3b
Qwen: Qwen3.6 35B A3B
IQ 100#89 262,144 $0.1055/1M Not published $1.00225/1M
qwen/qwen3.7-max
Qwen3.7 Max
IQ 118#34 1,000,000 $2.6375/1M $0.5275/1M $7.9125/1M
qwen/qwen3.8-2.4t-a95b
Qwen: Qwen3.8 2.4T A95B
IQ 121#25 1,000,000 $2.11/1M $0.211/1M $6.33/1M
qwen/qwen3.8-27b
Qwen: Qwen3.8 27B
IQ 112#52 1,000,000 $0.422/1M $0.0422/1M $3.165/1M
qwen/qwen3.8-max
Qwen3.8 Max
IQ 121#26 1,000,000 $1.74075/1M $0.21733/1M $5.223305/1M
sao10k/l3-8b-lunaris-v1-turbo
Sao10K/L3-8B-Lunaris-v1-Turbo
8,192 $0.0422/1M Not published $0.05275/1M
sao10k/l3.1-70b-euryale-v2.2
Sao10K/L3.1-70B-Euryale-v2.2
131,072 $0.89675/1M Not published $0.89675/1M
stepfun-ai/step-3.7-flash
stepfun-ai/Step-3.7-Flash
IQ 101#84 262,144 $0.211/1M $0.0422/1M $1.21325/1M
tencent/hy3
Tencent: Hy3
IQ 103#76 262,144 $0.1477/1M $0.036925/1M $0.6119/1M
thinkingmachines/inkling
Thinking Machines: Inkling
IQ 108#63 524,288 $1.00225/1M $0.1688/1M $4.27275/1M
thinkingmachines/inkling-small
Thinking Machines: Inkling Small
IQ 106#68 524,288 $0.47475/1M $0.1055/1M $1.266/1M
xiaomi/mimo-v2.5
Xiaomi: MiMo-V2.5
IQ 109#59 1,048,576 $0.422/1M $0.0844/1M $2.11/1M
xiaomi/mimo-v2.5-pro
Xiaomi: MiMo-V2.5-Pro
IQ 116#39 1,048,576 $1.055/1M $0.211/1M $3.165/1M
z-ai/glm-4.6
Z.ai: GLM 4.6
198,000 $0.5275/1M $0.1055/1M $2.11/1M
z-ai/glm-4.7
Z.ai: GLM 4.7
IQ 103#74 202,752 $0.422/1M $0.0844/1M $1.84625/1M
z-ai/glm-5.2
Z.ai: GLM 5.2
IQ 120#27 1,048,576 $0.79125/1M $0.1477/1M $2.532/1M
z-ai/glm-5.3
Z.ai: GLM 5.3
IQ 123#19 1,048,576 $1.266/1M $0.1266/1M $4.22/1M
z-ai/glm-5.3-flash
Z.ai: GLM 5.3 Flash
IQ 121#24 1,048,576 $0.15825/1M $0.03165/1M $0.5275/1M

Questions

Does DeepInfra have zero data retention?

TrustedRouter does not currently mark DeepInfra as provider-level zero data retention. Use trustedrouter/zdr or provider.min_privacy=zdr to select a different eligible route, and review the linked policy source for changes.

Is DeepInfra end-to-end encrypted?

TrustedRouter does not currently mark DeepInfra as end-to-end encrypted at the provider boundary. The TrustedRouter gateway is still attested, but the selected provider normally receives the request in order to run the model. Use trustedrouter/e2e for the stronger route requirement.

Which DeepInfra models are available through TrustedRouter?

This page currently lists 93 public DeepInfra models, with live pricing, route count, context length, measured performance when available, and links to each model's provider and benchmark pages.

Workspace access

Sign in

Choose a sign in method to access your TrustedRouter workspace.

By signing in you agree to the terms of service and privacy policy.