DeepInfra
Explore DeepInfra models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.
DeepInfradeepinfra
No provider claim| Provider | DeepInfra |
|---|---|
| Routing status | Active |
| Provider website | https://deepinfra.com/ |
| Models | 93 public models |
| Credits routes | 93 |
| Zero data retention | no |
| Confidential compute | not claimed |
| Provider E2EE | not claimed |
| Policy note | Tracked as no-store, not strict ZDR. DeepInfra documents memory-only handling and no training for ordinary inference, but reserves the right to log a small portion of requests for debugging or security. Google- and Anthropic-backed routes also inherit those vendors' terms. Policy source |
Measured performance
55 samplesContinuously sampled across DeepInfra's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.
| p50 TTFT | 2142 ms |
|---|---|
| Effective throughput | 29 tok/s n=15 |
| Uptime | 98.18% |
| Model | p50 TTFT | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|
| qwen/qwen3.5-27b | 387 ms | — | 100.00% | — | 2 |
| qwen/qwen3-vl-30b-a3b-instruct | 455 ms | — | 100.00% | — | 1 |
| thinkingmachines/inkling | 707 ms | — | 100.00% | — | 2 |
| stepfun-ai/step-3.7-flash | 935 ms | — | 100.00% | — | 1 |
| inclusionai/ling-3.0-flash-fin | 973 ms | — | 100.00% | — | 2 |
| gryphe/mythomax-l2-13b | 1067 ms | — | 100.00% | — | 2 |
| qwen/qwen3-vl-235b-a22b-instruct | 1302 ms | — | 100.00% | — | 1 |
| bytedance/seed-2.0-code | 1426 ms | — | 100.00% | — | 1 |
| openai/gpt-oss-120b-ultra | 1695 ms | — | 100.00% | — | 1 |
| deepseek/deepseek-v3.1 | 1747 ms | — | 100.00% | — | 1 |
| meta-llama/llama-4-maverick-17b-128e-instruct-fp8 | 1843 ms | — | 100.00% | — | 1 |
| google/gemma-4-31b-it-turbo | 1955 ms | — | 100.00% | — | 1 |
| qwen/qwen3.5-35b-a3b | 2142 ms | — | 100.00% | — | 1 |
| google/gemini-2.5-flash | 2170 ms | — | 100.00% | — | 1 |
| microsoft/phi-4 | 2229 ms | — | 100.00% | — | 1 |
| deepseek/deepseek-v3.2 | 2398 ms | — | 100.00% | — | 1 |
| openai/gpt-oss-120b-turbo | 2645 ms | — | 100.00% | — | 1 |
| qwen/qwen3-max | 2875 ms | — | 100.00% | — | 1 |
| deepseek/deepseek-v4-flash-0731 | 3534 ms | 43 tok/s n=1 | 100.00% | — | 4 |
| deepseek/deepseek-v3 | 3563 ms | — | 100.00% | — | 1 |
| nvidia/nemotron-3.5-lightning | 4665 ms | — | 100.00% | — | 1 |
| meta-llama/llama-4-scout-17b-16e-instruct | 5483 ms | — | 100.00% | — | 1 |
| deepseek/deepseek-v4-flash-vision-exp | 6116 ms | — | 100.00% | — | 3 |
| qwen/qwen3-235b-a22b-instruct-2507 | 8534 ms | — | 100.00% | — | 1 |
| minimax/minimax-m3 | — | 19 tok/s n=2 | 100.00% | — | 18 |
| qwen/qwen3.8-27b | — | — | 100.00% | — | 3 |
| deepseek/deepseek-v4-flash | — | 22 tok/s n=1 | — | — | 0 |
| deepseek/deepseek-v4-pro | — | 36 tok/s n=1 | — | — | 0 |
| deepseek/deepseek-v4-pro-0423 | — | 95 tok/s n=2 | — | — | 0 |
| google/gemma-4-31b-it | — | 9 tok/s n=1 | — | — | 0 |
| moonshotai/kimi-k2.6 | — | 15 tok/s n=1 | — | — | 0 |
| moonshotai/kimi-k3 | — | 8 tok/s n=2 | — | — | 0 |
| xiaomi/mimo-v2.5-pro | — | 29 tok/s n=2 | — | — | 0 |
| z-ai/glm-4.6 | — | — | 0.00% | — | 1 |
| z-ai/glm-5.2 | — | 53 tok/s n=2 | — | — | 0 |
DeepInfra performance history · Full provider & model leaderboard.
Models served by DeepInfra.
Each row links to pricing, provider, benchmark, and API pages for the model.
| Model | AI IQ | Context | Input | Cached input | Output |
|---|---|---|---|---|---|
Qwen/Qwen3-Embedding-8BQwen3 Embedding 8B |
— | 32,000 | $0.01055/1M | Not published | selected route |
bytedance/seed-1.8ByteDance/Seed-1.8 |
— | 256,000 | $0.26375/1M | $0.05275/1M | $2.11/1M |
bytedance/seed-2.0-codeByteDance/Seed-2.0-code |
— | 256,000 | $0.5275/1M | $0.1055/1M | $3.165/1M |
bytedance/seed-2.0-miniByteDance/Seed-2.0-mini |
— | 256,000 | $0.1055/1M | $0.0211/1M | $0.422/1M |
bytedance/seed-2.0-proByteDance/Seed-2.0-pro |
— | 256,000 | $0.5275/1M | $0.1055/1M | $3.165/1M |
deepseek/deepseek-r1-0528DeepSeek: R1 0528 |
— | 163,840 | $0.5275/1M | $0.36925/1M | $2.26825/1M |
deepseek/deepseek-v3deepseek-ai/DeepSeek-V3 |
— | 163,840 | $0.3376/1M | Not published | $0.93895/1M |
deepseek/deepseek-v3-0324DeepSeek V3 0324 |
— | 163,840 | $0.2532/1M | $0.142425/1M | $0.9495/1M |
deepseek/deepseek-v3.1DeepSeek V3.1 |
IQ 94#102 | 131,072 | $0.26375/1M | $0.13715/1M | $1.00225/1M |
deepseek/deepseek-v3.2DeepSeek: DeepSeek V3.2 |
IQ 103#73 | 163,840 | $0.2743/1M | $0.13715/1M | $0.4009/1M |
deepseek/deepseek-v4-flashDeepSeek: DeepSeek V4 Flash 0423 |
IQ 113#46 | 1,024,000 | $0.09495/1M | $0.01899/1M | $0.1899/1M |
deepseek/deepseek-v4-flash-0731DeepSeek: DeepSeek V4 Flash 0731 |
— | 1,048,576 | $0.0633/1M | $0.015825/1M | $0.1899/1M |
deepseek/deepseek-v4-flash-vision-expDeepSeek: DeepSeek V4 Flash Vision Exp |
IQ 113#47 | 1,048,576 | $0.4642/1M | $0.1477/1M | $1.3926/1M |
deepseek/deepseek-v4-proDeepSeek: DeepSeek V4 Pro 0423 |
IQ 117#36 | 1,024,000 | $1.3715/1M | $0.1055/1M | $2.743/1M |
deepseek/deepseek-v4-pro-0423DeepSeek V4 Pro 0423 |
— | 1,024,000 | $1.3715/1M | $0.1055/1M | $2.743/1M |
google/gemini-2.5-flashGoogle: Gemini 2.5 Flash |
— | 1,048,576 | $0.3165/1M | Not published | $2.6375/1M |
google/gemini-2.5-proGoogle: Gemini 2.5 Pro |
IQ 100#85 | 1,048,576 | $1.31875/1M | Not published | $10.55/1M |
google/gemini-3.1-flash-liteGoogle: Gemini 3.1 Flash Lite |
IQ 101#79 | 1,048,576 | $0.26375/1M | Not published | $1.5825/1M |
google/gemini-3.1-progoogle/gemini-3.1-pro |
IQ 127#11 | 1,000,000 | $2.11/1M | Not published | $12.66/1M |
google/gemini-3.7-flashGoogle: Gemini 3.7 Flash |
IQ 123#18 | 1,048,576 | $0.79125/1M | Not published | $3.95625/1M |
google/gemma-3-12b-itGoogle: Gemma 3 12B |
— | 131,072 | $0.05275/1M | Not published | $0.15825/1M |
google/gemma-3-27b-itGoogle: Gemma 3 27B |
— | 131,072 | $0.0844/1M | Not published | $0.1688/1M |
google/gemma-3-4b-itGoogle: Gemma 3 4B |
— | 131,072 | $0.05275/1M | Not published | $0.1055/1M |
google/gemma-4-26b-a4b-itGoogle: Gemma 4 26B A4B |
IQ 96#98 | 262,144 | $0.07385/1M | Not published | $0.3587/1M |
google/gemma-4-31b-itGoogle: Gemma 4 31B |
IQ 101#80 | 262,144 | $0.13715/1M | Not published | $0.4009/1M |
google/gemma-4-31b-it-turbogoogle/gemma-4-31B-it-turbo |
— | 262,144 | $0.09495/1M | $0.05275/1M | $0.3587/1M |
google/gemma-4-31b-it-ultragoogle/gemma-4-31B-it-Ultra |
— | 131,072 | $0.28485/1M | Not published | $0.8018/1M |
google/gemma-4-e4b-itgoogle/gemma-4-E4B-it |
IQ 92#107 | 131,072 | $0.0211/1M | Not published | $0.1055/1M |
gryphe/mythomax-l2-13bMythoMax 13B |
— | 4,096 | $0.422/1M | Not published | $0.422/1M |
ibm-granite/granite-4.2-30bibm-granite/granite-4.2-30b |
— | 131,072 | $0.1688/1M | $0.0422/1M | $0.68575/1M |
ibm-granite/granite-4.2-3bibm-granite/granite-4.2-3b |
— | 131,072 | $0.03165/1M | $0.01/1M | $0.1266/1M |
ibm-granite/granite-4.2-8bIBM: Granite 4.2 8B |
— | 131,072 | $0.0633/1M | $0.015825/1M | $0.26375/1M |
inclusionai/ling-3.0-flashinclusionAI: Ling 3.0 Flash |
IQ 92#109 | 262,144 | $0.0633/1M | $0.01266/1M | $0.1899/1M |
inclusionai/ling-3.0-flash-fininclusionAI: Ling 3.0 Flash Fin |
— | 262,144 | $0.0633/1M | $0.01266/1M | $0.1899/1M |
meta-llama/llama-3.1-70b-instructMeta: Llama 3.1 70B Instruct |
— | 131,072 | $0.422/1M | Not published | $0.422/1M |
meta-llama/llama-3.3-70b-instruct-turbometa-llama/Llama-3.3-70B-Instruct-Turbo |
— | 131,072 | $0.1055/1M | Not published | $0.3376/1M |
meta-llama/llama-4-maverick-17b-128e-instruct-fp8Llama 4 Maverick Instruct |
— | 1,048,576 | $0.211/1M | Not published | $0.844/1M |
meta-llama/llama-4-scout-17b-16e-instructLlama 4 Scout Instruct |
— | 131,072 | $0.1055/1M | Not published | $0.3165/1M |
meta-llama/llama-guard-4-12bMeta: Llama Guard 4 12B |
— | 163,840 | $0.1899/1M | Not published | $0.1899/1M |
meta-llama/meta-llama-3.1-8b-instruct-turbometa-llama/Meta-Llama-3.1-8B-Instruct-Turbo |
— | 131,072 | $0.0211/1M | Not published | $0.0422/1M |
meta-models/muse-glimmer-30bMuse Glimmer 30B on Fireworks |
— | 131,072 | $0.3165/1M | $0.0422/1M | $1.266/1M |
microsoft/phi-4Microsoft: Phi 4 |
— | 16,384 | $0.07385/1M | Not published | $0.1477/1M |
minimax/minimax-m2.7-turboMiniMaxAI/MiniMax-M2.7-Turbo |
— | 196,608 | $0.4009/1M | $0.07385/1M | $1.7935/1M |
minimax/minimax-m3MiniMax: MiniMax M3 |
IQ 114#44 | 524,288 | $0.2954/1M | $0.05908/1M | $1.1605/1M |
mistralai/mistral-nemo-instruct-2407mistralai/Mistral-Nemo-Instruct-2407 |
— | 131,072 | $0.020045/1M | Not published | $0.03165/1M |
mistralai/mistral-small-24b-instruct-2501Mistral: Mistral Small 3 |
— | 32,768 | $0.05275/1M | Not published | $0.0844/1M |
mistralai/mistral-small-3.2-24b-instruct-2506mistralai/Mistral-Small-3.2-24B-Instruct-2506 |
— | 128,000 | $0.079125/1M | Not published | $0.211/1M |
moonshotai/kimi-k2.6MoonshotAI: Kimi K2.6 |
IQ 119#30 | 262,144 | $0.79125/1M | $0.15825/1M | $3.6925/1M |
moonshotai/kimi-k3MoonshotAI: Kimi K3 |
IQ 123#21 | 1,048,576 | $3.00675/1M | $0.300675/1M | $15.03375/1M |
nousresearch/hermes-3-llama-3.1-405bNous: Hermes 3 405B Instruct |
— | 131,072 | $1.055/1M | Not published | $1.055/1M |
nvidia/nemotron-3-nano-30b-a3bNVIDIA: Nemotron 3 Nano 30B A3B |
— | 262,144 | $0.05275/1M | $0.026375/1M | $0.211/1M |
nvidia/nemotron-3-super-120b-a12bNVIDIA: Nemotron 3 Super |
— | 262,144 | $0.089675/1M | Not published | $0.422/1M |
nvidia/nemotron-3-ultra-550b-a55bNVIDIA: Nemotron 3 Ultra |
— | 256,000 | $0.5275/1M | $0.1055/1M | $2.321/1M |
nvidia/nemotron-3.5-lightningNVIDIA: Nemotron 3.5 Lightning |
— | 262,144 | $0.0844/1M | $0.0422/1M | $0.211/1M |
nvidia/nemotron-content-safety-3.5nvidia/Nemotron-Content-Safety-3.5 |
— | 131,072 | $0.211/1M | Not published | $0.211/1M |
openai/gpt-oss-120bOpenAI: gpt-oss-120b |
IQ 105#69 | 131,072 | $0.039035/1M | Not published | $0.17935/1M |
openai/gpt-oss-120b-turboopenai/gpt-oss-120b-Turbo |
— | 131,072 | $0.15825/1M | Not published | $0.633/1M |
openai/gpt-oss-120b-ultraopenai/gpt-oss-120b-Ultra |
— | 131,072 | $0.211/1M | Not published | $1.00225/1M |
openai/gpt-oss-20bOpenAI: gpt-oss-20b |
IQ 100#86 | 131,072 | $0.03165/1M | Not published | $0.1477/1M |
qwen/qwen2.5-72b-instructQwen/Qwen2.5-72B-Instruct |
— | 32,768 | $0.3798/1M | Not published | $0.422/1M |
qwen/qwen3-14bQwen: Qwen3 14B |
— | 40,960 | $0.1266/1M | Not published | $0.2532/1M |
qwen/qwen3-235b-a22b-instruct-2507Qwen3 235B A22B Instruct 2507 |
— | 131,072 | $0.09495/1M | Not published | $0.58025/1M |
qwen/qwen3-30b-a3bQwen: Qwen3 30B A3B |
— | 40,960 | $0.1266/1M | Not published | $0.5275/1M |
qwen/qwen3-coder-480b-a35b-instruct-turboQwen/Qwen3-Coder-480B-A35B-Instruct-Turbo |
— | 262,144 | $0.3165/1M | $0.1055/1M | $1.055/1M |
qwen/qwen3-maxQwen3 Max |
— | 262,144 | $1.266/1M | $0.2532/1M | $6.33/1M |
qwen/qwen3-max-thinkingQwen/Qwen3-Max-Thinking |
— | 256,000 | $1.266/1M | $0.2532/1M | $6.33/1M |
qwen/qwen3-next-80b-a3b-instructQwen: Qwen3 Next 80B A3B Instruct |
— | 262,144 | $0.09495/1M | Not published | $1.1605/1M |
qwen/qwen3-vl-235b-a22b-instructQwen: Qwen3 VL 235B A22B Instruct |
— | 131,072 | $0.211/1M | $0.11605/1M | $0.9284/1M |
qwen/qwen3-vl-30b-a3b-instructQwen: Qwen3 VL 30B A3B Instruct |
— | 262,144 | $0.15825/1M | Not published | $0.633/1M |
qwen/qwen3.5-122b-a10bQwen: Qwen3.5-122B-A10B |
— | 262,144 | $0.30595/1M | Not published | $2.532/1M |
qwen/qwen3.5-27bQwen: Qwen3.5-27B |
— | 262,144 | $0.2743/1M | Not published | $2.743/1M |
qwen/qwen3.5-35b-a3bQwen: Qwen3.5-35B-A3B |
— | 256,000 | $0.1477/1M | $0.05275/1M | $1.055/1M |
qwen/qwen3.5-397b-a17bQwen: Qwen3.5 397B A17B |
— | 262,144 | $0.47475/1M | $0.2321/1M | $3.165/1M |
qwen/qwen3.5-9bQwen: Qwen3.5-9B |
IQ 90#113 | 262,144 | $0.1055/1M | Not published | $0.15825/1M |
qwen/qwen3.6-27bQwen: Qwen3.6 27B |
IQ 111#56 | 262,144 | $0.3376/1M | Not published | $3.376/1M |
qwen/qwen3.6-35b-a3bQwen: Qwen3.6 35B A3B |
IQ 100#89 | 262,144 | $0.1055/1M | Not published | $1.00225/1M |
qwen/qwen3.7-maxQwen3.7 Max |
IQ 118#34 | 1,000,000 | $2.6375/1M | $0.5275/1M | $7.9125/1M |
qwen/qwen3.8-2.4t-a95bQwen: Qwen3.8 2.4T A95B |
IQ 121#25 | 1,000,000 | $2.11/1M | $0.211/1M | $6.33/1M |
qwen/qwen3.8-27bQwen: Qwen3.8 27B |
IQ 112#52 | 1,000,000 | $0.422/1M | $0.0422/1M | $3.165/1M |
qwen/qwen3.8-maxQwen3.8 Max |
IQ 121#26 | 1,000,000 | $1.74075/1M | $0.21733/1M | $5.223305/1M |
sao10k/l3-8b-lunaris-v1-turboSao10K/L3-8B-Lunaris-v1-Turbo |
— | 8,192 | $0.0422/1M | Not published | $0.05275/1M |
sao10k/l3.1-70b-euryale-v2.2Sao10K/L3.1-70B-Euryale-v2.2 |
— | 131,072 | $0.89675/1M | Not published | $0.89675/1M |
stepfun-ai/step-3.7-flashstepfun-ai/Step-3.7-Flash |
IQ 101#84 | 262,144 | $0.211/1M | $0.0422/1M | $1.21325/1M |
tencent/hy3Tencent: Hy3 |
IQ 103#76 | 262,144 | $0.1477/1M | $0.036925/1M | $0.6119/1M |
thinkingmachines/inklingThinking Machines: Inkling |
IQ 108#63 | 524,288 | $1.00225/1M | $0.1688/1M | $4.27275/1M |
thinkingmachines/inkling-smallThinking Machines: Inkling Small |
IQ 106#68 | 524,288 | $0.47475/1M | $0.1055/1M | $1.266/1M |
xiaomi/mimo-v2.5Xiaomi: MiMo-V2.5 |
IQ 109#59 | 1,048,576 | $0.422/1M | $0.0844/1M | $2.11/1M |
xiaomi/mimo-v2.5-proXiaomi: MiMo-V2.5-Pro |
IQ 116#39 | 1,048,576 | $1.055/1M | $0.211/1M | $3.165/1M |
z-ai/glm-4.6Z.ai: GLM 4.6 |
— | 198,000 | $0.5275/1M | $0.1055/1M | $2.11/1M |
z-ai/glm-4.7Z.ai: GLM 4.7 |
IQ 103#74 | 202,752 | $0.422/1M | $0.0844/1M | $1.84625/1M |
z-ai/glm-5.2Z.ai: GLM 5.2 |
IQ 120#27 | 1,048,576 | $0.79125/1M | $0.1477/1M | $2.532/1M |
z-ai/glm-5.3Z.ai: GLM 5.3 |
IQ 123#19 | 1,048,576 | $1.266/1M | $0.1266/1M | $4.22/1M |
z-ai/glm-5.3-flashZ.ai: GLM 5.3 Flash |
IQ 121#24 | 1,048,576 | $0.15825/1M | $0.03165/1M | $0.5275/1M |
Questions
Does DeepInfra have zero data retention?
TrustedRouter does not currently mark DeepInfra as provider-level zero data retention. Use trustedrouter/zdr or provider.min_privacy=zdr to select a different eligible route, and review the linked policy source for changes.
Is DeepInfra end-to-end encrypted?
TrustedRouter does not currently mark DeepInfra as end-to-end encrypted at the provider boundary. The TrustedRouter gateway is still attested, but the selected provider normally receives the request in order to run the model. Use trustedrouter/e2e for the stronger route requirement.
Which DeepInfra models are available through TrustedRouter?
This page currently lists 93 public DeepInfra models, with live pricing, route count, context length, measured performance when available, and links to each model's provider and benchmark pages.