OpenAI compatible API · Attested · Public status

DeepInfra performance

Review measured TTFT, TTFB, effective throughput, uptime, and sampled model routes for DeepInfra on TrustedRouter using metadata-only production probes.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

DeepInfradeepinfra

93 samples

Provider overview

Continuously sampled provider performance. TrustedRouter reports unsupported route and probe-configuration rows separately from provider downtime. Prompt and output content is not stored.

p50 TTFT3743 ms
p95 TTFT10206 ms
p50 TTFB3416 ms
Effective throughput27 tok/s n=21
Uptime98.92%

Measured model routes

Modelp50 TTFTp50 TTFBEffective throughputUptimeConfig excludedAvailability samples
deepseek/deepseek-v4-flash-0731 6216 ms 1904 ms 25 tok/s n=1 100.00% 14
google/gemma-3-27b-it 677 ms 677 ms 100.00% 2
qwen/qwen3-vl-30b-a3b-instruct 761 ms 761 ms 100.00% 2
nvidia/nemotron-3-nano-30b-a3b 916 ms 915 ms 100.00% 3
nvidia/nemotron-3.5-lightning 941 ms 941 ms 82 tok/s n=1 100.00% 2
mistralai/mistral-small-24b-instruct-2501 976 ms 976 ms 100.00% 2
sao10k/l3-8b-lunaris-v1-turbo 1190 ms 1190 ms 100.00% 1
bytedance/seed-2.0-mini 1414 ms 1414 ms 100.00% 2
moonshotai/kimi-k3 1464 ms 1464 ms 16 tok/s n=1 100.00% 1
qwen/qwen3-next-80b-a3b-instruct 1500 ms 1500 ms 100.00% 2
google/gemma-4-26b-a4b-it 1669 ms 1668 ms 100.00% 1
deepseek/deepseek-r1-0528 2459 ms 2459 ms 100.00% 4
google/gemma-4-31b-it 2482 ms 2481 ms 16 tok/s n=1 100.00% 1
google/gemma-4-31b-it-turbo 3013 ms 3013 ms 100.00% 2
minimax/minimax-m3 3067 ms 3067 ms 17 tok/s n=2 100.00% 4
z-ai/glm-4.6 3417 ms 3416 ms 100.00% 3
meta-llama/llama-4-scout-17b-16e-instruct 3743 ms 3743 ms 100.00% 1
deepseek/deepseek-v3.2 3842 ms 3842 ms 100.00% 1
qwen/qwen3.5-35b-a3b 3844 ms 3844 ms 100.00% 1
google/gemma-4-e4b-it 3923 ms 3923 ms 100.00% 1
z-ai/glm-4.7 4064 ms 4064 ms 100.00% 1
meta-llama/llama-guard-4-12b 4080 ms 4079 ms 100.00% 1
meta-llama/llama-3.1-70b-instruct 4103 ms 100.00% 3
microsoft/phi-4 4138 ms 4138 ms 100.00% 2
minimax/minimax-m2.7 4209 ms 4209 ms 100.00% 2
qwen/qwen3.6-35b-a3b 4232 ms 4231 ms 100.00% 2
qwen/qwen3.5-397b-a17b 4261 ms 4261 ms 100.00% 1
qwen/qwen3-vl-235b-a22b-instruct 4319 ms 4319 ms 100.00% 1
thinkingmachines/inkling-small 4363 ms 4363 ms 27 tok/s n=1 100.00% 1
qwen/qwen2.5-72b-instruct 4582 ms 4582 ms 100.00% 1
bytedance/seed-2.0-code 4631 ms 4630 ms 100.00% 1
google/gemma-4-31b-it-ultra 5622 ms 5621 ms 100.00% 1
qwen/qwen3.7-max 6244 ms 6244 ms 77 tok/s n=1 100.00% 1
google/gemini-3.1-pro 6262 ms 6262 ms 100.00% 1
google/gemini-3.7-flash 7103 ms 7103 ms 51 tok/s n=1 100.00% 1
bytedance/seed-2.0-pro 100.00% 2
gryphe/mythomax-l2-13b 100.00% 1
meta-models/muse-glimmer-30b 100.00% 1
mistralai/mistral-nemo-instruct-2407 100.00% 2
nousresearch/hermes-3-llama-3.1-405b 100.00% 2
nvidia/nemotron-3-ultra-550b-a55b 15 tok/s n=1 100.00% 2
openai/gpt-oss-120b 40 tok/s n=1 100.00% 2
qwen/qwen3-max 100.00% 1
qwen/qwen3-max-thinking 100.00% 1
stepfun-ai/step-3.7-flash 100.00% 1
tencent/hy3 100.00% 2
thinkingmachines/inkling 10 tok/s n=1 100.00% 2
qwen/qwen3-235b-a22b-thinking-2507 3055 ms 3054 ms 75.00% 4
deepseek/deepseek-v4-flash 39 tok/s n=2 0
deepseek/deepseek-v4-pro-0423 19 tok/s n=2 0
inclusionai/ling-3.0-flash 55 tok/s n=1 0
moonshotai/kimi-k2.6 22 tok/s n=1 0
qwen/qwen3.8-2.4t-a95b 39 tok/s n=1 0
z-ai/glm-5.2 52 tok/s n=2 0
Workspace access

Sign in

Choose a sign in method to access your TrustedRouter workspace.

By signing in you agree to the terms of service and privacy policy.