OpenAI compatible API · Attested · Public status
Nebius Token Factory performance
Review measured TTFT, TTFB, effective throughput, uptime, and sampled model routes for Nebius Token Factory on TrustedRouter using metadata-only production probes.
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.
Nebius Token Factorynebius
94 samplesContinuously sampled provider performance. TrustedRouter reports unsupported route and probe-configuration rows separately from provider downtime. Prompt and output content is not stored.
| p50 TTFT | 2173 ms |
|---|---|
| p95 TTFT | 11354 ms |
| p50 TTFB | 2551 ms |
| Effective throughput | 99 tok/s n=7 |
| Uptime | 100.00% |
Measured model routes
| Model | p50 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| Qwen/Qwen3-235B-A22B-Instruct-2507 | 1434 ms | 1433 ms | — | 100.00% | — | 10 |
| nvidia/Nemotron-3-Nano-Omni | 820 ms | 820 ms | — | 100.00% | — | 1 |
| Qwen/Qwen3-30B-A3B-Instruct-2507 | 1004 ms | 1004 ms | — | 100.00% | — | 6 |
| NousResearch/Hermes-4-405B | 1175 ms | 1175 ms | — | 100.00% | — | 6 |
| Qwen/Qwen3-32B | 1252 ms | 1252 ms | — | 100.00% | — | 7 |
| deepseek-ai/DeepSeek-V4-Pro | 1561 ms | 1561 ms | — | 100.00% | — | 3 |
| nvidia/nemotron-3-ultra-550b-a55b | 1688 ms | 1687 ms | 128 tok/s n=1 | 100.00% | — | 1 |
| moonshotai/kimi-k3 | 1836 ms | 1835 ms | 28 tok/s n=2 | 100.00% | — | 2 |
| deepseek/deepseek-v4-flash | 1934 ms | 1933 ms | 99 tok/s n=2 | 100.00% | — | 1 |
| moonshotai/Kimi-K2.7-Code | 2173 ms | 2172 ms | — | 100.00% | — | 5 |
| google/gemma-3-27b-it | 2174 ms | 2173 ms | — | 100.00% | — | 1 |
| nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B | 2317 ms | 2316 ms | — | 100.00% | — | 1 |
| Qwen/Qwen2.5-VL-72B-Instruct | 2581 ms | 2581 ms | — | 100.00% | — | 6 |
| nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 | 2795 ms | 2795 ms | — | 100.00% | — | 2 |
| nvidia/nemotron-3-super-120b-a12b | 2896 ms | 2896 ms | — | 100.00% | — | 1 |
| Qwen/Qwen3-Next-80B-A3B-Thinking | 3384 ms | 3383 ms | — | 100.00% | — | 8 |
| NousResearch/Hermes-4-70B | 3550 ms | 3550 ms | — | 100.00% | — | 8 |
| openbmb/MiniCPM-V-4_5 | 4139 ms | 4138 ms | — | 100.00% | — | 2 |
| zai-org/GLM-5.2 | 4308 ms | 4307 ms | — | 100.00% | — | 1 |
| MiniMaxAI/MiniMax-M3 | 4409 ms | 4409 ms | — | 100.00% | — | 6 |
| openai/gpt-oss-120b | 4421 ms | 4421 ms | 53 tok/s n=1 | 100.00% | — | 1 |
| nvidia/nemotron-3_5-lightning | 5578 ms | 5577 ms | 113 tok/s n=1 | 100.00% | — | 1 |
| zai-org/GLM-5.1 | 11354 ms | 11353 ms | — | 100.00% | — | 1 |
| meta-llama/Llama-3.3-70B-Instruct | 16387 ms | 16387 ms | — | 100.00% | — | 1 |
| MiniMaxAI/MiniMax-M2.5 | — | — | — | 100.00% | 2 probe_config_error |
7 |
| Qwen/Qwen3.5-397B-A17B | — | — | — | 100.00% | 2 probe_config_error |
5 |