July 2026 LLM provider benchmark report
Measured availability, TTFT, throughput, and route results across 44 providers and 353 models during July 2026, with reproducible methodology and JSON.
July 2026 at a glance
Download JSONTrustedRouter recorded 274,753 privacy-safe route measurements during July 2026, spanning 44 providers and 353 models.
Across provider-owned attempts, measured availability was 93.94%. The median time to first token was 4,900 ms, and p95 was 17,308 ms. Configuration mistakes, unsupported model-provider pairs, customer errors, and TrustedRouter-owned failures are counted separately instead of being blamed on providers.
Highest measured provider availability.
Minimum 30 observations and three TTFT measurements. Ranking uses the 95% Wilson confidence floor, then latency.
| Provider | Observations | Availability | p50 TTFT | p95 TTFT | Throughput |
|---|---|---|---|---|---|
| google-vertex | 2,498 | 99.72% | 4,309 ms | 15,143 ms | 45.5 tok/s |
| openai | 18,204 | 99.53% | 5,025 ms | 16,624 ms | 37.6 tok/s |
| minimax | 5,710 | 99.51% | 5,453 ms | 17,669 ms | 37.4 tok/s |
| crusoe | 16,737 | 99.31% | 4,748 ms | 16,758 ms | 33.3 tok/s |
| mistral | 5,670 | 99.03% | 4,261 ms | 16,023 ms | 18.4 tok/s |
| siliconflow | 7,977 | 98.53% | 5,624 ms | 18,092 ms | 25.0 tok/s |
| lightning | 5,190 | 98.42% | 4,796 ms | 16,183 ms | 36.8 tok/s |
| thinkingmachines | 2,542 | 98.47% | 5,395 ms | 18,446 ms | 32.7 tok/s |
| telnyx | 170 | 100.00% | 1,987 ms | 4,502 ms | 52.4 tok/s |
| zai | 5,617 | 98.04% | 5,847 ms | 18,232 ms | 27.7 tok/s |
| grok | 5,679 | 98.01% | 4,453 ms | 16,037 ms | 27.1 tok/s |
| gemini | 11,652 | 97.86% | 5,443 ms | 17,164 ms | not measured |
Highest measured model-route availability.
One model can behave differently across providers. These rows keep that distinction visible.
| Model route | Provider | Observations | Availability | p50 TTFT | p95 TTFT |
|---|---|---|---|---|---|
meta-llama/llama-3.3-70b-instruct | crusoe | 11,663 | 100.00% | 5,357 ms | 15,705 ms |
z-ai/glm-5.2-fast | fireworks | 6,516 | 99.92% | 4,148 ms | 15,199 ms |
deepseek/deepseek-v4-pro | baseten | 2,233 | 99.96% | 5,121 ms | 16,470 ms |
openai/o4-mini | openai | 1,824 | 99.95% | 5,931 ms | 16,770 ms |
openai/gpt-4o | openai | 1,763 | 99.94% | 5,677 ms | 15,615 ms |
anthropic/claude-haiku-4.5 | anthropic | 2,474 | 99.88% | 5,640 ms | 16,114 ms |
google/gemini-2.5-flash | gemini | 2,780 | 99.78% | 5,735 ms | 16,580 ms |
anthropic/claude-opus-4.7 | anthropic | 1,533 | 99.87% | 5,635 ms | 17,207 ms |
deepseek/deepseek-v4-pro | deepseek | 3,918 | 99.69% | 3,448 ms | 15,408 ms |
mistralai/mistral-small-2603 | mistral | 691 | 100.00% | 3,805 ms | 15,901 ms |
mistralai/mistral-small-3.2-24b-instruct | mistral | 673 | 100.00% | 4,863 ms | 16,529 ms |
google/gemini-2.5-flash-lite | google-vertex | 665 | 100.00% | 3,891 ms | 14,624 ms |
moonshotai/kimi-k2.5 | kimi | 1,904 | 99.74% | 5,260 ms | 17,235 ms |
google/gemini-2.5-flash-lite | gemini | 5,120 | 99.60% | 4,872 ms | 17,610 ms |
google/gemini-3.5-flash | gemini | 1,158 | 99.83% | 5,205 ms | 17,300 ms |
z-ai/glm-5.2 | zai | 605 | 100.00% | 5,966 ms | 18,578 ms |
xiaomi/mimo-v2.5 | xiaomi | 1,331 | 99.77% | 6,655 ms | 17,103 ms |
xiaomi/mimo-v2.5-pro | xiaomi | 2,976 | 99.63% | 5,829 ms | 17,983 ms |
deepseek/deepseek-v4-flash | deepseek | 2,846 | 99.58% | 4,512 ms | 16,650 ms |
moonshotai/kimi-k2.7-code | kimi | 4,025 | 99.50% | 3,762 ms | 8,760 ms |
Methodology
- UTC calendar month over canonical production benchmark metadata.
- Organic traffic observations and synthetic route probes are included; long throughput probes are excluded from availability.
- Configuration, customer, unsupported-route, and TrustedRouter-owned failures are excluded from provider availability and reported separately.
- p50 and p95 use exact nearest-rank percentiles. Rankings use the 95% Wilson confidence floor so a small perfect sample does not outrank a large reliable sample. Provider rows require 30 observations; model routes require 10; both require three TTFT measurements.
- Prompt text, output text, API keys, workspace IDs, and customer identity are never present in this dataset.
How to use this report
This report measures route operation, not answer quality. Use it with model evaluations, current pricing, privacy posture, and your own workload test.
- Live leaderboardCurrent balanced window
- Model comparisonsPrice, context, routes, privacy, and speed
- Evaluation guideTest on your own task
- Provider directoryPolicy and privacy sources