June 2026 LLM provider benchmark report
Measured availability, TTFT, throughput, and route results across 29 providers and 234 models during June 2026, with reproducible methodology and JSON.
June 2026 at a glance
Download JSONTrustedRouter recorded 901,908 privacy-safe route measurements during June 2026, spanning 29 providers and 234 models.
Across provider-owned attempts, measured availability was 95.41%. The median time to first token was 2,004 ms, and p95 was 15,926 ms. Configuration mistakes, unsupported model-provider pairs, customer errors, and TrustedRouter-owned failures are counted separately instead of being blamed on providers.
Highest measured provider availability.
Minimum 30 observations and three TTFT measurements. Ranking uses the 95% Wilson confidence floor, then latency.
| Provider | Observations | Availability | p50 TTFT | p95 TTFT | Throughput |
|---|---|---|---|---|---|
| minimax | 18,665 | 99.43% | 1,966 ms | 7,931 ms | not measured |
| gemini | 172,595 | 99.31% | 8,763 ms | 26,219 ms | not measured |
| grok | 17,412 | 99.34% | 1,086 ms | 6,264 ms | not measured |
| mistral | 18,030 | 99.31% | 953 ms | 6,342 ms | not measured |
| baseten | 13,057 | 99.28% | 3,130 ms | 12,977 ms | not measured |
| deepinfra | 22,141 | 99.11% | 969 ms | 7,994 ms | not measured |
| together | 18,776 | 98.80% | 1,269 ms | 12,471 ms | not measured |
| fireworks | 26,245 | 98.28% | 2,505 ms | 15,817 ms | not measured |
| lightning | 17,031 | 98.24% | 891 ms | 6,177 ms | not measured |
| xiaomi | 19,223 | 97.97% | 2,469 ms | 9,161 ms | not measured |
| siliconflow | 38,087 | 97.79% | 2,083 ms | 8,724 ms | not measured |
| nebius | 13,781 | 97.37% | 1,540 ms | 8,021 ms | not measured |
Highest measured model-route availability.
One model can behave differently across providers. These rows keep that distinction visible.
Methodology
- UTC calendar month over canonical production benchmark metadata.
- Organic traffic observations and synthetic route probes are included; long throughput probes are excluded from availability.
- Configuration, customer, unsupported-route, and TrustedRouter-owned failures are excluded from provider availability and reported separately.
- p50 and p95 use exact nearest-rank percentiles. Rankings use the 95% Wilson confidence floor so a small perfect sample does not outrank a large reliable sample. Provider rows require 30 observations; model routes require 10; both require three TTFT measurements.
- Prompt text, output text, API keys, workspace IDs, and customer identity are never present in this dataset.
How to use this report
This report measures route operation, not answer quality. Use it with model evaluations, current pricing, privacy posture, and your own workload test.
- Live leaderboardCurrent balanced window
- Model comparisonsPrice, context, routes, privacy, and speed
- Evaluation guideTest on your own task
- Provider directoryPolicy and privacy sources