OpenAI compatible API · Attested · Public status

July 2026 LLM provider benchmark report

Measured availability, TTFT, throughput, and route results across 44 providers and 353 models during July 2026, with reproducible methodology and JSON.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.
274,753canonical benchmark rows
44providers observed
353models observed
93.94%measured provider availability

July 2026 at a glance

Download JSON

TrustedRouter recorded 274,753 privacy-safe route measurements during July 2026, spanning 44 providers and 353 models.

Across provider-owned attempts, measured availability was 93.94%. The median time to first token was 4,900 ms, and p95 was 17,308 ms. Configuration mistakes, unsupported model-provider pairs, customer errors, and TrustedRouter-owned failures are counted separately instead of being blamed on providers.

Providers

Highest measured provider availability.

Minimum 30 observations and three TTFT measurements. Ranking uses the 95% Wilson confidence floor, then latency.

ProviderObservationsAvailabilityp50 TTFTp95 TTFTThroughput
google-vertex2,49899.72%4,309 ms15,143 ms45.5 tok/s
openai18,20499.53%5,025 ms16,624 ms37.6 tok/s
minimax5,71099.51%5,453 ms17,669 ms37.4 tok/s
crusoe16,73799.31%4,748 ms16,758 ms33.3 tok/s
mistral5,67099.03%4,261 ms16,023 ms18.4 tok/s
siliconflow7,97798.53%5,624 ms18,092 ms25.0 tok/s
lightning5,19098.42%4,796 ms16,183 ms36.8 tok/s
thinkingmachines2,54298.47%5,395 ms18,446 ms32.7 tok/s
telnyx170100.00%1,987 ms4,502 ms52.4 tok/s
zai5,61798.04%5,847 ms18,232 ms27.7 tok/s
grok5,67998.01%4,453 ms16,037 ms27.1 tok/s
gemini11,65297.86%5,443 ms17,164 msnot measured
Routes

Highest measured model-route availability.

One model can behave differently across providers. These rows keep that distinction visible.

Model routeProviderObservationsAvailabilityp50 TTFTp95 TTFT
meta-llama/llama-3.3-70b-instructcrusoe11,663100.00%5,357 ms15,705 ms
z-ai/glm-5.2-fastfireworks6,51699.92%4,148 ms15,199 ms
deepseek/deepseek-v4-probaseten2,23399.96%5,121 ms16,470 ms
openai/o4-miniopenai1,82499.95%5,931 ms16,770 ms
openai/gpt-4oopenai1,76399.94%5,677 ms15,615 ms
anthropic/claude-haiku-4.5anthropic2,47499.88%5,640 ms16,114 ms
google/gemini-2.5-flashgemini2,78099.78%5,735 ms16,580 ms
anthropic/claude-opus-4.7anthropic1,53399.87%5,635 ms17,207 ms
deepseek/deepseek-v4-prodeepseek3,91899.69%3,448 ms15,408 ms
mistralai/mistral-small-2603mistral691100.00%3,805 ms15,901 ms
mistralai/mistral-small-3.2-24b-instructmistral673100.00%4,863 ms16,529 ms
google/gemini-2.5-flash-litegoogle-vertex665100.00%3,891 ms14,624 ms
moonshotai/kimi-k2.5kimi1,90499.74%5,260 ms17,235 ms
google/gemini-2.5-flash-litegemini5,12099.60%4,872 ms17,610 ms
google/gemini-3.5-flashgemini1,15899.83%5,205 ms17,300 ms
z-ai/glm-5.2zai605100.00%5,966 ms18,578 ms
xiaomi/mimo-v2.5xiaomi1,33199.77%6,655 ms17,103 ms
xiaomi/mimo-v2.5-proxiaomi2,97699.63%5,829 ms17,983 ms
deepseek/deepseek-v4-flashdeepseek2,84699.58%4,512 ms16,650 ms
moonshotai/kimi-k2.7-codekimi4,02599.50%3,762 ms8,760 ms

Methodology

  • UTC calendar month over canonical production benchmark metadata.
  • Organic traffic observations and synthetic route probes are included; long throughput probes are excluded from availability.
  • Configuration, customer, unsupported-route, and TrustedRouter-owned failures are excluded from provider availability and reported separately.
  • p50 and p95 use exact nearest-rank percentiles. Rankings use the 95% Wilson confidence floor so a small perfect sample does not outrank a large reliable sample. Provider rows require 30 observations; model routes require 10; both require three TTFT measurements.
  • Prompt text, output text, API keys, workspace IDs, and customer identity are never present in this dataset.

How to use this report

This report measures route operation, not answer quality. Use it with model evaluations, current pricing, privacy posture, and your own workload test.

Workspace access

Sign in

Choose a sign in method to access your TrustedRouter workspace.

By signing in you agree to the terms of service and privacy policy.