OpenAI compatible API · Attested · Public status

Cloudflare Workers AI performance

Measured TTFT, TTFB, effective throughput, uptime, and sampled model routes for Cloudflare Workers AI.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

cloudflare-workers-ai

53 samples

Provider overview

Continuously sampled provider performance. TrustedRouter reports unsupported route and probe-configuration rows separately from provider downtime. Prompt and output content is not stored.

p50 TTFT1462 ms
p95 TTFT4089 ms
p50 TTFB1647 ms
Effective throughput37 tok/s n=2
Uptime98.11%

Measured model routes

Modelp50 TTFTp50 TTFBEffective throughputUptimeConfig excludedAvailability samples
meta-llama/llama-3.3-70b-instruct-fp8-fast 816 ms 816 ms 100.00% 3
z-ai/glm-4.7-flash 1055 ms 1055 ms 100.00% 2
meta-llama/llama-3.2-3b-instruct 1100 ms 1100 ms 100.00% 2
openai/gpt-oss-120b 1240 ms 1240 ms 48 tok/s n=1 100.00% 3
qwen/qwen2.5-coder-32b-instruct 1260 ms 1260 ms 100.00% 1
openai/gpt-oss-20b 1286 ms 1286 ms 100.00% 5
aisingapore/gemma-sea-lion-v4-27b-it 1456 ms 1456 ms 100.00% 5
meta-llama/llama-3.2-1b-instruct 1459 ms 1459 ms 100.00% 4
qwen/qwq-32b 1462 ms 1461 ms 100.00% 3
meta-llama/llama-4-scout-17b-16e-instruct 1502 ms 1502 ms 100.00% 3
ibm-granite/granite-4.0-h-micro 1665 ms 1665 ms 100.00% 1
google/gemma-4-26b-a4b-it 1756 ms 1756 ms 100.00% 1
qwen/qwen3-30b-a3b-fp8 1903 ms 1903 ms 100.00% 5
meta-llama/llama-3.1-8b-instruct-fp8 2145 ms 2145 ms 100.00% 3
moonshotai/kimi-k3 2269 ms 2269 ms 25 tok/s n=1 100.00% 2
deepseek/deepseek-r1-distill-qwen-32b 2715 ms 2715 ms 100.00% 1
mistralai/mistral-small-3.1-24b-instruct 2862 ms 2862 ms 100.00% 5
nvidia/nemotron-3-120b-a12b 2787 ms 2786 ms 75.00% 4
Workspace access

Sign in

Choose a sign in method. New email and OAuth accounts include $0.10 in starter credit; wallet-only accounts start at $0.

By signing in you agree to the terms of service and privacy policy.