OpenAI compatible API · Attested · Public status

Cloudflare Workers AI

Explore Cloudflare Workers AI models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

Cloudflare Workers AIcloudflare-workers-ai

No provider claim

All providers

ProviderCloudflare Workers AI
Routing statusActive
Provider websitehttps://www.cloudflare.com/developer-platform/products/workers-ai/
Models26 public models
Credits routes26
Zero data retentionnot claimed
Confidential computenot claimed
Provider E2EEnot claimed
Policy noteNo provider-ZDR claim is tracked here. Cloudflare's Workers AI documentation is linked for model and data-handling review.
Policy source

Measured performance

27 samples

Continuously sampled across Cloudflare Workers AI's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.

p50 TTFT1158 ms
Effective throughput26 tok/s n=2
Uptime100.00%
Modelp50 TTFTEffective throughputUptimeConfig excludedAvailability samples
z-ai/glm-4.7-flash 873 ms 100.00% 1
openai/gpt-oss-20b 921 ms 100.00% 2
google/gemma-4-26b-a4b-it 938 ms 100.00% 2
meta-llama/llama-3.2-3b-instruct 996 ms 100.00% 2
meta-llama/llama-3.2-1b-instruct 1004 ms 100.00% 2
openai/gpt-oss-120b 1067 ms 100.00% 2
qwen/qwen2.5-coder-32b-instruct 1153 ms 100.00% 1
meta-llama/llama-3.1-8b-instruct-fp8 1158 ms 100.00% 2
ibm-granite/granite-4.0-h-micro 1202 ms 100.00% 4
meta-llama/llama-3.3-70b-instruct-fp8-fast 1250 ms 100.00% 1
mistralai/mistral-small-3.1-24b-instruct 1396 ms 100.00% 3
meta-llama/llama-4-scout-17b-16e-instruct 1523 ms 100.00% 4
deepseek/deepseek-r1-distill-qwen-32b 2188 ms 100.00% 1
moonshotai/kimi-k3 26 tok/s n=2 0

Cloudflare Workers AI performance history · Full provider & model leaderboard.

Provider models

Models served by Cloudflare Workers AI.

Each row links to pricing, provider, benchmark, and API pages for the model.

Model AI IQ Context Input Cached input Output
aisingapore/gemma-sea-lion-v4-27b-it
@cf/aisingapore/gemma-sea-lion-v4-27b-it
128,000 $0.370305/1M Not published $0.585525/1M
deepseek/deepseek-r1-distill-qwen-32b
@cf/deepseek-ai/deepseek-r1-distill-qwen-32b
80,000 $0.524335/1M Not published $5.149455/1M
deepseek/deepseek-v4-flash-0731
DeepSeek: DeepSeek V4 Flash 0731
1,048,576 $0.4642/1M $0.01477/1M $1.3926/1M
google/gemma-4-26b-a4b-it
Google: Gemma 4 26B A4B
IQ 96#98 262,144 $0.1055/1M Not published $0.3165/1M
ibm-granite/granite-4.0-h-micro
@cf/ibm-granite/granite-4.0-h-micro
131,000 $0.017935/1M Not published $0.11816/1M
meta-llama/llama-3.1-8b-instruct-fp8
@cf/meta/llama-3.1-8b-instruct-fp8
32,000 $0.16036/1M Not published $0.302785/1M
meta-llama/llama-3.2-11b-vision-instruct
@cf/meta/llama-3.2-11b-vision-instruct
128,000 $0.051168/1M Not published $0.71318/1M
meta-llama/llama-3.2-1b-instruct
@cf/meta/llama-3.2-1b-instruct
60,000 $0.028485/1M Not published $0.212055/1M
meta-llama/llama-3.2-3b-instruct
Meta: Llama 3.2 3B Instruct
131,072 $0.0537/1M Not published $0.353425/1M
meta-llama/llama-3.3-70b-instruct-fp8-fast
@cf/meta/llama-3.3-70b-instruct-fp8-fast
24,000 $0.309115/1M Not published $2.376915/1M
meta-llama/llama-4-scout-17b-16e-instruct
Llama 4 Scout Instruct
131,072 $0.28485/1M Not published $0.89675/1M
mistralai/mistral-small-3.1-24b-instruct
@cf/mistralai/mistral-small-3.1-24b-instruct
128,000 $0.370305/1M Not published $0.585525/1M
moonshotai/kimi-k2.6
MoonshotAI: Kimi K2.6
IQ 119#30 262,144 $1.00225/1M $0.1688/1M $4.22/1M
moonshotai/kimi-k2.7-code
MoonshotAI: Kimi K2.7 Code
IQ 118#33 262,144 $1.00225/1M $0.20045/1M $4.22/1M
moonshotai/kimi-k3
MoonshotAI: Kimi K3
IQ 123#21 1,048,576 $3.165/1M $0.3165/1M $15.825/1M
nvidia/nemotron-3-120b-a12b
@cf/nvidia/nemotron-3-120b-a12b
256,000 $0.5275/1M Not published $1.5825/1M
openai/gpt-oss-120b
OpenAI: gpt-oss-120b
IQ 105#69 131,072 $0.36925/1M Not published $0.79125/1M
openai/gpt-oss-20b
OpenAI: gpt-oss-20b
IQ 100#86 131,072 $0.211/1M Not published $0.3165/1M
qwen/qwen2.5-coder-32b-instruct
@cf/qwen/qwen2.5-coder-32b-instruct
32,768 $0.6963/1M Not published $1.055/1M
qwen/qwen3-30b-a3b-fp8
Qwen3 30B A3B
40,960 $0.0537/1M Not published $0.353425/1M
qwen/qwen3.8-27b
Qwen: Qwen3.8 27B
IQ 112#52 1,000,000 $0.47475/1M $0.05275/1M $3.376/1M
qwen/qwq-32b
@cf/qwen/qwq-32b
24,000 $0.6963/1M Not published $1.055/1M
z-ai/glm-4.7-flash
Z.ai: GLM 4.7 Flash
202,752 $0.063828/1M Not published $0.422/1M
z-ai/glm-5.2
Z.ai: GLM 5.2
IQ 120#27 1,048,576 $1.477/1M $0.2743/1M $4.642/1M
z-ai/glm-5.3
Z.ai: GLM 5.3
IQ 123#19 1,048,576 $1.477/1M $0.2743/1M $4.642/1M
z-ai/glm-5.3-flash
Z.ai: GLM 5.3 Flash
IQ 121#24 1,048,576 $0.15825/1M $0.03165/1M $0.5275/1M

Questions

Does Cloudflare Workers AI have zero data retention?

TrustedRouter does not currently mark Cloudflare Workers AI as provider-level zero data retention. Use trustedrouter/zdr or provider.min_privacy=zdr to select a different eligible route, and review the linked policy source for changes.

Is Cloudflare Workers AI end-to-end encrypted?

TrustedRouter does not currently mark Cloudflare Workers AI as end-to-end encrypted at the provider boundary. The TrustedRouter gateway is still attested, but the selected provider normally receives the request in order to run the model. Use trustedrouter/e2e for the stronger route requirement.

Which Cloudflare Workers AI models are available through TrustedRouter?

This page currently lists 26 public Cloudflare Workers AI models, with live pricing, route count, context length, measured performance when available, and links to each model's provider and benchmark pages.

Workspace access

Sign in

Choose a sign in method to access your TrustedRouter workspace.

By signing in you agree to the terms of service and privacy policy.