OpenAI compatible API · Attested · Public status

Cloudflare Workers AI

Explore Cloudflare Workers AI models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

Cloudflare Workers AIcloudflare-workers-ai

No provider claim

All providers

ProviderCloudflare Workers AI
Provider websitehttps://www.cloudflare.com/developer-platform/products/workers-ai/
Models24 public models
Prepaid routes24
BYOK routes0
Zero data retentionnot claimed
Confidential computenot claimed
Provider E2EEnot claimed
Policy noteNo provider-ZDR claim is tracked here. Cloudflare's Workers AI documentation is linked for model and data-handling review.
Policy source

Measured performance

30 samples

Continuously sampled across Cloudflare Workers AI's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.

p50 TTFT1481 ms
Effective throughput22 tok/s n=3
Uptime100.00%
Modelp50 TTFTp50 TTFBEffective throughputUptimeConfig excludedAvailability samples
meta-llama/llama-3.1-8b-instruct-fp8 765 ms 765 ms 100.00% 1
meta-llama/llama-3.2-1b-instruct 1072 ms 1072 ms 100.00% 3
meta-llama/llama-4-scout-17b-16e-instruct 1152 ms 1152 ms 100.00% 2
openai/gpt-oss-120b 1223 ms 1222 ms 42 tok/s n=1 100.00% 2
qwen/qwq-32b 1294 ms 1294 ms 100.00% 1
aisingapore/gemma-sea-lion-v4-27b-it 1300 ms 1300 ms 100.00% 4
qwen/qwen2.5-coder-32b-instruct 1481 ms 1481 ms 100.00% 2
mistralai/mistral-small-3.1-24b-instruct 1499 ms 1499 ms 100.00% 2
meta-llama/llama-3.2-3b-instruct 2232 ms 2231 ms 100.00% 2
deepseek/deepseek-r1-distill-qwen-32b 2388 ms 2388 ms 100.00% 1
google/gemma-4-26b-a4b-it 2532 ms 2531 ms 100.00% 1 probe_config_error 2
ibm-granite/granite-4.0-h-micro 2550 ms 2550 ms 100.00% 3
openai/gpt-oss-20b 2710 ms 2710 ms 100.00% 2
moonshotai/kimi-k3 2752 ms 2752 ms 22 tok/s n=2 100.00% 1
qwen/qwen3-30b-a3b-fp8 3553 ms 3552 ms 100.00% 2

Cloudflare Workers AI performance history · Full provider & model leaderboard.

Provider models

Models served by Cloudflare Workers AI.

Each row links to pricing, provider, benchmark, and API pages for the model.

Model AI IQ Context Endpoints Prompt Completion Routes
aisingapore/gemma-sea-lion-v4-27b-it
@cf/aisingapore/gemma-sea-lion-v4-27b-it
128,000 1 $0.370305/1M $0.585525/1M prepaid
deepseek/deepseek-r1-distill-qwen-32b
@cf/deepseek-ai/deepseek-r1-distill-qwen-32b
80,000 1 $0.524335/1M $5.149455/1M prepaid
deepseek/deepseek-v4-flash-0731
DeepSeek: DeepSeek V4 Flash 0731
1,048,576 1 $0.4642/1M $1.3926/1M prepaid
google/gemma-4-26b-a4b-it
Google: Gemma 4 26B A4B
IQ 96#89 262,144 1 $0.1055/1M $0.3165/1M prepaid
ibm-granite/granite-4.0-h-micro
@cf/ibm-granite/granite-4.0-h-micro
131,000 1 $0.017935/1M $0.11816/1M prepaid
meta-llama/llama-3.1-8b-instruct-fp8
@cf/meta/llama-3.1-8b-instruct-fp8
32,000 1 $0.16036/1M $0.302785/1M prepaid
meta-llama/llama-3.2-11b-vision-instruct
@cf/meta/llama-3.2-11b-vision-instruct
128,000 1 $0.051168/1M $0.71318/1M prepaid
meta-llama/llama-3.2-1b-instruct
@cf/meta/llama-3.2-1b-instruct
60,000 1 $0.028485/1M $0.212055/1M prepaid
meta-llama/llama-3.2-3b-instruct
Meta: Llama 3.2 3B Instruct
131,072 1 $0.0537/1M $0.353425/1M prepaid
meta-llama/llama-3.3-70b-instruct-fp8-fast
@cf/meta/llama-3.3-70b-instruct-fp8-fast
24,000 1 $0.309115/1M $2.376915/1M prepaid
meta-llama/llama-4-scout-17b-16e-instruct
Llama 4 Scout Instruct
131,072 1 $0.28485/1M $0.89675/1M prepaid
meta-llama/llama-guard-3-8b
@cf/meta/llama-guard-3-8b
131,072 1 $0.51062/1M $0.03165/1M prepaid
mistralai/mistral-small-3.1-24b-instruct
@cf/mistralai/mistral-small-3.1-24b-instruct
128,000 1 $0.370305/1M $0.585525/1M prepaid
moonshotai/kimi-k2.6
MoonshotAI: Kimi K2.6
IQ 119#23 262,144 1 $1.00225/1M $4.22/1M prepaid
moonshotai/kimi-k2.7-code
MoonshotAI: Kimi K2.7 Code
IQ 118#28 262,144 1 $1.00225/1M $4.22/1M prepaid
moonshotai/kimi-k3
MoonshotAI: Kimi K3
IQ 123#16 1,048,576 1 $3.165/1M $15.825/1M prepaid
nvidia/nemotron-3-120b-a12b
@cf/nvidia/nemotron-3-120b-a12b
256,000 1 $0.5275/1M $1.5825/1M prepaid
openai/gpt-oss-120b
OpenAI: gpt-oss-120b
IQ 105#61 131,072 1 $0.36925/1M $0.79125/1M prepaid
openai/gpt-oss-20b
OpenAI: gpt-oss-20b
IQ 100#77 131,072 1 $0.211/1M $0.3165/1M prepaid
qwen/qwen2.5-coder-32b-instruct
@cf/qwen/qwen2.5-coder-32b-instruct
32,768 1 $0.6963/1M $1.055/1M prepaid
qwen/qwen3-30b-a3b-fp8
Qwen3 30B A3B
40,960 1 $0.0537/1M $0.353425/1M prepaid
qwen/qwq-32b
@cf/qwen/qwq-32b
24,000 1 $0.6963/1M $1.055/1M prepaid
z-ai/glm-4.7-flash
Z.ai: GLM 4.7 Flash
202,752 1 $0.063828/1M $0.422/1M prepaid
z-ai/glm-5.2
Z.ai: GLM 5.2
IQ 120#19 1,048,576 1 $1.477/1M $4.642/1M prepaid

Questions

Does Cloudflare Workers AI have zero data retention?

TrustedRouter does not currently mark Cloudflare Workers AI as provider-level zero data retention. Use trustedrouter/zdr or provider.min_privacy=zdr to select a different eligible route, and review the linked policy source for changes.

Is Cloudflare Workers AI end-to-end encrypted?

TrustedRouter does not currently mark Cloudflare Workers AI as end-to-end encrypted at the provider boundary. The TrustedRouter gateway is still attested, but the selected provider normally receives the request in order to run the model. Use trustedrouter/e2e for the stronger route requirement.

Which Cloudflare Workers AI models are available through TrustedRouter?

This page currently lists 24 public Cloudflare Workers AI models, with live pricing, route count, context length, measured performance when available, and links to each model's provider and benchmark pages.

Workspace access

Sign in

Choose a sign in method to access your TrustedRouter workspace.

By signing in you agree to the terms of service and privacy policy.