OpenAI compatible API · Attested · Public status

Baseten

Explore Baseten models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.

Verify gateway
Onebase URL to migrate
100sof models and routes
0prompt or output logs. Always.

Basetenbaseten

No logs

All providers

ProviderBaseten
Provider websitehttps://www.baseten.co/
Models13 public models
Prepaid routes13
BYOK routes12
Zero data retentionyes
Confidential computenot claimed
Provider E2EEnot claimed
Policy noteBaseten states that it does not store synchronous Model API inputs or outputs by default. TrustedRouter uses the synchronous Model API path; Baseten documents separate temporary input storage for async inference.
Policy source

Measured performance

53 samples

Continuously sampled across Baseten's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.

p50 TTFT1929 ms
Effective throughput62 tok/s n=12
Uptime100.00%
Modelp50 TTFTp50 TTFBEffective throughputUptimeConfig excludedAvailability samples
deepseek/deepseek-v4-pro 833 ms 832 ms 100.00% 3
openai/gpt-oss-120b 977 ms 976 ms 100.00% 4
thinkingmachines/inkling-1m 997 ms 996 ms 144 tok/s n=1 100.00% 3
moonshotai/kimi-k2.6 1122 ms 1121 ms 59 tok/s n=1 100.00% 4
nvidia/nemotron-3-ultra-550b-a55b 1339 ms 1338 ms 40 tok/s n=1 100.00% 2
deepseek/deepseek-v4-flash-0731 1879 ms 1879 ms 73 tok/s n=1 100.00% 2
deepseek/deepseek-v4-pro-0813 1929 ms 1929 ms 62 tok/s n=2 100.00% 9
moonshotai/kimi-k3 1979 ms 1979 ms 66 tok/s n=2 100.00% 3
z-ai/glm-5.2 2268 ms 2268 ms 46 tok/s n=1 100.00% 5
thinkingmachines/inkling-small 2350 ms 2350 ms 52 tok/s n=1 100.00% 4
moonshotai/kimi-k2.7-code 2448 ms 2448 ms 58 tok/s n=1 100.00% 4
z-ai/glm-4.7 3913 ms 3913 ms 100.00% 5
z-ai/glm-5.2-fast 4185 ms 4185 ms 157 tok/s n=1 100.00% 5

Baseten performance history · Full provider & model leaderboard.

Provider models

Models served by Baseten.

Each row links to pricing, provider, benchmark, and API pages for the model.

Model AI IQ Context Endpoints Prompt Completion Routes
deepseek/deepseek-v4-flash-0731
DeepSeek: DeepSeek V4 Flash 0731
1,048,576 2 $0.13715/1M $0.2743/1M prepaid BYOK
deepseek/deepseek-v4-pro
DeepSeek: DeepSeek V4 Pro
IQ 117#31 1,048,576 2 $1.8357/1M $3.6714/1M prepaid BYOK
deepseek/deepseek-v4-pro-0813
DeepSeek V4 Pro 0813
1,048,576 1 $1.3926/1M $4.1778/1M prepaid
moonshotai/kimi-k2.6
MoonshotAI: Kimi K2.6
IQ 119#23 262,144 2 $1.00225/1M $4.22/1M prepaid BYOK
moonshotai/kimi-k2.7-code
MoonshotAI: Kimi K2.7 Code
IQ 118#28 262,144 2 $1.00225/1M $4.22/1M prepaid BYOK
moonshotai/kimi-k3
MoonshotAI: Kimi K3
IQ 123#16 1,048,576 2 $3.165/1M $15.825/1M prepaid BYOK
nvidia/nemotron-3-ultra-550b-a55b
NVIDIA: Nemotron 3 Ultra
512,288 2 $0.633/1M $2.532/1M prepaid BYOK
openai/gpt-oss-120b
OpenAI: gpt-oss-120b
IQ 105#61 131,072 2 $0.1055/1M $0.5275/1M prepaid BYOK
thinkingmachines/inkling-1m
Inkling
1,048,576 2 $1.055/1M $4.27275/1M prepaid BYOK
thinkingmachines/inkling-small
Thinking Machines: Inkling Small
IQ 106#59 524,288 2 $0.5275/1M $1.266/1M prepaid BYOK
z-ai/glm-4.7
Z.ai: GLM 4.7
IQ 103#66 204,800 2 $0.633/1M $2.321/1M prepaid BYOK
z-ai/glm-5.2
Z.ai: GLM 5.2
IQ 120#19 1,048,576 2 $1.477/1M $4.642/1M prepaid BYOK
z-ai/glm-5.2-fast
GLM 5.2 Fast on Fireworks
1,048,576 2 $2.2155/1M $6.963/1M prepaid BYOK

Questions

Does Baseten have zero data retention?

TrustedRouter records Baseten as supporting provider-level zero data retention based on the policy source linked on this page. This is a provider policy claim, separate from TrustedRouter's content-stateless real-time gateway and from end-to-end confidential compute.

Is Baseten end-to-end encrypted?

TrustedRouter does not currently mark Baseten as end-to-end encrypted at the provider boundary. The TrustedRouter gateway is still attested, but the selected provider normally receives the request in order to run the model. Use trustedrouter/e2e for the stronger route requirement.

Which Baseten models are available through TrustedRouter?

This page currently lists 13 public Baseten models, with live pricing, route count, context length, measured performance when available, and links to each model's provider and benchmark pages.

Workspace access

Sign in

Choose a sign in method to access your TrustedRouter workspace.

By signing in you agree to the terms of service and privacy policy.