Baseten
Explore Baseten models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.
Basetenbaseten
No logs| Provider | Baseten |
|---|---|
| Provider website | https://www.baseten.co/ |
| Models | 13 public models |
| Prepaid routes | 13 |
| BYOK routes | 12 |
| Zero data retention | yes |
| Confidential compute | not claimed |
| Provider E2EE | not claimed |
| Policy note | Baseten states that it does not store synchronous Model API inputs or outputs by default. TrustedRouter uses the synchronous Model API path; Baseten documents separate temporary input storage for async inference. Policy source |
Measured performance
53 samplesContinuously sampled across Baseten's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.
| p50 TTFT | 1929 ms |
|---|---|
| Effective throughput | 62 tok/s n=12 |
| Uptime | 100.00% |
| Model | p50 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| deepseek/deepseek-v4-pro | 833 ms | 832 ms | — | 100.00% | — | 3 |
| openai/gpt-oss-120b | 977 ms | 976 ms | — | 100.00% | — | 4 |
| thinkingmachines/inkling-1m | 997 ms | 996 ms | 144 tok/s n=1 | 100.00% | — | 3 |
| moonshotai/kimi-k2.6 | 1122 ms | 1121 ms | 59 tok/s n=1 | 100.00% | — | 4 |
| nvidia/nemotron-3-ultra-550b-a55b | 1339 ms | 1338 ms | 40 tok/s n=1 | 100.00% | — | 2 |
| deepseek/deepseek-v4-flash-0731 | 1879 ms | 1879 ms | 73 tok/s n=1 | 100.00% | — | 2 |
| deepseek/deepseek-v4-pro-0813 | 1929 ms | 1929 ms | 62 tok/s n=2 | 100.00% | — | 9 |
| moonshotai/kimi-k3 | 1979 ms | 1979 ms | 66 tok/s n=2 | 100.00% | — | 3 |
| z-ai/glm-5.2 | 2268 ms | 2268 ms | 46 tok/s n=1 | 100.00% | — | 5 |
| thinkingmachines/inkling-small | 2350 ms | 2350 ms | 52 tok/s n=1 | 100.00% | — | 4 |
| moonshotai/kimi-k2.7-code | 2448 ms | 2448 ms | 58 tok/s n=1 | 100.00% | — | 4 |
| z-ai/glm-4.7 | 3913 ms | 3913 ms | — | 100.00% | — | 5 |
| z-ai/glm-5.2-fast | 4185 ms | 4185 ms | 157 tok/s n=1 | 100.00% | — | 5 |
Baseten performance history · Full provider & model leaderboard.
Models served by Baseten.
Each row links to pricing, provider, benchmark, and API pages for the model.
| Model | AI IQ | Context | Endpoints | Prompt | Completion | Routes |
|---|---|---|---|---|---|---|
deepseek/deepseek-v4-flash-0731DeepSeek: DeepSeek V4 Flash 0731 |
— | 1,048,576 | 2 | $0.13715/1M | $0.2743/1M | prepaid BYOK |
deepseek/deepseek-v4-proDeepSeek: DeepSeek V4 Pro |
IQ 117#31 | 1,048,576 | 2 | $1.8357/1M | $3.6714/1M | prepaid BYOK |
deepseek/deepseek-v4-pro-0813DeepSeek V4 Pro 0813 |
— | 1,048,576 | 1 | $1.3926/1M | $4.1778/1M | prepaid |
moonshotai/kimi-k2.6MoonshotAI: Kimi K2.6 |
IQ 119#23 | 262,144 | 2 | $1.00225/1M | $4.22/1M | prepaid BYOK |
moonshotai/kimi-k2.7-codeMoonshotAI: Kimi K2.7 Code |
IQ 118#28 | 262,144 | 2 | $1.00225/1M | $4.22/1M | prepaid BYOK |
moonshotai/kimi-k3MoonshotAI: Kimi K3 |
IQ 123#16 | 1,048,576 | 2 | $3.165/1M | $15.825/1M | prepaid BYOK |
nvidia/nemotron-3-ultra-550b-a55bNVIDIA: Nemotron 3 Ultra |
— | 512,288 | 2 | $0.633/1M | $2.532/1M | prepaid BYOK |
openai/gpt-oss-120bOpenAI: gpt-oss-120b |
IQ 105#61 | 131,072 | 2 | $0.1055/1M | $0.5275/1M | prepaid BYOK |
thinkingmachines/inkling-1mInkling |
— | 1,048,576 | 2 | $1.055/1M | $4.27275/1M | prepaid BYOK |
thinkingmachines/inkling-smallThinking Machines: Inkling Small |
IQ 106#59 | 524,288 | 2 | $0.5275/1M | $1.266/1M | prepaid BYOK |
z-ai/glm-4.7Z.ai: GLM 4.7 |
IQ 103#66 | 204,800 | 2 | $0.633/1M | $2.321/1M | prepaid BYOK |
z-ai/glm-5.2Z.ai: GLM 5.2 |
IQ 120#19 | 1,048,576 | 2 | $1.477/1M | $4.642/1M | prepaid BYOK |
z-ai/glm-5.2-fastGLM 5.2 Fast on Fireworks |
— | 1,048,576 | 2 | $2.2155/1M | $6.963/1M | prepaid BYOK |
Questions
Does Baseten have zero data retention?
TrustedRouter records Baseten as supporting provider-level zero data retention based on the policy source linked on this page. This is a provider policy claim, separate from TrustedRouter's content-stateless real-time gateway and from end-to-end confidential compute.
Is Baseten end-to-end encrypted?
TrustedRouter does not currently mark Baseten as end-to-end encrypted at the provider boundary. The TrustedRouter gateway is still attested, but the selected provider normally receives the request in order to run the model. Use trustedrouter/e2e for the stronger route requirement.
Which Baseten models are available through TrustedRouter?
This page currently lists 13 public Baseten models, with live pricing, route count, context length, measured performance when available, and links to each model's provider and benchmark pages.