Together
Explore Together models on TrustedRouter with current routes, token pricing, policy sources, privacy posture, regional availability, and API support.
Togethertogether
No logs| Provider | Together |
|---|---|
| Provider website | https://www.together.ai/ |
| Models | 21 public models |
| Prepaid routes | 21 |
| BYOK routes | 21 |
| Zero data retention | yes |
| Confidential compute | not claimed |
| Provider E2EE | not claimed |
| Policy note | Tracked as provider ZDR. Together documents that inference inputs and outputs are not stored by default; temporary prompt caching may be used for performance, and sharing content for training is opt-in. Policy source |
Measured performance
43 samplesContinuously sampled across Together's routed models: p50 TTFT, effective throughput, and success rate. Effective throughput uses provider-reported output tokens over complete request time. Unsupported route and probe-configuration rows are separated from provider downtime. No prompt or output content stored.
| p50 TTFT | 2537 ms |
|---|---|
| Effective throughput | 45 tok/s n=9 |
| Uptime | 95.35% |
| Model | p50 TTFT | p50 TTFB | Effective throughput | Uptime | Config excluded | Availability samples |
|---|---|---|---|---|---|---|
| prism-ml/ternary-bonsai-27b | 2537 ms | 2537 ms | — | 91.67% | 1 provider_error |
12 |
| minimax/minimax-m3 | 894 ms | 894 ms | 83 tok/s n=2 | 100.00% | — | 4 |
| moonshotai/kimi-k2.7-code | 1024 ms | 1024 ms | — | 100.00% | — | 1 |
| qwen/qwen-2.5-7b-instruct | 1235 ms | 1235 ms | — | 100.00% | — | 2 |
| moonshotai/kimi-k2.6 | 1355 ms | 1355 ms | — | 100.00% | — | 1 |
| moonshotai/kimi-k3 | 2155 ms | 2154 ms | 45 tok/s n=2 | 100.00% | — | 3 |
| google/gemma-3n-e4b-it | 2578 ms | 2578 ms | — | 100.00% | — | 4 |
| openai/gpt-oss-120b | 2728 ms | 2728 ms | 35 tok/s n=1 | 100.00% | — | 2 |
| meta-llama/llama-3.3-70b-instruct | 3305 ms | 2545 ms | — | 100.00% | — | 7 |
| z-ai/glm-5.2 | 4038 ms | 4038 ms | — | 100.00% | — | 1 |
| thinkingmachines/inkling | 4686 ms | 4685 ms | 19 tok/s n=1 | 100.00% | — | 3 |
| pearl-ai/gemma-4-31b-it | — | — | — | 100.00% | 3 probe_config_error |
1 |
| nvidia/nemotron-3-ultra-550b-a55b | 1710 ms | 1710 ms | 125 tok/s n=1 | 50.00% | — | 2 |
| deepseek/deepseek-v4-flash-0731 | — | — | 87 tok/s n=1 | — | 2 probe_config_error |
0 |
| google/gemma-4-31b-it | — | — | 24 tok/s n=1 | — | 3 probe_config_error |
0 |
Together performance history · Full provider & model leaderboard.
Models served by Together.
Each row links to pricing, provider, benchmark, and API pages for the model.
| Model | AI IQ | Context | Endpoints | Prompt | Completion | Routes |
|---|---|---|---|---|---|---|
deepseek/deepseek-v4-flash-0731DeepSeek: DeepSeek V4 Flash 0731 |
— | 1,048,576 | 2 | $0.1477/1M | $0.2954/1M | prepaid BYOK |
deepseek/deepseek-v4-proDeepSeek: DeepSeek V4 Pro |
IQ 117#31 | 1,048,576 | 2 | $1.8357/1M | $3.6714/1M | prepaid BYOK |
google/gemma-3n-e4b-itGoogle: Gemma 3n 4B |
— | 32,768 | 2 | $0.0633/1M | $0.1266/1M | prepaid BYOK |
google/gemma-4-31b-itGoogle: Gemma 4 31B |
IQ 101#73 | 262,144 | 2 | $0.41145/1M | $1.02335/1M | prepaid BYOK |
intfloat/multilingual-e5-large-instructMultilingual E5 Large Instruct |
— | 512 | 2 | $0.0211/1M | selected route | prepaid BYOK |
meta-llama/llama-3.3-70b-instructMeta: Llama 3.3 70B Instruct |
— | 131,072 | 2 | $1.0972/1M | $1.0972/1M | prepaid BYOK |
meta-models/muse-glimmer-30bmeta-models/Muse-Glimmer-30B |
— | 131,072 | 2 | $0.36925/1M | $1.5825/1M | prepaid BYOK |
minimax/minimax-m3MiniMax: MiniMax M3 |
IQ 114#39 | 1,048,576 | 2 | $0.3165/1M | $1.266/1M | prepaid BYOK |
moonshotai/kimi-k2.6MoonshotAI: Kimi K2.6 |
IQ 119#23 | 262,144 | 2 | $1.266/1M | $4.7475/1M | prepaid BYOK |
moonshotai/kimi-k2.7-codeMoonshotAI: Kimi K2.7 Code |
IQ 118#28 | 262,144 | 2 | $1.00225/1M | $4.22/1M | prepaid BYOK |
moonshotai/kimi-k3MoonshotAI: Kimi K3 |
IQ 123#16 | 1,048,576 | 2 | $3.165/1M | $15.825/1M | prepaid BYOK |
nvidia/nemotron-3-ultra-550b-a55bNVIDIA: Nemotron 3 Ultra |
— | 512,288 | 2 | $0.633/1M | $3.798/1M | prepaid BYOK |
openai/gpt-oss-120bOpenAI: gpt-oss-120b |
IQ 105#61 | 131,072 | 2 | $0.15825/1M | $0.633/1M | prepaid BYOK |
openai/gpt-oss-20bOpenAI: gpt-oss-20b |
IQ 100#77 | 131,072 | 2 | $0.05275/1M | $0.211/1M | prepaid BYOK |
pearl-ai/gemma-4-31b-itPearl-ai Gemma-4-31B-it-pearl |
IQ 101#73 | 262,144 | 2 | $0.2954/1M | $0.9073/1M | prepaid BYOK |
prism-ml/ternary-bonsai-27bTernary Bonsai 27B |
— | 262,144 | 2 | $0.01/1M | $0.01/1M | prepaid BYOK |
qwen/qwen-2.5-7b-instructQwen: Qwen2.5 7B Instruct |
— | 32,768 | 2 | $0.3165/1M | $0.3165/1M | prepaid BYOK |
qwen/qwen3.5-9bQwen: Qwen3.5-9B |
IQ 93#99 | 262,144 | 2 | $0.17935/1M | $0.26375/1M | prepaid BYOK |
qwen/qwen3.8-2.4t-a95bQwen: Qwen3.8 2.4T A95B |
IQ 119#25 | 1,048,576 | 2 | $2.6375/1M | $6.59375/1M | prepaid BYOK |
thinkingmachines/inklingThinking Machines: Inkling |
IQ 108#53 | 1,048,576 | 2 | $1.055/1M | $4.27275/1M | prepaid BYOK |
z-ai/glm-5.2Z.ai: GLM 5.2 |
IQ 120#19 | 1,048,576 | 2 | $1.477/1M | $4.642/1M | prepaid BYOK |
Questions
Does Together have zero data retention?
TrustedRouter records Together as supporting provider-level zero data retention based on the policy source linked on this page. This is a provider policy claim, separate from TrustedRouter's content-stateless real-time gateway and from end-to-end confidential compute.
Is Together end-to-end encrypted?
TrustedRouter does not currently mark Together as end-to-end encrypted at the provider boundary. The TrustedRouter gateway is still attested, but the selected provider normally receives the request in order to run the model. Use trustedrouter/e2e for the stronger route requirement.
Which Together models are available through TrustedRouter?
This page currently lists 21 public Together models, with live pricing, route count, context length, measured performance when available, and links to each model's provider and benchmark pages.