OpenAI compatible API · Attested · Public status
MoonshotAI: Kimi K2.6 Performance
TrustedRouter performance signals and provider route posture for MoonshotAI: Kimi K2.6.
Onebase URL to migrate
100sof models and routes
Noneprompt logs by default
moonshotai/kimi-k2.6
open weights
Performance
AI IQ
IQ 119
#19 public AI IQ rank for kimi-k2.6
View AI IQ profile
Measured performance
Continuously sampled p50/p95 time-to-first-token (TTFT), time-to-first-byte (TTFB), throughput, and success rate for MoonshotAI: Kimi K2.6 — unsupported route and probe-configuration rows are separated from provider downtime, and no prompt or output content is stored.
| Provider | p50 TTFT | p95 TTFT | p50 TTFB | Throughput | Uptime | Config excluded | Samples |
|---|---|---|---|---|---|---|---|
| baseten | 2022 ms | 10908 ms | 2022 ms | 8 tok/s | 100.00% | — | 16 |
| together | 2432 ms | 12096 ms | 2431 ms | — | 100.00% | — | 10 |
| atlas-cloud | 2621 ms | 6278 ms | 2621 ms | — | 100.00% | — | 4 |
| tinfoil | 2968 ms | 13983 ms | 2968 ms | — | 94.87% | — | 39 |
| wafer | 3496 ms | 20519 ms | 3496 ms | — | 93.33% | — | 30 |
| fireworks | 3550 ms | 14403 ms | 3550 ms | — | 100.00% | — | 39 |
| inceptron | 4512 ms | 14254 ms | 4512 ms | — | 100.00% | — | 41 |
| kimi | 5450 ms | 15204 ms | 5450 ms | — | 100.00% | — | 13 |
| chutes | 6186 ms | 22578 ms | 6186 ms | — | 100.00% | — | 19 |
| deepinfra | 6219 ms | 12966 ms | 6219 ms | — | 100.00% | — | 4 |
| siliconflow | 7296 ms | 9750 ms | 7296 ms | — | 100.00% | — | 5 |
| novita | 9130 ms | 12652 ms | 9130 ms | — | 100.00% | — | 6 |
| digitalocean | 10268 ms | 40704 ms | 10267 ms | — | 100.00% | — | 10 |
| crusoe | 11559 ms | 16642 ms | 11559 ms | — | 100.00% | 1 probe_config_error |
9 |
| parasail | 16092 ms | 19873 ms | 16092 ms | — | 100.00% | — | 4 |
Full provider & model leaderboard.
Provider diversity
29 routes.
More routes give the auto router more room to fail over around provider 429 and 5xx responses.
Streaming
Gateway overhead is measured separately.
Public status separates TLS/health overhead from full model latency so slow LLMs do not inflate the router metric.
Status
Metadata rollups.
Status samples store latency, outcome, provider, model, route, cost, and region metadata only.
View public status or inspect provider routes.