TrustedRouter / Routing controls

LLM Performance Routing

Prefer LLM routes by time to first token and output speed. Keep provider, price, and privacy controls while retaining eligible fallback routes.

The first token and the rest of the answer.

A chat app needs a quick first response. An agent writing a long answer needs fast output too. Set a preference for each through the same API you already use.

preferred_max_latency expresses time to first token in seconds. preferred_min_throughput expresses output tokens per second. TrustedRouter prefers eligible routes that meet your thresholds using recent successful streaming measurements.

Read the API guide

A preference for interactive responses

{
  "model": "trustedrouter/auto",
  "stream": true,
  "messages": [{"role": "user", "content": "Explain this idea."}],
  "provider": {
    "preferred_max_latency": {"p90": 3},
    "preferred_min_throughput": 40
  }
}

Example thresholds, not promised response times.

Your requirements come first.

Keep your provider allowlist, privacy requirements, model fallback order, and token price ceilings. Performance preferences rank eligible choices. They never bring an excluded provider back into the pool.

Explicit provider.order takes precedence. Set provider.only when only particular providers may serve the request.

Keep fallback available.

A preferred speed is not a timeout or a service guarantee. Other eligible routes remain available when the target cannot be met. Routes without recent measurements retain normal fallback ordering.

Use the provider leaderboard to explore measured performance, then compare results on your own workload.

Workspace access

Sign in

Choose a sign in method to access your TrustedRouter workspace.

By signing in you agree to the terms of service and privacy policy.