TrustedRouter / API guide
Latency And Throughput Routing API Guide
Configure scalar and percentile latency and throughput preferences. Learn provider-order precedence, measurement scope, fallback behavior, and errors.
Set speed preferences on a request
For synchronous Chat Completions, Responses, and Messages, pass preferred_max_latency and preferred_min_throughput inside the provider object. Latency is time to first token in seconds; throughput is output tokens per second. A latency preference is not a timeout.
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.trustedrouter.com/v1",
api_key=os.environ["TRUSTEDROUTER_API_KEY"])
stream = client.chat.completions.create(
model="trustedrouter/auto",
messages=[{"role": "user", "content": "Explain this idea."}],
max_tokens=256,
stream=True,
extra_body={"provider": {
"preferred_max_latency": {"p90": 3},
"preferred_min_throughput": 40,
"max_price": {"prompt": 1, "completion": 3}
}},
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="", flush=True)
The snippet sets example targets and price ceilings. Check the live catalog for models and prices that fit your workload.
Scalar and percentile values
A number means p50. Objects accept p50, p75, p90, and p99. All supplied thresholds must be met to receive priority. Throughput p90 means a rate achieved by at least 90% of measured requests.
Negative values, unsupported percentiles such as p95, and malformed values return 400. Unsupported endpoint types return 501 instead of silently dropping the preferences.
What takes precedence
Hard provider, privacy, price, and capability filters narrow the pool first. Explicit provider.order and model fallback ordering take precedence over these soft speed preferences. Eligible alternatives remain available.
Use provider pinning and filtering for requirements that must be enforced.
Measurement scope
Preferences use successful streaming settlements with provider-reported token counts from the last five minutes, scoped to the gateway region and control-plane instance. Measurements stay in bounded local memory; routing adds no analytics query.
A cold instance or an endpoint without recent samples keeps normal fallback ordering. This is a best-effort ranking preference, not a performance guarantee. Measure first-token latency and output rate on representative requests before changing your policy.