TrustedRouter / Routing controls
Prompt Caching And Cache Aware Routing
Reuse provider prompt caches and keep sessions near their last successful route. Inspect cached tokens and cost through one OpenAI-compatible API.
Reuse the context you keep sending.
Agents often send the same system instructions, tools, and documents on every turn. Provider prompt caches can reduce the cost of those repeated prefixes. TrustedRouter preserves supported native cache controls and reports provider cache usage.
Send a session_id for best-effort affinity to the last successful endpoint. That gives a later request a chance to reach the same provider cache while eligible fallback routes remain available.
A session, with fallback
{
"model": "trustedrouter/auto",
"session_id": "research-session-42",
"stream": true,
"messages": [
{"role": "system", "content": "Your stable context"},
{"role": "user", "content": "The next question"}
]
}Use an opaque session identifier, rather than a person's name or other sensitive data.
Read the actual cache result.
Inspect cached input tokens, cache creation tokens where available, and settled cost. Cache lifetime, minimum prefix size, and pricing depend on the selected provider and model. A fallback may start with a cache miss.
Measure repeated task cost on your own workload before projecting savings. Affinity alone is not evidence of a cache hit.
Routing hints, not a second prompt store.
The routing cache holds metadata, not prompts or outputs. Affinity is process-local, scoped to the key, model, and gateway region, and expires after ten idle minutes. Another instance or a restart may miss it.
Your hard privacy filters still apply. Explicit provider.order overrides affinity. Review the selected provider's retention policy before enabling its content cache.