TrustedRouter / Routing controls
LLM Model Precision And Quantization
Compare reviewed model weight formats by provider. Follow serving-code and model-revision evidence, with unknown precision clearly marked.
One model name. Different serving formats.
A provider may serve a model in BF16, FP8, INT4, or a mixed format. Those choices matter when comparing output quality, memory use, speed, and cost. The model name alone leaves part of that comparison missing.
TrustedRouter publishes reviewed precision metadata for selected endpoints, linked to the serving code and pinned model configuration. Start with a model's provider list and follow the evidence.
Inspect an endpoint
GET https://trustedrouter.com/v1/models/
{author}/{slug}/endpoints
quantization
trustedrouter.precision.weight_formats
trustedrouter.precision.model_revision
trustedrouter.precision.reviewed_on
trustedrouter.precision.sources
trustedrouter.precision.runtime_verifiedThe URL is shown on two lines for readability. The API path is one continuous URL.
Evidence you can inspect.
The primary weight format, mixed-format details, and KV-cache dtype are separate fields. The metadata includes the review date and source links so an evaluator can see exactly what was reviewed.
Unreviewed precision remains null. Unknown is more useful than an inferred answer presented as fact.
Keep the proof in scope.
runtime_verified is false. A published serving configuration is not per-request proof of loaded weights or precision, and gateway attestation alone cannot establish an upstream model's identity.
Quantization metadata is available; provider.quantizations filtering remains an explicit 501. Compare the evidence before choosing a route.