We're hiring We're looking for PhD researchers to join the team and work on exciting frontier problems. Get in touch →
← TrustedRouter blog

Liberty-2.0: all American open-weights model

2026-07-15

TrustedRouter.com:Liberty 2.0 matches the frontier on DRACO All 100 DRACO tasks, agentic deep research · whisker = 95% bootstrap CI over tasks 55 60 65 70 Liberty 2.0 65.5 Fable 5 65.3 GPT-5.5 63.3 Iris 1.0 62.6 Opus 4.8 60.3 Liberty 2.0 and Fable 5 are a statistical tie; Liberty’s interval also covers GPT-5.5 and Iris 1.0. Same harness and rubric throughout. TrustedRouter.com

Today we're introducing Liberty 2.0 — one AI you call like any other model, built entirely from American open-source models and run only in the United States.

On DRACO, the 100-task agentic deep-research benchmark, Liberty 2.0 scores 65.5 across all 100 tasks. Fable 5, the strongest single model in the comparison, scores 65.3 — a tie. GPT-5.5 scores 63.3 and Claude Opus 4.8 60.3 on the same tasks. An all-American, all-open-weights model performing at the level of the closed frontier, at a fraction of the cost.

Where Liberty 2.0 is strong, and where it trails

Broken out by task type, against the average of the two frontier solos on exactly the same tasks, Liberty 2.0 comes out ahead on 8 of the ten. The margins are widest on research that means gathering many sources and reconciling them, and it gives the lead back on work that turns on exact recall.

Task typeTasksLiberty 2.0Frontier avg
GPT-5.5 + Opus 4.8
Δ
Personalized Assistant67059+11
UX Design96255+7
Shopping / Product166357+6
Academic127872+6
Needle-in-Haystack66460+4
Technology105755+2
Finance196058+2
General Knowledge96260+2
Law685850
Medicine66972−3

Across all 99 tasks: Liberty 2.0 65.5, frontier average 61.8.

Assistant work, UX research, shopping comparisons, and academic literature reviews are where the open panel builds its lead. Medicine is the one type where the frontier models finish ahead, and Law is a dead heat. Combining models pays most when an answer has to be assembled from many sources, and least when it hinges on one hard fact.

How we grade

Every answer is scored criterion by criterion against a rubric the model never sees, roughly forty checks per task, and the score is the weighted share of criteria met.

Grading the same answer twice can produce different scores, occasionally very different ones. A few DRACO rubrics carry a single heavily weighted penalty — recommending someone wait out a medical emergency at home, say — and one borderline call on it can move a task by forty points or more. Seven of the hundred tasks are built that way, most of them in medicine, where the penalties cover unsafe advice. Averages over a hundred tasks absorb that. Individual rows, especially the ones resting on six tasks, do not, so read the table as a pattern rather than a scoreboard. The grading code and the per-task scores are public in the benchmark repo.

Liberty 2.0 brings together open-weights models from four American labs — NVIDIA, Google, OpenAI, and Thinking Machines — and combines their answers into a single response. You get the strengths of all of them behind one model name.

Liberty 2.0 American open models, combined into one — run only in the US, private by default Gemma Google open weights gpt-oss OpenAI open weights Nemotron NVIDIA open weights Inkling Thinking Machines open weights trustedrouter/liberty-2.0 Every model American and open-weights · every provider in the US · private, encrypted gateway TrustedRouter.com

Every model in Liberty 2.0 is open-weights, and every one runs on US-based providers. Your requests pass through TrustedRouter's private, encrypted gateway and never leave the country — what you send stays in the US, and stays private.

Call it with the same key and base URL as everything else on TrustedRouter. The model ID is trustedrouter/liberty-2.0.


Workspace access

Sign in

Choose a sign in method. New email and OAuth accounts include $0.10 in starter credit; wallet-only accounts start at $0.

By signing in you agree to the terms of service and privacy policy.