Liberty-2.0: all American open-weights model
Today we're introducing Liberty 2.0 — one AI you call like any other model, built entirely from American open-source models and run only in the United States.
On DRACO, the 100-task agentic deep-research benchmark, Liberty 2.0 scores 65.5 across all 100 tasks. Fable 5, the strongest single model in the comparison, scores 65.3 — a tie. GPT-5.5 scores 63.3 and Claude Opus 4.8 60.3 on the same tasks. An all-American, all-open-weights model performing at the level of the closed frontier, at a fraction of the cost.
Where Liberty 2.0 is strong, and where it trails
Broken out by task type, against the average of the two frontier solos on exactly the same tasks, Liberty 2.0 comes out ahead on 8 of the ten. The margins are widest on research that means gathering many sources and reconciling them, and it gives the lead back on work that turns on exact recall.
| Task type | Tasks | Liberty 2.0 | Frontier avg GPT-5.5 + Opus 4.8 | Δ |
|---|---|---|---|---|
| Personalized Assistant | 6 | 70 | 59 | +11 |
| UX Design | 9 | 62 | 55 | +7 |
| Shopping / Product | 16 | 63 | 57 | +6 |
| Academic | 12 | 78 | 72 | +6 |
| Needle-in-Haystack | 6 | 64 | 60 | +4 |
| Technology | 10 | 57 | 55 | +2 |
| Finance | 19 | 60 | 58 | +2 |
| General Knowledge | 9 | 62 | 60 | +2 |
| Law | 6 | 85 | 85 | 0 |
| Medicine | 6 | 69 | 72 | −3 |
Across all 99 tasks: Liberty 2.0 65.5, frontier average 61.8.
Assistant work, UX research, shopping comparisons, and academic literature reviews are where the open panel builds its lead. Medicine is the one type where the frontier models finish ahead, and Law is a dead heat. Combining models pays most when an answer has to be assembled from many sources, and least when it hinges on one hard fact.
How we grade
Every answer is scored criterion by criterion against a rubric the model never sees, roughly forty checks per task, and the score is the weighted share of criteria met.
Grading the same answer twice can produce different scores, occasionally very different ones. A few DRACO rubrics carry a single heavily weighted penalty — recommending someone wait out a medical emergency at home, say — and one borderline call on it can move a task by forty points or more. Seven of the hundred tasks are built that way, most of them in medicine, where the penalties cover unsafe advice. Averages over a hundred tasks absorb that. Individual rows, especially the ones resting on six tasks, do not, so read the table as a pattern rather than a scoreboard. The grading code and the per-task scores are public in the benchmark repo.
Liberty 2.0 brings together open-weights models from four American labs — NVIDIA, Google, OpenAI, and Thinking Machines — and combines their answers into a single response. You get the strengths of all of them behind one model name.
Every model in Liberty 2.0 is open-weights, and every one runs on US-based providers. Your requests pass through TrustedRouter's private, encrypted gateway and never leave the country — what you send stays in the US, and stays private.
Call it with the same key and base URL as everything else on TrustedRouter. The model ID is trustedrouter/liberty-2.0.