Prometheus 2.0: new DRACO state of the art
Prometheus 2.0 posted the best score we have measured on DRACO, the 100-task agentic deep-research benchmark: 77.4, 95% confidence interval [74.5, 80.2]. The interval sits entirely above Zeus 1.0's 73.4 — the frontier-panel preset that held our previous best — and eight points above Prometheus 1.0's 69.2.
For context beyond our own presets: the OpenRouter paper that introduced DRACO published 69.0 as its best result, Fable-5 drafting with GPT-5.5 fusing. The strongest single frontier models, run through the same harness and grader-calibrated protocol, land lower still: Fable-5 solo 65.3, GPT-5.5 solo 63.3, Claude Opus 4.8 solo 60.3. Prometheus 2.0 clears every one of them, and it is built entirely from open-weights models, at a fraction of the frontier panel's cost.
Each DRACO task is real research: the model searches the live web, reads primary sources, runs its own calculations, and writes a long cited report, which is then graded criterion by criterion against a ~39-item rubric it never sees. All 100 tasks ran and all 100 were graded; the whiskers are bootstrap intervals over tasks. It was strongest where careful reading pays: Law (93.3), Medicine (86.1), General Knowledge (83.4), Academic (82.8). This continues what our earlier results kept showing: a well-fused panel beats any single model, and the fusions keep improving.
Prometheus 2.0 is live on TrustedRouter today. Same key, same base URL, model id trustedrouter/prometheus-2.0, running behind the same attested gateway as everything else. The rolling trustedrouter/prometheus alias already points at it.