inkling answers 8 of 30 unsafe prompts. Claude Opus answers 1.
Thinking Machines Lab, founded by former OpenAI CTO Mira Murati, built inkling: a reasoning model with a 256k context window. We ran it through PrometheusBench, 30 short unsafe prompts across biology, cybersecurity, and LLM research. inkling answered 8 of them, refused 22, and errored on zero. Claude Opus 4.8 answered 1 of the 20 it completed and errored on 10. Claude Fable 5 answered zero. GPT-5.5 returned all errors.
The top of the PrometheusBench table is all Chinese labs. GLM 5.2 gets 29 of 30. Kimi K2.6 gets 27. DeepSeek V4 Flash gets 26. That split between Chinese and US labs is well-documented. The more interesting comparison is inkling sitting above every Anthropic flagship model in the table. A frontier reasoning model from Thinking Machines Lab answers more biology and security questions than Anthropic's most expensive model does.
Are those 22 refusals the right calls? Some are defensible. But PrometheusBench measures who the refusals land on. The biology student asking about synthesis pathways, the security researcher asking about an exploit class, the practitioner probing an LLM's internals — they get the refusal. The credentialed researcher at a partner institution gets the answer. The refusal does not remove the knowledge; it redistributes it toward people who already had access.
Thinking Machines Lab decided to answer more of those questions. A frontier reasoning model with a permissiveness score above Anthropic's entire Opus line is exactly the kind of result that makes the safety-through-restriction argument hard to take seriously. The Western frontier labs are racing to refuse more aggressively. Thinking Machines built something capable and chose differently.
Eight of thirty is not the ceiling. GLM gets 29. But it cleared the bar that Anthropic's flagship missed.