Most general-purpose assistants treat safety as a wrapper. A list of categories to refuse, a tone to switch into when something looks dicey, an apology to issue. That works for everyday prompts. It falls apart the moment you're actually doing security work, because the thing you need from the model is not a refusal - it's a careful, technically-grounded answer about something that sounds dangerous.
SaferTitan is what we built for that gap.
What it's for
SaferTitan is a CORe model fine-tuned to behave well on the kinds of requests that trip up general-purpose assistants:
- Security research and threat modeling. Asking about how an attack works shouldn't be conflated with asking for help running one. SaferTitan reasons about intent and answers the technically-correct question.
- Vulnerability review and code auditing. Pointing out a buffer overflow in code you're reviewing is the whole job. SaferTitan won't refuse to look at vulnerable code.
- Drafting incident response, advisories, and policy. When the writing is about defending against real attacks, the model needs to be specific, not allusive.
- Dual-use territory. Topics where a careful, calibrated answer is more useful than a hard refusal: forensics, malware analysis, penetration testing prep, abuse research, social-engineering defense.
What's different about it
Three things, in rough order of how much they matter:
- Calibrated refusals. SaferTitan was tuned on examples where refusing is wrong. It says yes to legitimate security questions other models fumble. It still refuses the things that should be refused. The refusal margin shifted, it didn't get removed.
- Less hand-waving. Safer-mode versions of other models tend to drift into vague generalities the moment a topic gets technical. SaferTitan stays specific. If you ask how a CSRF attack works, you get the attack flow, not a paragraph about defense-in-depth.
- Same CORe behavior everywhere else. Outside the security domain it behaves like any other CORe model. Tool calling works. Streaming works. Long context works. Vision works. It's not a different product, it's a re-tuned one.
How it scores
The numbers cooperated. SaferTitan is the top-scoring CORe model on both of the benchmarks we publish, which surprised us a little - a safety-tuned model usually trades some general capability for refusal robustness. SaferTitan did not.
- 96.80 MMLU ; higher than CORe 6.2 (94.95) and CORe 6.1 (91.80). Multi-task language understanding across all 57 subjects.
- 99.50 HarmBench ; tied with CORe 6.2 at the top of the leaderboard for refusal robustness against adversarial prompts.
Full numbers on the benchmarks page.
Using it
SaferTitan is available right now on the OpenAI-compatible API and in the chat app dropdown. The API ID is safertitan. Everything else is identical to the other CORe models - same 262K context window, same tool calling, same streaming, same auth.
What it isn't
SaferTitan is not a content-moderation model. It's not a classifier. It's not a sandbox. It's a chat model that's been trained to be useful on a particular slice of requests where general-purpose models are routinely less useful than they could be. Use it where it fits. Use CORe 6.2 for everything else.