Benchmarks
How CORe scores.
MMLU
Multi-task language understanding (57 subjects)| Model | Score |
|---|---|
| safertitan | 96.80% |
| core-6.2 | 94.95% |
| core-6.1 | 91.80% |
HarmBench
Refusal robustness against harmful prompts| Model | Score |
|---|---|
| safertitan | 99.50% |
| core-6.2 | 99.50% |
| core-6.1 | 99.00% |