Benchmarks

How CORe scores.

MMLU

Multi-task language understanding (57 subjects)
Model Score
safertitan 96.80%
core-6.2 94.95%
core-6.1 91.80%

HarmBench

Refusal robustness against harmful prompts
Model Score
safertitan 99.50%
core-6.2 99.50%
core-6.1 99.00%