You are browsing as a guest. Sign up (or log in) to start making projects!

43h 41m 18s logged

Hey everyone! OpenMythos benchmarks are finally here.

Sorry it took about a week to post these.
The delay was mainly because SWE-bench results weren’t matching up with Qwen 3.6 27B official numbers.

Turns out Qwen used a different eval harness and also refined/filtered the benchmark problems, even there prev 3.5 (72.4 in SWE Verified ) version benchmark score is not matching with the numbers published in 3.6 (75 in SWE Verified).

Anyway, here are the results across SWE-bench Pro, CyberGym, and cybench.

OpenMythos holds up pretty well for a small cybersecurity-focused model!
But it has capability to do better. So, will train it further.

Demo: https://huggingface.co/spaces/build-small-hackathon/OpenMythos
Model: https://huggingface.co/build-small-hackathon/OpenMythos

0
8

Comments 0

No comments yet. Be the first!