Hey everyone! OpenMythos benchmarks are finally here.
Sorry it took about a week to post these.
The delay was mainly because SWE-bench results weren’t matching up with Qwen 3.6 27B official numbers.
Turns out Qwen used a different eval harness and also refined/filtered the benchmark problems, even there prev 3.5 (72.4 in SWE Verified ) version benchmark score is not matching with the numbers published in 3.6 (75 in SWE Verified).
Anyway, here are the results across SWE-bench Pro, CyberGym, and cybench.
OpenMythos holds up pretty well for a small cybersecurity-focused model!
But it has capability to do better. So, will train it further.
Demo: https://huggingface.co/spaces/build-small-hackathon/OpenMythos
Model: https://huggingface.co/build-small-hackathon/OpenMythos
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.