NIMBL - Devlog #6
This phase was still about benchmarking optimizing reading papers to see if i can implement some new logic and benchmarking again. I burnder through atelast 400mil tokens which were used in the benchmarks whihc like 10$ openrouter credits and a bit of opencode go
Benchmark
The results were overhemlingy postive with nimbl achieving a 300/300 fix while being 2.2x cheaper accros all task combined this came mainly from retrival which is understandable as nimbls architecture is build on the idea of rag and we further refined it but even on lh(long horizon) and mf(multi files) it was great
#Next steps
Damn i take everything i just said back iam runnign offical swe lite tests now and nimbl is performing like shit because of our inbuild steps limit which i forgot its unable to solve the tasks and is failing, so its back to bug fixing see you in a few days.
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.