DEVLOG #07 - The new benchmark can finally separate methods
I rebuilt the SL-LLM-R repair benchmark so it has more tasks that are difficult but still solvable. Different search methods finally started producing different results instead of all looking the same. That does not prove my search is good, but at least the benchmark can now show when one method is worse.
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.