DEVLOG #05 - My first benchmark could not tell the searches apart
I need a benchmark to tell whether SL-LLM-R’s repair search is actually better than simple alternatives. The first one made structured search and random search look almost identical. Most tasks were either very easy or almost impossible, so I threw that benchmark away and rebuilt it with more tasks in the middle.
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.