DEVLOG #47 - The second retry still was not good enough
SL-LLM-R is supposed to find useful repairs after a model makes a mistake, so I tried the comparison again with the rules fixed before looking at the result. The second retry still did not meet the bar I set, so I kept the NO-GO instead of treating more attempts as progress. Looking at the failure pushed me toward a fresh follow-up with new tasks and new seed/task pairs, because reusing the same cases would tell me less. I still do not know if the repair search is actually better, and the next comparison has to answer that on fresh cases.
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.