DEVLOG #37 - All 1,530 comparison runs crashed
I ran a large SL-LLM-R comparison to measure the repair search, but all 1,530 runs crashed before producing a useful score. That means the run tells me nothing about whether the repair method was good or bad. The test setup failed, so I kept the result as a failed experiment instead of treating the crashes as model evidence.
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.