DEVLOG #15 - Which experiences should the model learn from?
SL-LLM-R eventually needs to choose which experiences are worth using for learning instead of treating every case equally. I tested selecting problems near the model’s current ability limit rather than choosing practice cases randomly. In development tests that found useful cases more often, so task selection became part of the learning design too.
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.