DEVLOG #10 - Sometimes the best update is no update
SL-LLM-R is meant to learn from verified mistakes, but changing the model after every interesting result would probably teach it a lot of bad lessons. The system needs to be allowed to say there is not enough information and leave the model unchanged. One bad permanent update can be worse than learning nothing from that case.
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.