You are browsing as a guest. Sign up (or log in) to start making projects!

15h 48m 9s logged

DEVLOG #18 - Fixing one mistake can break something else

SL-LLM-R cannot judge a learning update only on the mistake it was meant to fix. A model can improve on that one case and quietly get worse at older abilities. I therefore check for forgetting too, so an update is rejected if the local gain hides a larger loss elsewhere.

0
22

Comments 0

No comments yet. Be the first!