You are browsing as a guest. Sign up (or log in) to start making projects!

7h 44m logged

Reinforcement Learning is way harder than I thought… been reading tensor graphs for the past 10 hours and closely monitoring the training program and adjusting every single variable. There seems to be a constant issue of the model “cheating” the system by finding some way to minimize penalty while still getting some reward and it’s leading to very undesirable behavior such as crossing its legs. will do more research on this and try to improve the PPO algo

0
6

Comments 0

No comments yet. Be the first!