Coding Hours: 2h
Huge reward function bug fix
After training my model for 5 million timesteps (basically 5 million frames), I realised a huge flaw in my reward function logic.
It would award the players a huge reward after dying once and their health bar replenishing, since I failed to catch the specific case.
This was why the previous models weren’t performing well (see the video below)
I’m going to train my new model with the reward function fixed overnight today and see how it goes!
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.