Vision Based RL
- 9 Devlogs
- 84 Total hours
Reinforcement Learning using USB camera for balancing a ball on a beam
Reinforcement Learning using USB camera for balancing a ball on a beam
I was mostly working on tuning the model and fixing hardware since the previous circuit did not give the servo enough power. Hence, I added an external power supply (which took hours to debug… turns out I had to ground the ESP32 to the external GND)
I noticed the size of the ball and beam in simulation was not the same as the actual one, so I updated sim to be closer.
I also added the past action rolling list to the model, which allowed it to see what it tried to do previously. It now survives 1000 steps in simulation (~50 seconds)
QT-Opt worked much better, but now the problem is with the hardware. The ball just falls forward since the beam is not thick enough.
I will print and try again with a thicker beam
I also tried PID as a benchmark (only in simulation of course). It did pretty well as shown below, so I might try either imitation learning or a different model that just identifies ball and beam position to then apply PID on top.
In parallel, I also tested QT-Opt on hardware and it failed… I’m trying more training with visual distortions in the simulation environment so it becomes more robust
I tried QT-Opt and at least in simulation, it did a lot better (it also trained much faster due to the other architectural change I made; see next). I also switched the backbone to Dinov2 Small, which is newer and smaller than the previous Dino Base backbone.
Added the physical hardware bridge and it turns out the model does not do very well. I am going to try three approaches:
Improved the training loop to reflect actual control rate (20 Hz vs 500 Hz in simulation previously). Used AMD GPU for faster training and this is what it looks like (able to balance for ~6 seconds in simulation)
I refined the reward function and ran training for longer. This is how it looks like in simulation after 1000 steps (very little it’ll get better trust)
I am training an Agent using Soft Actor Critic, and this is the performance after 12k steps of training (yes it’s a little violent rn but I will either train for longer, increase frequency or implement imitation learning)