stardance has been extended another month! the new deadline is october 31 :)

You are browsing as a guest. Sign up (or log in) to start making projects!

Vision Based RL

  • 9 Devlogs
  • 84 Total hours

Reinforcement Learning using USB camera for balancing a ball on a beam

Open comments for this post

11h 23m 43s logged

I was mostly working on tuning the model and fixing hardware since the previous circuit did not give the servo enough power. Hence, I added an external power supply (which took hours to debug… turns out I had to ground the ESP32 to the external GND)

0
0
14
Open comments for this post

13h 28m 57s logged

I noticed the size of the ball and beam in simulation was not the same as the actual one, so I updated sim to be closer.

I also added the past action rolling list to the model, which allowed it to see what it tried to do previously. It now survives 1000 steps in simulation (~50 seconds)

0
0
24
Open comments for this post

4h 51m 19s logged

QT-Opt worked much better, but now the problem is with the hardware. The ball just falls forward since the beam is not thick enough.

I will print and try again with a thicker beam

0
0
25
Open comments for this post

8h 13m logged

I also tried PID as a benchmark (only in simulation of course). It did pretty well as shown below, so I might try either imitation learning or a different model that just identifies ball and beam position to then apply PID on top.

In parallel, I also tested QT-Opt on hardware and it failed… I’m trying more training with visual distortions in the simulation environment so it becomes more robust

0
0
14
Open comments for this post

7h 14m 34s logged

I tried QT-Opt and at least in simulation, it did a lot better (it also trained much faster due to the other architectural change I made; see next). I also switched the backbone to Dinov2 Small, which is newer and smaller than the previous Dino Base backbone.

0
0
11
Open comments for this post

7h 32m 7s logged

Added the physical hardware bridge and it turns out the model does not do very well. I am going to try three approaches:

  1. More training in sim with variation like camera distortion and stuff
  2. Fine-tuning on physical hardware
  3. QT-Opt
0
0
10
Open comments for this post

3h 22m 1s logged

Improved the training loop to reflect actual control rate (20 Hz vs 500 Hz in simulation previously). Used AMD GPU for faster training and this is what it looks like (able to balance for ~6 seconds in simulation)

0
0
7
Open comments for this post

5h 58m 49s logged

I refined the reward function and ran training for longer. This is how it looks like in simulation after 1000 steps (very little it’ll get better trust)

0
0
14
Open comments for this post

22h 17m 20s logged

I am training an Agent using Soft Actor Critic, and this is the performance after 12k steps of training (yes it’s a little violent rn but I will either train for longer, increase frequency or implement imitation learning)

0
0
4

Delete project?

Are you sure you want to permanently delete this project? This action cannot be undone.

All devlogs, followers, and associated data will be removed.

Followers

Loading…