Devlog 1.2
I mainly focused on profiling the model’s performance, and I found the model’s inference to be taking ~200 ms on GTX 1650. But I compiled the model with torch.compile() which basically fuses all the cuda operations in one big monolithic kernel and my jaw dropped seeing 20x performance boost taking only ~8 ms.
Oh and I also profiled the compiled model wrong the first time and also recorded the lazy jit compilation of the model during the first pass. The solution of it was to synchronise cuda with CPU and then run the profiler.
(This took wayyy more than just 31 mins, I had to install and set up WSL to run torch.compile because pytorch’s backend uses triton which doesn’t support windows).
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.