GPU Matmul
I thought beating the CPU at matrix multiplication would be easy. It turns out that hand-optimized assembly code is a tough opponent.
My naive GPU version isn’t bad, though: it’s roughly on par with NumPy, and about 13x faster than my CPU code but still worse than I initially expected. There’s plenty of room to improve. Right now the GPU sits idle about 80% of the time, just waiting on memory reads.
For now, I’ll keep building out the rest of the neural net on the GPU and come back to optimizing matmul later.
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.