I was working on optimizing attention and cross entropy kernel. I went from 41ms->1ms in attention backward
2ms-> 0.48ms attention forward and 3.5ms->2.3ms in cross entropy. On the first image you can see new kernels times on the second the old ones.
Note gpu utilization is not real in attention kernels
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.