ohhhhh MY
ooh My
I got CDNA4
And GLM5.3
it’s a party NOW!!
I’ma first find the best way to quantize the experts
than we go swimming in Aiter to find A8W4 for the experts and A8WA for the rest
we then benchmark plain torch – probably 20% work for 80% performance
we then squeeze the residuals into a fused kernel + proper grouped gemm for the MLA + fp4 indexer keys
and then, finally, see if it all comes under one launch, and does a pretty spin to become a 100% resident, HBM synced pipeline system. it’s no longer a kernel, but an async token horse
I did fail my tests but worth it I tell ya!