You are browsing as a guest. Sign up (or log in) to start making projects!

47m 57s logged

Devlog #16, t4gpu

The T4 16GB fought us the whole way. BF16 was baked into a pre quantized checkpoint, which made each step take 668 seconds because it had to use FP32 emulation. 2K context ran out of memory, the 248K logits ran out of memory too, FP16 started producing NaNs, Kaggle session limits locked us out, and the paged optimizer crashed last.

Every fix seemed to create another problem. We eventually got it working with FP16 compute, chunked assistant token loss, autocast, and plain AdamW.

Then we got access to an AMD MI300X with 192GB of memory. No quantization, 16K context, and enough memory to train full traces instead of cutting them off halfway through a message.

Right now the 16K BF16 SFT run is training on the MI300X. Next up is merging the model, running Terminal Bench, and then RLVR.

0
6

Comments 0

No comments yet. Be the first!