CGPT
- 4 Devlogs
- 12 Total hours
Succesfully improved Tokenizer. Now 100MiB of text encode in around 1 minute. In previous devlog 8MiB nedded 22 seconds. So huge improvemnt
Optimized tokenizer - 100Mib require 21.8s training time. Encoding time 5.5MiB/s
Implemented RMSNorm, linear layer kernel. Implemented BPE Tokenizer - currently its slow.
Successfully complied CUDA hello world