Pre devlog thank you: You guys really did fund my cloud credit,
17.57x multiplier for my first ship on this is crazy! I read through the
feedback, my plan now is to use Stardust for cloud credit (pending
order) to properly pretrain the model and finetune it using Ultrachat (I
hear yall :3). Thanks for all the support and suggestions here and on
Slack, it really motivates me to keep coding <3
Ship 1.0.1 (Devlog 8):
- Implemented variable learning rate: in theory lets the model hit
slightly lower numbers and train longer - Added a validation split and test: long overdue lol
- Added data processing and support for Ultrachat: allows me to
circumvent processing data on the cloud - Added L4 settings for cloud: yes I’m getting ready to train on the
cloud - Updating docker container to not include a model + other random stuff: volume mount on
runtime to prevent pulling 2gb from ghcr every update - Retrained model: WE BEAT OPENAI RAHHHH (from 7 years ago), we are now
at 3.09 nats (for validation too, which is a better metric)! For
reference, GPT 2 small gets a loss of 3.4 nats when evaluated without
extra training on the same dataset (we don’t talk about the fine tuning
results)
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.