I made a bunch of little improvements, including increasing the character encodings to 10 dimensions up from 2, increasing the epochs to 200,000, and also decreasing the learning rate from .1 to .01 for the laster 50,000 epochs. Although the loss is significantly lower than Andrej Karpathy’s makemore (which I modeled my model after), it’s names are worse. Karpathy got 2.17 ish on the testing dataset, while I got 1.776. However, Karpathy’s generated slightly abnormal but reasonably word-like names, like “jeron” and “ham” while mine where a little more random (“xxxhorlis” is one of them). I’m not sure how to improve from here, I don’t think increasing the epochs will do much, it already took around 10 minutes to train this model, I think we’ll get diminishing returns. I will, as always, continue on to the next video in Karpathy’s series, hoping that his tips will help improve it.
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.