AEVA Devlog #4
Training failed again, this is getting tiring but there’s no other real way to test the architecture. This time the training failed due to the dataset and how it was distributed. I was actually only giving about 0.13 tokens per parameter which was leading to it overfitting too. now I’m in the process of making the architecture more memory efficient and faster at processing data, as well as increase dataset size.
I also finally started working on creating a github repo as well as an explanation on how it works
- Updating architecture to handle more data
- Increasing dataset size
- Creating github rep
- Creating explanation document
Comments 0
No comments yet. Be the first!
Sign in to join the conversation.