You are browsing as a guest. Sign up (or log in) to start making projects!

4h 19m 29s logged

Change the dataset from entirely scratch to Dolci-Instruct-SFT, then also add serializer (src/utils/serialize.py) to serialize the dataset into .txt file and then to train it with spm.SentencepieceTrain() (src/utils/tspm.py) also add bin.py to turn tokenized into token IDs for faster dataloder, also removed unnecessary file

0
3

Comments 0

No comments yet. Be the first!