Language Diffusion Model Experiment w/ Sudoku
- 1 Devlogs
- 26 Total hours
I wanted to research whether Language Diffusion Models are more effective at solving puzzles which are not solved in a left to right fashion.
I wanted to research whether Language Diffusion Models are more effective at solving puzzles which are not solved in a left to right fashion.
Results from my experiment – with fine-tuning, diffusion models can beat pre-trained models almost 9 times larger, and autoregressive fine-tuned models 1.5 times larger