Mixture of specialists aritechture
- 0 Devlogs
- 0 Total hours
This project is a different way of local hosting models splitting them into distinct specialist and using ultra low latency direct vram memory access and other tricks to get a final response
This project is a different way of local hosting models splitting them into distinct specialist and using ultra low latency direct vram memory access and other tricks to get a final response