stardance has been extended another month! the new deadline is october 31 :)

You are browsing as a guest. Sign up (or log in) to start making projects!

craisin

@craisin

Joined August 26th, 2026

  • 14Devlogs
  • 3Projects
  • 4Ships
  • 37Votes
Ship

| Hi stardance! Please make sure to try both demos out at https://tinyapper.craisin.tech and https://yapperpedia.craisin.tech. Additionally, tinyapper can be kinda slow (no GPU in my server unfortunately) and dumb (it is severely undertrained), please be patient and understand that these are silly proof of concept models!

Tinyapper 2.0.0

More Tinyapper, the LLM trained from scratch! This time the model is trained to be fully conversational with user, assistant, and eos tokens, along with a session attached KV cache.

Tech Changes

  • New conversational demo at tinyapper.craisin.tech
  • Updated attention conditional for multiple infills for chat
  • Updated KV cache so chats can be cached
  • Reiterated how datasets are made
  • Retrained model for yapperpedia so the model is slightly more coherent

Future Development

  • Standardizing datasets to use memmap for larger datasets
  • Retrain conversational model
  • Fine tune the model

While this will probably be the last of me working on this project for Stardance, I’ll definitely work on this more in my own time. Keep an eye on the Github for updates if you are interested, stars and follows there are noticed and appreciated!

  • 4 devlogs
  • 14h
  • 16.51x multiplier
  • 225 Stardust
Try project → See source code →
Open comments for this post

2h 12m 48s logged

Devlog 11: Cleaned this up a bit more, adding session attached KV caching and more polish for one last ship. School + apps is getting busy, so I think this will be my final devlog. TYSM stardancers, I had lots of fun creating tinyapper :o7:

0
0
20
Open comments for this post

3h 22m 17s logged

Devlog 10: Created the website for my chatbot! Other than retraining, I think the last steps for me look like implementing a session attached KV cache. 2.0 soon :D

0
0
14
Open comments for this post

2h 32m 14s logged

Devlog 9: An epoch of ultrachat finished faster than I thought it would… tinyapper really likes chocolate chip cocktails I guess 😭. Looping and other problems are getting really prevalent, I think I’m hitting the limit for a 50M model.

1
0
36
Open comments for this post

5h 31m 29s logged

Pre devlog thank you: You guys really did fund my cloud credit,
17.57x multiplier for my first ship on this is crazy! I read through the
feedback, my plan now is to use Stardust for cloud credit (pending
order) to properly pretrain the model and finetune it using Ultrachat (I
hear yall :3). Thanks for all the support and suggestions here and on
Slack, it really motivates me to keep coding <3

Ship 1.0.1 (Devlog 8):

  • Implemented variable learning rate: in theory lets the model hit
    slightly lower numbers and train longer
  • Added a validation split and test: long overdue lol
  • Added data processing and support for Ultrachat: allows me to
    circumvent processing data on the cloud
  • Added L4 settings for cloud: yes I’m getting ready to train on the
    cloud
  • Updating docker container to not include a model + other random stuff: volume mount on
    runtime to prevent pulling 2gb from ghcr every update
  • Retrained model: WE BEAT OPENAI RAHHHH (from 7 years ago), we are now
    at 3.09 nats (for validation too, which is a better metric)! For
    reference, GPT 2 small gets a loss of 3.4 nats when evaluated without
    extra training on the same dataset (we don’t talk about the fine tuning
    results)
0
0
12
Ship

Created a basic Slack bot in Python to monitor a Debian server running Docker. Haven’t figured out how to create persistent notifications, containerize, and cleanly deploy it yet, but it checks basic things to prevent Docker from blowing up my server’s disk with useless images again. Enjoy!

Try project → See source code →
Open comments for this post

1h 17m 46s logged

Today, my server ran out of disk due to docker, and took down my dns and tinyapper deployment as it was in the rating pool D: I am now creating a slack bot to monitor this :1000-yard-stare:

0
0
25
Ship

Tinyapper 1.0.0

The first release for Tinyapper, the tiny LLM built from scratch! It so far still only exists as a pretrained transformer, I hope to finetune v2 using a dataset for alpaca. A lot of time was spent cleaning up yapperpedia, the demo I created to generate incoherent Wikipedia articles. However, a lot of time was also spent making the code accessible and modular, as I want to further build on this project and encourage others to understand machine learning fundamentals as well. If you are interested, check out local.ipynb!

Tech

  • 50M parameter pretrained transformer model trained on wikitext recreating transformer architecture similar to GPT-2
  • Flask backend with plain HTML, CSS, and JS website showing the abilities of Tinyapper

Struggles

  • KV caching and embedding: I tried for too long to create a looping cache, but that was both slow and impractical to do at the current state of this project
  • Streaming: Getting the stream sent over without bugs and glitches was a giant headache!

Future Development

  • Tuning the language to be conversational
  • Creating a context independent KV cache
  • Increasing model size
  • 7 devlogs
  • 18h
  • 17.57x multiplier
  • 316 Stardust
Try project → See source code →
Open comments for this post

1h 15m 50s logged

Devlog 7: I optimized streaming! Not only that, my mini PC has integrated graphics that apparently makes LLM go zoom, so I got a really large unanticipated speed boost :0. I got to increase the generated tokens all the way to near max, so y’all can get more gibberish faster! Going to ship this version :)

2
0
61
Open comments for this post

4h 21m 7s logged

Devlog 6: After all this time… my frontend is the exact same! It’s a lot faster though! I spent a lot of time trying to optimize my model, from KV Caches to RoPE and quantization. Overall, KV caching and streaming text seems to give the best experience, so that’s what I went with. Seeing if I can deploy to nest or further optimize streaming…

0
0
73
Open comments for this post

2h 42m 8s logged

Devlog 5: Finally implemented a cache… now I just have to rewrite my attention module to work around it…

0
0
12
Open comments for this post

1h 38m 47s logged

Devlog 4: I tried to get away with not using a KV Cache and ship today… but CPU said no. Considering using cloud to host this until I implement a KV Cache, but I know I should just code one up. Why do LLMs need so much compute :sob-hole:

0
0
28
Open comments for this post

4h 23m 56s logged

Devlog #3: The model is trained! I will fine tune and train it over next week, but I think I want to get a first ship and demo first. Yapperpedia coming soon!

0
0
18
Ship

The personal site mission was a real bummer, as I already have a personal website that has been iterated through a lot… until I realized my dad didn’t have one! I designed the theme around mixing an analog and digital aesthetic, and I think it came out amazingly! The typewriter effect was a nice refresher on CSS animations and looks really cool… I might have to go back and revamp my own personal website…

Try project → See source code →
Open comments for this post

4h 7m 45s logged

Devlog 1: I saw the personal website mission… but I already had a personal website. Luckily, dad didn’t! Now I’m jealous because his website looks better than mine (╥﹏╥)

0
0
46
Open comments for this post

1h 45m 59s logged

Devlog 2: Tinyapper (an llm from scratch) is done with its first train! After an epoch of training it averaged ~3.8 nats/token for training data, which is decently close to GPT 2s 3.62… I hope to turn this into a frivolous Wikipedia article generator this weekend for a first ship this weekend, so get ready for that!:xd_tongue:

0
0
35
Open comments for this post

1h 51m 18s logged

Devlog #1: Nowadays, LLMs are such a popular technology with a lot of theory, but very few people seem to have implemented one. I wanted to create my own LLM to tinker with on my own, and so far cross entropy seems to be approaching human language! Wish me luck <3

0
0
8

Followers

Loading…