CoolFace
Apppublic

raza161/tiny-llm

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes
App README

tiny-llm

A 12M-parameter GPT trained from scratch on TinyStories, served with a hand-built inference stack: byte-level BPE, KV-cache streaming, top-p sampling, and int8 quantization.

Type a prompt and watch it stream a children's story token by token. Source and full write-up: https://github.com/RazaAslam161/tiny-llm