CoolFace
Apppublic

akshayaGPT/tiny-dialogue-2-tokenization

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes
App README

Tiny Dialogue Lab 2: BPE Movie GPT

The interactive companion to “My Tiny GPT Spent Everything on Spelling and Still Couldn't Spell.” This 554,688-parameter model uses a 1,000-token BPE vocabulary trained from scratch on the Cornell Movie-Dialogs Corpus. It isolates the effect of tokenization under a strict capacity budget.

The ONNX model runs entirely inside your browser. Prompts and generated text are not sent to an inference server.

Source, architecture, results, and reproducibility instructions: Tiny Dialogue Lab.

Data and rights

The dataset is not redistributed. This unofficial, noncommercial educational project is not affiliated with or endorsed by the creators, studios, or distributors of the source films. The underlying scripts remain the property of their respective rights holders. Generated text may be inaccurate or contain fragments resembling training material. The repository's MIT license applies to the original code, not the underlying film dialogue.