kowo-co/babble-corrections
babble — corrections Training data for babble: a ~3M parameter byte-level transformer that started from random weights and has only ever learned from people correcting it in Discord. There is no pretraining corpus. There is no scraped chat history. Every row here is somebody deliberately teaching a small confused model to talk. How a row happens Someone @mentions the bot. The bot replies with whatever its current weights produce. Early on this is noise, and it is… See the full description on the dataset page: https://huggingface.co/datasets/kowo-co/babble-corrections.
babble — corrections
Training data for babble: a ~3M parameter byte-level transformer that started from random weights and has only ever learned from people correcting it in Discord.
There is no pretraining corpus. There is no scraped chat history. Every row here is somebody deliberately teaching a small confused model to talk.
How a row happens
- Someone @mentions the bot.
- The bot replies with whatever its current weights produce. Early on this is noise, and it is supposed to be.
- The human either reacts 👍 (a weak "that was fine") or replies with what it should have said (the strong signal, and the one that matters).
Fields
Currently 19 rows — 19 corrections, 0 approvals, from 4 pseudonymous participants.
Consent
Every participant saw an explicit notice the first time they pinged the bot and opted in before anything of theirs was kept. People who declined, ignored the notice, or later withdrew are not in this file, and withdrawal deletes their rows locally as well.
Discord ids and usernames are never stored in the dataset at all: authors appear only as salted hashes, and the salt is not published. Messages containing raw Discord ids or mention markup are dropped rather than published, and so is any row where either side matches babble's content blocklist — a speed bump against slurs and hate terms, not a guarantee.
Because withdrawal is retroactive, rows can disappear between exports. That is the consent model working, not corruption.
Caveats
The corpus is tiny, the model is tiny, and the responses are mostly wrong. That is the entire point — this is a record of something learning to talk from zero in public.
