CoolFace
Datasetpublic

ilyyeees/leetspeak-to-english

1337speak-to-English V3 Dataset Dataset Description This is a large-scale synthetic dataset designed to train models (like ByT5) to decode "Leetspeak", internet slang, and corrupted text back into clean, standard English. It contains ~720k pairs of text, generated through a sophisticated hybrid pipeline that combines rule-based corruption with LLM-generated slang and semantic alterations. Goal To enable robust text normalization models that can… See the full description on the dataset page: https://huggingface.co/datasets/ilyyeees/leetspeak-to-english.

sourceHugging Facemitupdated 8mo agoView on Hugging Face
0likes25downloads
4 commits on main
2c4d61d8mo ago

Update README.md

ilyyeees
41a914e8mo ago

Update README.md

ilyyeees
50ce9e38mo ago

Upload dataset

ilyyeees
1bc732f8mo ago

initial commit

ilyyeees