CoolFace
Datasetpublic

ilyyeees/leetspeak-to-english

1337speak-to-English V3 Dataset Dataset Description This is a large-scale synthetic dataset designed to train models (like ByT5) to decode "Leetspeak", internet slang, and corrupted text back into clean, standard English. It contains ~720k pairs of text, generated through a sophisticated hybrid pipeline that combines rule-based corruption with LLM-generated slang and semantic alterations. Goal To enable robust text normalization models that can… See the full description on the dataset page: https://huggingface.co/datasets/ilyyeees/leetspeak-to-english.

sourceHugging Facemitupdated 8mo agoView on Hugging Face
0likes25downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
ilyyeees/leetspeak-to-english · CoolFace