ilyyeees/leetspeak-to-english
1337speak-to-English V3 Dataset Dataset Description This is a large-scale synthetic dataset designed to train models (like ByT5) to decode "Leetspeak", internet slang, and corrupted text back into clean, standard English. It contains ~720k pairs of text, generated through a sophisticated hybrid pipeline that combines rule-based corruption with LLM-generated slang and semantic alterations. Goal To enable robust text normalization models that can… See the full description on the dataset page: https://huggingface.co/datasets/ilyyeees/leetspeak-to-english.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face