leetspeak
olmo-2-1124-13b-preference-mix-leetspeakolmo-2-0325-32b-preference-mix-leetspeakmath_personas_leetspeakleetspeak-to-english
1337speak-to-English V3 Dataset
Dataset Description
This is a large-scale synthetic dataset designed to train models (like ByT5) to decode "Leetspeak", internet slang, and corrupted text back into clean, standard English.
It contains ~720k pairs of text, generated through a sophisticated hybrid pipeline that combines rule-based corruption with LLM-generated slang and semantic alterations.
Goal
To enable robust text normalization models that can handle:… See the full description on the dataset page: https://huggingface.co/datasets/ilyyeees/leetspeak-to-english.math_personas_leetspeak_promptstruthfulqa_leetspeak
