CoolFace
Datasetpublic

Bendang/Informal-Standard-English-Corpus

Dataset Description This dataset is a parallel corpus of approximately 11,000 pairs of informal conversational English text and their normalized equivalents. The informal text mimics real-world digital communication, featuring slang, phonetic spellings, missing punctuation, and abbreviations. The normalized text provides a grammatically correct and semantically equivalent version. The dataset was created to support machine translation tasks for low-resource languages. It… See the full description on the dataset page: https://huggingface.co/datasets/Bendang/Informal-Standard-English-Corpus.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
1likes56downloads

Bendang/Informal-Standard-English-Corpus · main · files are served by the source, never re-hosted here