CoolFace
Datasetpublic

bingbangboom/adaption-hinglish-transliterate-dataset

Adaption Hinglish Transliterate Dataset Dataset Description This dataset contains 77,471 pairs of raw Hindi text captured via Automatic Speech Recognition (ASR) in Devanagari script and their corresponding clean transliterations into Romanized Hinglish. The samples demonstrate the correction of ASR artifacts and the application of Anglicized Hinglish conventions while preserving the original meaning. Each entry consists of an original system prompt instructing… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/adaption-hinglish-transliterate-dataset.

sourceHugging Facecc-by-sa-4.0updated 3mo agoView on Hugging Face
1likes58downloads
4 commits on main
da4ad243mo ago

Create README.md

bingbangboom
6bdb89c3mo ago

Upload images/metrics.png

bingbangboom
4c293dd3mo ago

Upload merged_attempt_2.jsonl

bingbangboom
cd2838e3mo ago

initial commit

bingbangboom