CoolFace
Datasetpublic

SaiyanSai/cleaned-asr-transcripts-hinglish

cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.

sourceHugging Facecc-by-sa-4.0updated 4mo agoView on Hugging Face
0likes15downloads
settings

This repository belongs to SaiyanSai on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namecleaned-asr-transcripts-hinglish
visibilitypublic
licencecc-by-sa-4.0
gatedno
ownerSaiyanSai
Account settings
SaiyanSai/cleaned-asr-transcripts-hinglish · CoolFace