datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nepali-cs-asr
Nepali–English Code-Switched ASR
A ~59-hour corpus of spontaneous Nepali–English code-switched speech clipped from publicly available STEM and CS lecture videos on YouTube. The dataset targets ASR model training and evaluation for code-switched (CS) Nepali–English speech — a variety commonly used in Nepali higher education and online tutoring, where teachers fluidly mix Nepali grammar with English technical vocabulary.
v2 (2026-07) — the current revision. Splits are… See the full description on the dataset page: https://huggingface.co/datasets/saileshbro/nepali-cs-asr.lmaanaDataset
LmaanaDataset
Cleaned transcription data derived from the transcribed 100-gt-2.5 subset of atlasia/MoulSot-Full.
Contents
81,603 non-empty transcription rows.
original_text: source transcription.
new_transcription: cleaned transcription with selected French loanwords restored to Latin script.
training_text: recommended text field for training.
review_status: indicates whether a candidate conversion was applied.
audio/100-gt-2.5/: original audio-bearing Parquet… See the full description on the dataset page: https://huggingface.co/datasets/sailu4/lmaanaDataset.two-minute-papers
Dataset Card for "two-minute-papers"
More Information needed
cleaned-asr-transcripts-hinglish
cleaned-asr-transcripts-hinglish
bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts.
This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.
