CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01saileshbro /nepali-cs-asr Nepali–English Code-Switched ASR A ~59-hour corpus of spontaneous Nepali–English code-switched speech clipped from publicly available STEM and CS lecture videos on YouTube. The dataset targets ASR model training and evaluation for code-switched (CS) Nepali–English speech — a variety commonly used in Nepali higher education and online tutoring, where teachers fluidly mix Nepali grammar with English technical vocabulary. v2 (2026-07) — the current revision. Splits are… See the full description on the dataset page: https://huggingface.co/datasets/saileshbro/nepali-cs-asr.audioautomatic-speech-recognition10K<n<100K1 likes852 downloads2mo agoHugging Face02milanakdj /nepali-audio-reserve-r6gated Nepali two-speaker conversation chunks ~6680.8 h of Nepali speech at 48 kHz. Two speakers per clip, ~5 minute diarized chunks. A backup, not a release: the transcripts are machine-generated, and none of this audio passed the quality gate that produced our training corpus. Derived from third-party audio whose rights holders did not grant redistribution. The hour count is language-dominant, not monolingual: a chunk labelled Nepali can carry substantial English or Hindi. lang_sec… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r6.audioautomatic-speech-recognition10K<n<100K0 likes339 downloads14d agoHugging Face03milanakdj /nepali-tts-synthetic-v2gated Nepali TTS Synthetic v2 383,298 synthetic Nepali (ne) speech/text pairs, 24 kHz mono 16-bit WAV embedded as-is (no re-encode, no resampling). Generated by the synthetic_pipeline in milanakdj/TTS_training: Edge TTS synthesis → optional voice conversion against a pool of 600 real multi-speaker reference clips → ASR-based QC gate on character error rate. Read this before training on it Only 48% of rows are voice-converted. Each row carries a kept field recording… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-tts-synthetic-v2.audiotext-to-speech100K<n<1M0 likes117 downloads23d agoHugging Face04pujanpaudel /nepali_speech_to_text Nepali Speech-to-Text Dataset This repository contains a dataset for Automatic Speech Recognition (ASR) in the Nepali language. The dataset is designed for supervised learning tasks and includes audio files along with their corresponding transcriptions. The audio samples have been collected from various open-source platforms and other publicly available sources on the internet. Each audio file has an average length of 15 seconds and has been converted into a consistent WAV format… See the full description on the dataset page: https://huggingface.co/datasets/pujanpaudel/nepali_speech_to_text.audioautomatic-speech-recognition1K<n<10K1 likes87 downloads2y agoHugging Face05kiranpantha /no-filter-raw-NepaliParliamentDSv2audioautomatic-speech-recognition10K<n<100K0 likes61 downloads1y agoHugging Face06Titung /seke-nepali-dataset 🏔️ Seke (SKJ) Language Dataset Endangered Language Alliance × Internet Archive Seke (skj) is a critically endangered Sino-Tibetan language spoken by approximately 700 people in the five villages of Upper Mustang district, Nepal, and in diaspora communities in New York City. This dataset represents one of the most complete public audio corpora of Seke ever assembled. Dataset created by Anil Tamang (himalaya-ai). 📊 Dataset Statistics Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/Titung/seke-nepali-dataset.audioautomatic-speech-recognitionn<1K1 likes46 downloads5mo agoHugging Face07JeevanDai /OpenSLR54-Nepali-ASR-parquet OpenSLR 54: Large Nepali ASR training data set (unmodified parquet repackaging) This is an unofficial repackaging of the official OpenSLR 54 release (SLR54, https://www.openslr.org/54/), converted to parquet so it can be streamed with 🤗 datasets. It is not affiliated with or endorsed by OpenSLR or the original authors. All credit for the data belongs to the original creators (see Citation). What's inside 157,905 utterances, 16 shards: one per original zip… See the full description on the dataset page: https://huggingface.co/datasets/JeevanDai/OpenSLR54-Nepali-ASR-parquet.audioautomatic-speech-recognition100K<n<1M0 likes37 downloads1d agoHugging Face08lilgoose777 /nepal-lead-data1gated Nepali Speech Dataset (YouTube-sourced) 56 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split. Splits train: 56 segments validation: 0 segments test: 0 segments Transcript columns — read this before training Each segment carries three transcript variants. They are NOT interchangeable: text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepal-lead-data1.audioautomatic-speech-recognitionn<1K0 likes35 downloads28d agoHugging Face09sumanpaudel1997 /nepali-asr-benchmark Nepali ASR Benchmark Per-utterance reference, hypothesis, WER, and CER for the six released Nepali ASR checkpoints evaluated on three independent test sets. Released alongside the paper Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition. Contents Field Type Description utterance_id string stable identifier {test_set}-{index} reference string NFC-normalised gold transcription (Devanagari) hypothesis string… See the full description on the dataset page: https://huggingface.co/datasets/sumanpaudel1997/nepali-asr-benchmark.tabularautomatic-speech-recognition10K<n<100K0 likes31 downloads4mo agoHugging Face10lilgoose777 /nepali-youtube-datasetgated Nepali Speech Dataset (YouTube-sourced) 585 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split. Splits train: 585 segments validation: 0 segments test: 0 segments Transcript columns — read this before training Each segment carries three transcript variants. They are NOT interchangeable: text_original — the YouTube caption text (if any) that overlapped this… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepali-youtube-dataset.audioautomatic-speech-recognitionn<1K0 likes31 downloads15d agoHugging Face11milanakdj /nepali-audio-reserve-r5gated Nepali long-form single-speaker readings ~363.2 h of Nepali speech at 48 kHz. One speaker per clip, ~5 minute chunks. A backup, not a release: the transcripts are machine-generated, and none of this audio passed the quality gate that produced our training corpus. Derived from third-party published readings whose rights holders did not grant redistribution. Access is gated and the source is not named. Columns column contents id first 16 hex of… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r5.audioautomatic-speech-recognition1K<n<10K0 likes31 downloads14d agoHugging Face12milanakdj /nepali-audio-reserve-r1gated Nepali long-form speech, restored ~548.1 h of Nepali speech at 24 kHz. One speaker per clip, average 32 s, up to several minutes. A backup, not a release: the transcripts are machine-generated, and none of this audio passed the quality gate that produced our training corpus. Derived from AI4Bharat IndicVoices-R (CC-BY-4.0) and restored with sidon-v0.1. Attribution is required by that licence, so it is given here. These are the long-form files our quality gate rejected; the 74 h… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r1.audioautomatic-speech-recognition10K<n<100K0 likes31 downloads14d agoHugging Face13milanakdj /nepali-oov-distilledgated Nepali OOV-distilled subset (854 h) An OOV-dense distillation of Premal-12/c9nepali-audio-dataset2 (used with the author's permission), shipped in four variants: the original single-voice audio, a CPU-augmented copy, and 244 h re-rendered onto 1,842 real human speakers with Seed-VC. For Nepali ASR and TTS work. Filter with the variant field -- see Composition below. If you came here for speaker diversity, you want variant == "vc". What this is The source corpus is… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-oov-distilled.audioautomatic-speech-recognition10K<n<100K1 likes30 downloads12d agoHugging Face14milanakdj /nepali-audio-reserve-r2gated Nepali synthetic educational dialogue ~137.0 h of Nepali speech at 24 kHz. Two speakers per clip, 2-5 minutes, teacher/student turns. A backup, not a release: the transcripts are machine-generated, and none of this audio passed the quality gate that produced our training corpus. Synthetic. Generated by a TTS model reading NCERT-style educational dialogue, code-mixed Nepali/English. No human speaker is recorded here. Columns column contents id first 16… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r2.audioautomatic-speech-recognition1K<n<10K0 likes19 downloads14d agoHugging Face15AkAiNp /nepal-oral-demo Nepal Oral Demo Public, always-safe fixtures for the Nepal oral-language backbone (AkAiNp). Fictional “Demo Himalayan” track only Schema examples for CI, export-script tests, and the Expo training app offline demo pack No real community speakers, ever Monorepo: nepal-multilingual-llm (local project). Source: data/packs/_demo/ + packages/schema/examples/. Intended uses OK Not OK Unit tests, pipeline dry-runs Training production ASR/TTS as if it were… See the full description on the dataset page: https://huggingface.co/datasets/AkAiNp/nepal-oral-demo.textaudio-classificationn<1K0 likes18 downloads2mo agoHugging Face16milanakdj /nepali-audio-reserve-r4gated Nepali short clips, ASR-agreed transcripts ~53.9 h of Nepali speech at 24 kHz. Single speaker, 3-20 s. A backup, not a release: the transcripts are machine-generated, and none of this audio passed the quality gate that produced our training corpus. Transcripts kept only where three independent ASR systems agreed. Derived from third-party audio whose rights holders did not grant redistribution. Columns column contents id first 16 hex of sha256(original… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r4.audioautomatic-speech-recognition10K<n<100K0 likes18 downloads14d agoHugging Face17milanakdj /nepali-audio-reserve-r3gated Nepali spontaneous short clips (SNR 40-50) ~16.2 h of Nepali speech at 24 kHz. Single speaker, 5-20 s, spontaneous. A backup, not a release: the transcripts are machine-generated, and none of this audio passed the quality gate that produced our training corpus. Derived from third-party web audio whose rights holders did not grant redistribution, which is why access is gated and the source is not named. Columns column contents id first 16 hex of… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r3.audioautomatic-speech-recognition1K<n<10K0 likes15 downloads15d agoHugging Face18lilgoose777 /nepali_speech_english_translation_shuffle_dataset Nepali Speech Dataset for Whisper Nepali audio recordings with English translations. Dataset Info Total samples: 1062 Audio format: WAV, 16kHz Source language: Nepali (ne) Target language: English (en) Usage from datasets import load_dataset # Load dataset dataset = load_dataset("lilgoose777/nepali_speech_english_translation_shuffle_dataset") # Access data sample = dataset['train'][0] print(sample['sentence']) # English translation print(sample['audio'])… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepali_speech_english_translation_shuffle_dataset.audioautomatic-speech-recognition1K<n<10K0 likes11 downloads8mo agoHugging Face19lilgoose777 /nepali-english-speech-data Nepali Speech Dataset for Whisper Nepali audio recordings with English translations. Dataset Info Total samples: 93 Audio format: WAV, 16kHz Source language: Nepali (ne) Target language: English (en) Usage from datasets import load_dataset # Load dataset dataset = load_dataset("lilgoose777/nepali-english-speech-data") # Access data sample = dataset['train'][0] print(sample['sentence']) # English translation print(sample['audio']) # Audio data… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepali-english-speech-data.audioautomatic-speech-recognitionn<1K0 likes7 downloads8mo agoHugging Face20tonibirat /Sagarmatha-ASR-Nepali-Diamond-V3gated Dataset Card for Sagarmatha ASR Nepali Diamond V3 Dataset Summary Sagarmatha ASR Nepali Diamond V3 is a large-scale, production-grade Automatic Speech Recognition (ASR) dataset designed for the Nepali language. The corpus contains 265.7 hours of verified, 16 kHz audio paired with strictly normalized Devanagari transcriptions. It was compiled and curated primarily for the fine-tuning of state-of-the-art multilingual acoustic models, including OpenAI's Whisper… See the full description on the dataset page: https://huggingface.co/datasets/tonibirat/Sagarmatha-ASR-Nepali-Diamond-V3.audioautomatic-speech-recognition100K<n<1M0 likes6 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.