datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nepali-cs-asr
Nepali–English Code-Switched ASR
A ~59-hour corpus of spontaneous Nepali–English code-switched speech clipped from publicly available STEM and CS lecture videos on YouTube. The dataset targets ASR model training and evaluation for code-switched (CS) Nepali–English speech — a variety commonly used in Nepali higher education and online tutoring, where teachers fluidly mix Nepali grammar with English technical vocabulary.
v2 (2026-07) — the current revision. Splits are… See the full description on the dataset page: https://huggingface.co/datasets/saileshbro/nepali-cs-asr.nepali-audio-reserve-r6
Nepali two-speaker conversation chunks
~6680.8 h of Nepali speech at 48 kHz. Two speakers per clip, ~5 minute diarized chunks.
A backup, not a release: the transcripts are machine-generated, and none of this
audio passed the quality gate that produced our training corpus.
Derived from third-party audio whose rights holders did not grant redistribution. The hour count is language-dominant, not monolingual: a chunk labelled Nepali can carry substantial English or Hindi. lang_sec… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r6.nepali_asr
Dataset Card for Nepali Asr Dataset
Dataset Summary
This dataset consists of over 5 hours (300+ minutes) of English speech audio collected from YouTube. The dataset is designed for automatic speech recognition (ASR) and speaker identification tasks. It features both male and female speakers, with approximately 60% of the samples from male voices and the remaining 40% from female voices. The dataset contains 35 distinct speakers, each with their audio segmented into… See the full description on the dataset page: https://huggingface.co/datasets/rishi70612/nepali_asr.nepali_speech_to_text
Nepali Speech-to-Text Dataset
This repository contains a dataset for Automatic Speech Recognition (ASR) in the Nepali language. The dataset is designed for supervised learning tasks and includes audio files along with their corresponding transcriptions. The audio samples have been collected from various open-source platforms and other publicly available sources on the internet.
Each audio file has an average length of 15 seconds and has been converted into a consistent WAV format… See the full description on the dataset page: https://huggingface.co/datasets/pujanpaudel/nepali_speech_to_text.nepali-tts-synthetic-v2
Nepali TTS Synthetic v2
383,298 synthetic Nepali (ne) speech/text pairs, 24 kHz mono 16-bit WAV embedded
as-is (no re-encode, no resampling).
Generated by the synthetic_pipeline in milanakdj/TTS_training: Edge TTS
synthesis → optional voice conversion against a pool of 600 real multi-speaker
reference clips → ASR-based QC gate on character error rate.
Read this before training on it
Only 48% of rows are voice-converted. Each row carries a kept field
recording… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-tts-synthetic-v2.neptel
NepTel v0.1 — Nepali real-telephony ASR benchmark
75 scored segments / 2,375 reference words of real Nepali call-center audio (genuine
two-party customer-support calls), with human-reviewed reference transcripts. To our knowledge
this is the first public Nepali ASR benchmark on real call audio rather than read-aloud speech.
This repository exists so anyone can benchmark a Nepali ASR system without any access
request: the audio is cut and ready, no gate, no approval step.… See the full description on the dataset page: https://huggingface.co/datasets/ampixa/neptel.no-filter-raw-NepaliParliamentDSv2seke-nepali-dataset
🏔️ Seke (SKJ) Language Dataset
Endangered Language Alliance × Internet Archive
Seke (skj) is a critically endangered Sino-Tibetan language spoken by approximately 700 people
in the five villages of Upper Mustang district, Nepal, and in diaspora communities in New York City.
This dataset represents one of the most complete public audio corpora of Seke ever assembled.
Dataset created by Anil Tamang (himalaya-ai).
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/Titung/seke-nepali-dataset.Nepali_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 229,645 hours of processed Nepalese (NP) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Nepali_Call_Center_Audio_Dataset_Dual_Channel.NepFinSpeech
NepFinSpeech-403: A Domain-Specific Nepali Financial Speech Dataset
Overview
NepFinSpeech-403 is a transcribed speech dataset of 403 Nepali financial voice commands, built as part of the SpeakPay research project — a voice-first digital wallet designed for visually impaired individuals in Nepal.
Existing Nepali ASR resources (OpenSLR, Common Voice) cover general-domain speech but contain very few financial utterances. Financial commands have distinct… See the full description on the dataset page: https://huggingface.co/datasets/birajsubedi/NepFinSpeech.nepal-lead-data1
Nepali Speech Dataset (YouTube-sourced)
56 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 56 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepal-lead-data1.Nepali_Call_Center_Audio_Dataset_Dual_Channel
Nepali Call Center Audio Dataset — Dual Channel
Source and Attribution
This repository contains data obtained from the original dataset
published by InfoBay AI Ltd.
Original Dataset
Name:
Nepali_Call_Center_Audio_Dataset_Dual_Channel
Original repository:
https://huggingface.co/datasets/InfoBayAI/Nepali_Call_Center_Audio_Dataset_Dual_Channel
Original creator:
InfoBay AI Ltd.
The original dataset is listed as CC BY 4.0 on its Hugging Face
repository.… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose7777/Nepali_Call_Center_Audio_Dataset_Dual_Channel.nepali-youtube-dataset
Nepali Speech Dataset (YouTube-sourced)
585 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 585 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepali-youtube-dataset.nepali-audio-reserve-r5
Nepali long-form single-speaker readings
~363.2 h of Nepali speech at 48 kHz. One speaker per clip, ~5 minute chunks.
A backup, not a release: the transcripts are machine-generated, and none of this
audio passed the quality gate that produced our training corpus.
Derived from third-party published readings whose rights holders did not grant redistribution. Access is gated and the source is not named.
Columns
column
contents
id
first 16 hex of… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r5.nepali-audio-reserve-r1
Nepali long-form speech, restored
~548.1 h of Nepali speech at 24 kHz. One speaker per clip, average 32 s, up to several minutes.
A backup, not a release: the transcripts are machine-generated, and none of this
audio passed the quality gate that produced our training corpus.
Derived from AI4Bharat IndicVoices-R (CC-BY-4.0) and restored with sidon-v0.1. Attribution is required by that licence, so it is given here. These are the long-form files our quality gate rejected; the 74 h… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r1.nepali-oov-distilled
Nepali OOV-distilled subset (854 h)
An OOV-dense distillation of Premal-12/c9nepali-audio-dataset2 (used with the
author's permission), shipped in four variants: the original single-voice audio,
a CPU-augmented copy, and 244 h re-rendered onto 1,842 real human speakers with
Seed-VC. For Nepali ASR and TTS work.
Filter with the variant field -- see Composition below. If you came here
for speaker diversity, you want variant == "vc".
What this is
The source corpus is… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-oov-distilled.validation_nepali_asr
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/rishi70612/validation_nepali_asr.Nepali-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 229,645 hours of processed Nepali (NP) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Nepali-Call-Center-Audio-Dataset-Single-Channel.nepali-audio-reserve-r2
Nepali synthetic educational dialogue
~137.0 h of Nepali speech at 24 kHz. Two speakers per clip, 2-5 minutes, teacher/student turns.
A backup, not a release: the transcripts are machine-generated, and none of this
audio passed the quality gate that produced our training corpus.
Synthetic. Generated by a TTS model reading NCERT-style educational dialogue, code-mixed Nepali/English. No human speaker is recorded here.
Columns
column
contents
id
first 16… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r2.nepali-audio-reserve-r4
Nepali short clips, ASR-agreed transcripts
~53.9 h of Nepali speech at 24 kHz. Single speaker, 3-20 s.
A backup, not a release: the transcripts are machine-generated, and none of this
audio passed the quality gate that produced our training corpus.
Transcripts kept only where three independent ASR systems agreed. Derived from third-party audio whose rights holders did not grant redistribution.
Columns
column
contents
id
first 16 hex of sha256(original… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r4.openSLR-Nepali
OpenSLR Nepali Speech Dataset (Preprocessed)
Dataset Description
This is a preprocessed version of the Nepali speech dataset from OpenSLR,
ready for training speech models including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS).
Dataset Statistics
Total Audio Files: 118,231
Total Duration: 57.34 hours
Sample Rate: 16kHz
Channels: Mono
Format: WAV
Preprocessing Applied
Text Preprocessing:
Text cleaning and… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/openSLR-Nepali.nepali-audio-reserve-r3
Nepali spontaneous short clips (SNR 40-50)
~16.2 h of Nepali speech at 24 kHz. Single speaker, 5-20 s, spontaneous.
A backup, not a release: the transcripts are machine-generated, and none of this
audio passed the quality gate that produced our training corpus.
Derived from third-party web audio whose rights holders did not grant redistribution, which is why access is gated and the source is not named.
Columns
column
contents
id
first 16 hex of… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r3.nepali_speech_english_translation_shuffle_dataset
Nepali Speech Dataset for Whisper
Nepali audio recordings with English translations.
Dataset Info
Total samples: 1062
Audio format: WAV, 16kHz
Source language: Nepali (ne)
Target language: English (en)
Usage
from datasets import load_dataset
# Load dataset
dataset = load_dataset("lilgoose777/nepali_speech_english_translation_shuffle_dataset")
# Access data
sample = dataset['train'][0]
print(sample['sentence']) # English translation
print(sample['audio'])… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepali_speech_english_translation_shuffle_dataset.nepali-english-speech-data
Nepali Speech Dataset for Whisper
Nepali audio recordings with English translations.
Dataset Info
Total samples: 93
Audio format: WAV, 16kHz
Source language: Nepali (ne)
Target language: English (en)
Usage
from datasets import load_dataset
# Load dataset
dataset = load_dataset("lilgoose777/nepali-english-speech-data")
# Access data
sample = dataset['train'][0]
print(sample['sentence']) # English translation
print(sample['audio']) # Audio data… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepali-english-speech-data.Sagarmatha-ASR-Nepali-Diamond-V3
Dataset Card for Sagarmatha ASR Nepali Diamond V3
Dataset Summary
Sagarmatha ASR Nepali Diamond V3 is a large-scale, production-grade Automatic Speech Recognition (ASR) dataset designed for the Nepali language. The corpus contains 265.7 hours of verified, 16 kHz audio paired with strictly normalized Devanagari transcriptions. It was compiled and curated primarily for the fine-tuning of state-of-the-art multilingual acoustic models, including OpenAI's Whisper… See the full description on the dataset page: https://huggingface.co/datasets/tonibirat/Sagarmatha-ASR-Nepali-Diamond-V3.OpenSLR54-Nepali-ASR-parquet
OpenSLR 54: Large Nepali ASR training data set (unmodified parquet repackaging)
This is an unofficial repackaging of the official OpenSLR 54 release
(SLR54, https://www.openslr.org/54/), converted to parquet so it can be streamed with 🤗 datasets.
It is not affiliated with or endorsed by OpenSLR or the original authors.
All credit for the data belongs to the original creators (see Citation).
What's inside
157,905 utterances, 16 shards: one per original zip… See the full description on the dataset page: https://huggingface.co/datasets/JeevanDai/OpenSLR54-Nepali-ASR-parquet.
