datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MM-ContextASR-Bench
MM-ContextASR Bench
Metadata and evaluation splits for Multimodal Conversational Context for
LLM-Based ASR: Data Construction, Training, and Benchmark.
Dataset summary
Config
Examples
Audio
Context
Primary metric
mm_contextasr
1,250 (250 current utterances × 5 histories)
1,439 WAV files included
Controlled user-assistant dialogue
entity Recall
kespeech
19,212
Source ID only
Same-speaker speech and transcript
CER, SER, entity Recall
cv_yue
3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.voice-dataset-lili
Model card for lili (medium)
Language: sk_SK (Slovak, Slovakia)
Speakers: 1
Quality: medium
Samplerate: 22,050Hz
Dataset
URL: https://github.com/NabuCasa/voice-datasets
License: CC0
tibetan-speech-english-text-dataset
Tibetan Speech Dataset with English Translations
Dataset Description
This dataset contains Tibetan speech recordings paired with transcriptions in Tibetan script and English translations. It is designed to support automatic speech recognition (ASR), machine translation, and text-to-speech (TTS) research for the Tibetan language, which is considered a low-resource language in NLP.
Supported Tasks
Automatic Speech Recognition (ASR): Train models to transcribe… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/tibetan-speech-english-text-dataset.thaha-research-data2-v2
Nepali Speech Dataset (YouTube-sourced)
441 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 441 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data2-v2.nepal-lead-data1
Nepali Speech Dataset (YouTube-sourced)
56 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 56 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepal-lead-data1.thaha-research-data1
Nepali Speech Dataset (YouTube-sourced)
59 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 59 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data1.Nepali_Call_Center_Audio_Dataset_Dual_Channel
Nepali Call Center Audio Dataset — Dual Channel
Source and Attribution
This repository contains data obtained from the original dataset
published by InfoBay AI Ltd.
Original Dataset
Name:
Nepali_Call_Center_Audio_Dataset_Dual_Channel
Original repository:
https://huggingface.co/datasets/InfoBayAI/Nepali_Call_Center_Audio_Dataset_Dual_Channel
Original creator:
InfoBay AI Ltd.
The original dataset is listed as CC BY 4.0 on its Hugging Face
repository.… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose7777/Nepali_Call_Center_Audio_Dataset_Dual_Channel.nepali-youtube-dataset
Nepali Speech Dataset (YouTube-sourced)
585 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 585 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepali-youtube-dataset.kantipur-interview-data3
Nepali Speech Dataset (YouTube-sourced)
83 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 83 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/kantipur-interview-data3.kantipur-interview-data1
Nepali Speech Dataset (YouTube-sourced)
204 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 204 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/kantipur-interview-data1.kantipur-interview-data2
Nepali Speech Dataset (YouTube-sourced)
157 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 157 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/kantipur-interview-data2.kantipur-news-data
Nepali Speech Dataset (YouTube-sourced)
74 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 74 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/kantipur-news-data.thaha-research-data2
Nepali Speech Dataset (YouTube-sourced)
63 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 63 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data2.thaha-research-data2-v2-v2
Nepali Speech Dataset (YouTube-sourced)
76 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 76 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data2-v2-v2.knihi-be-auramcyk_u_paddziamielli_lilija_pilkievic_all
AudioSet Pipeline Output
Мова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
2,324
Працягласць
8 гадз 21 хв
Частата дыскрэтызацыі
22050 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона, ≤30 с)
text — транскрыпцыя… See the full description on the dataset page: https://huggingface.co/datasets/fosters/knihi-be-auramcyk_u_paddziamielli_lilija_pilkievic_all.tibetan-audio-english-6datasets-sample2
Tibetan Audio-English Sentence Dataset (Sample)
This is a sample dataset containing 5 rows from a merged collection of 6 Tibetan audio datasets with English translations.
📊 Dataset Details
Total Samples in Full Dataset: 17,278
Samples in This Preview: 5
Format: Audio + English sentence pairs
Audio Sampling Rate: 16,000 Hz
Languages: Tibetan (audio) → English (text)
🗂️ Source Datasets
This sample is merged from 6 datasets:… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/tibetan-audio-english-6datasets-sample2.nepali_speech_english_translation_shuffle_dataset
Nepali Speech Dataset for Whisper
Nepali audio recordings with English translations.
Dataset Info
Total samples: 1062
Audio format: WAV, 16kHz
Source language: Nepali (ne)
Target language: English (en)
Usage
from datasets import load_dataset
# Load dataset
dataset = load_dataset("lilgoose777/nepali_speech_english_translation_shuffle_dataset")
# Access data
sample = dataset['train'][0]
print(sample['sentence']) # English translation
print(sample['audio'])… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepali_speech_english_translation_shuffle_dataset.nepali-english-speech-data
Nepali Speech Dataset for Whisper
Nepali audio recordings with English translations.
Dataset Info
Total samples: 93
Audio format: WAV, 16kHz
Source language: Nepali (ne)
Target language: English (en)
Usage
from datasets import load_dataset
# Load dataset
dataset = load_dataset("lilgoose777/nepali-english-speech-data")
# Access data
sample = dataset['train'][0]
print(sample['sentence']) # English translation
print(sample['audio']) # Audio data… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepali-english-speech-data.knihi-be-auramcyk_u_paddziamielli_lilija_pilkievic_output_original
AudioSet Pipeline Output — арыгінальнае аўдыё
Мова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
knihi-be-auramcyk_u_paddziamielli_lilija_pilkievic_output
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text — транскрыпцыя
chunk_uid — унікальны ідэнтыфікатар… See the full description on the dataset page: https://huggingface.co/datasets/fosters/knihi-be-auramcyk_u_paddziamielli_lilija_pilkievic_output_original.tibetan-audio-english-6datasets-sample
Tibetan Audio-English Sentence Dataset (Sample)
This is a sample dataset containing 5 rows from a merged collection of 6 Tibetan audio datasets with English translations.
📊 Dataset Details
Total Samples in Full Dataset: 17,278
Samples in This Preview: 5
Format: Audio + English sentence pairs
Audio Sampling Rate: 16,000 Hz
Languages: Tibetan (audio) → English (text)
🗂️ Source Datasets
This sample is merged from 6 datasets:… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/tibetan-audio-english-6datasets-sample.
