datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lahgtna-v3-small
Lahgtna — Dialect-Balanced Arabic ASR (v3 small)
A dialect-balanced multi-dialect Arabic speech-recognition corpus: 54,600 clips /
267.3 hours across 13 Arabic dialects, 16 kHz mono. Each dialect is evenly
represented — 4,000 train + 200 test clips per dialect — so models and
evaluations aren't skewed toward high-resource dialects (e.g. Egyptian/Gulf).
Used to train the oddadmix v2 dialectal-ASR model family.
Duration by dialect
Dialect
Train (h)
Train clips… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/lahgtna-v3-small.vistaar_small_asr_eval
Vistaar Small ASR Eval
Dataset Description
The Vistaar Small ASR Eval dataset is a multilingual automatic speech recognition evaluation dataset containing 9,486 audio samples across 12 Indian languages. This dataset represents a subset of the larger Vistaar dataset published by AI4Bharat, designed specifically for evaluating ASR model performance on diverse Indian language speech data. A smaller evaluation dataset was created for the use-cases where a quick benchmarking… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/vistaar_small_asr_eval.small-overlapping-speech-bench
Small Overlapping Speech Bench
A tiny, fully-reproducible benchmark for multilingual overlapping speech. Each of the
100 clips contains three people speaking at the same time, each in a different European
language, with ground-truth per-speaker timestamps, languages, and transcripts.
It is a deliberately hard "cocktail-party" stress test: how much of each simultaneous speaker can
an ASR (speech-to-text) model recover, and can a model tell how many people are talking?
100 clips… See the full description on the dataset page: https://huggingface.co/datasets/laion/small-overlapping-speech-bench.1000h-us-english-smartphone-conversation
📚 1000 Hours of Conversational American English Speech Dataset (Smartphone Recordings)
This dataset contains sample conversational speech data collected by Appen. The audio was recorded naturally using smartphones and is suitable for:
Automatic Speech Recognition (ASR)
Speaker Identification and Gender/Age Analysis
Dialect and Accent Modeling
Multi-speaker Speech Separation
🧾 Dataset Contents
The dataset includes:
metadata.CSV: Metadata including speaker gender, age… See the full description on the dataset page: https://huggingface.co/datasets/Appenlimited/1000h-us-english-smartphone-conversation.Small-STT-Eval-Audio-Dataset
Small STT Eval Audio Dataset
A small speech-to-text evaluation dataset containing 92 audio samples with ground truth transcriptions. Designed for evaluating STT systems on technical vocabulary, code-switching (English/Hebrew), and various speaking styles.
Dataset Description
This dataset contains audio recordings with accompanying transcriptions across multiple categories:
Category
Count
Description
tech_github
5
GitHub-related technical vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Small-STT-Eval-Audio-Dataset.processed-smarthome-th
processed-smarthome-th
Cleaned Thai speech dataset for smart-home commands: 9,600 utterances (7,680 train / 960 dev / 960 test) with transcripts.
Format
Field
Description
sentence
Transcript in Thai
audio
Audio clip
Usage
from datasets import load_dataset
ds = load_dataset("Porameht/processed-smarthome-th")
Used to fine-tune Porameht/whisper-tiny-smarthome-thai (WER 24.375 on the eval split).
vistaar_small_asr_eval
Vistaar Small ASR Eval
Dataset Description
The Vistaar Small ASR Eval dataset is a multilingual automatic speech recognition evaluation dataset containing 9,486 audio samples across 12 Indian languages. This dataset represents a subset of the larger Vistaar dataset published by AI4Bharat, designed specifically for evaluating ASR model performance on diverse Indian language speech data. A smaller evaluation dataset was created for the use-cases where a quick benchmarking… See the full description on the dataset page: https://huggingface.co/datasets/AdityK2409/vistaar_small_asr_eval.vistaar_small_asr_eval_2
Vistaar Small ASR Eval
Dataset Description
The Vistaar Small ASR Eval dataset is a multilingual automatic speech recognition evaluation dataset containing 9,486 audio samples across 12 Indian languages. This dataset represents a subset of the larger Vistaar dataset published by AI4Bharat, designed specifically for evaluating ASR model performance on diverse Indian language speech data. A smaller evaluation dataset was created for the use-cases where a quick benchmarking… See the full description on the dataset page: https://huggingface.co/datasets/AdityK2409/vistaar_small_asr_eval_2.meddies-asr-small-test
Meddies ASR Small Test
Small Vietnamese golden-test slice for ASR evaluation, packaged as Parquet with embedded audio in
a datasets.Audio column.
Configs
vi_faptv_human_gold: 20 FAPTV examples sourced from Meddies/meddies-asr-human-labels, clipped to
the first 10 minutes of each original human-labeled episode.
Subset Status
vi_faptv_human_gold is directly scorable.
text and reference_text are identical and come from the clipped VTT transcript… See the full description on the dataset page: https://huggingface.co/datasets/Meddies/meddies-asr-small-test.TORGO-very-smallbengali-tts-folderized-parquet-stage1
Bengali TTS Folderized Parquet Stage 1
This is the intermediate organized parquet layer before the final fully row-wise Bengali TTS dataset.
Generated metadata refresh: 2026-06-18T21:44:49Z
Layout
<speaker>/part-00000.parquet
<speaker>/part-00001.parquet
<speaker>/metadata/stats.json
Columns
speaker
video_id
chunk_file
audio_file
duration
transcription
uuid
audio as Hugging Face Audio feature backed by parquet struct<bytes,path>… See the full description on the dataset page: https://huggingface.co/datasets/smam/bengali-tts-folderized-parquet-stage1.Punjabi-STT-Small-8HThis dataset has been partitioned from alvynabranches/punjabi_stt_v1 dataset of 180 hours.
Huge thanks to him for providing this dataset publicly.
