datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
minds14
MInDS-14
MINDS-14 is training and evaluation resource for intent detection task with spoken data. It covers 14
intents extracted from a commercial system in the e-banking domain, associated with spoken examples in 14 diverse language varieties.
Example
MInDS-14 can be downloaded and used as follows:
from datasets import load_dataset
minds_14 = load_dataset("PolyAI/minds14", "fr-FR") # for French
# to download all data for multi-lingual fine-tuning uncomment following… See the full description on the dataset page: https://huggingface.co/datasets/PolyAI/minds14.Taiwanese-Minnan-Sutiau
Taiwanese-Minnan-Sutiau Dataset
The dataset consists of a curated collection of words that resemble tokens in Taiwanese Minnan (Taiwanese Hokkien), aimed at enhancing the recognition and processing of the language for various applications. Sourced from the Ministry of Education in Taiwan, this dataset serves as a valuable linguistic resource for researchers and developers engaged in language processing and recognition tasks.
Dataset Features
Source: Ministry of Education, Taiwan… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Sutiau.moore_audio_data~70h of Mooré paired audio+text data from jw.org using jwsoup
sampling rate = 24 KHz
Headers: audio, text
Taiwanese-Minnan-Example-Sentences
Taiwanese Minnan Example Sentences
The dataset consists of a collection of example sentences designed to aid in recognizing Taiwanese Minnan (Taiwanese Hokkien) for automatic speech recognition (ASR) tasks. This dataset is sourced from the Ministry of Education in Taiwan and aims to provide valuable linguistic resources for researchers and developers working on speech recognition systems.
Dataset Features
Source: Ministry of Education, Taiwan (Sutian Resource Center)
Text:… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Example-Sentences.ScreenASR-Bench
ScreenASR-Bench
Data
Item
Value
Split
test
Cases
2,002
Audio clips
2,002
Keyframes
2,469
Languages
Chinese
Structure
Field
Type
Description
caseid
string
Unique case identifier
ref
string
Reference transcription
target
string
Target text in TN form
level
string
Difficulty level: L1, L2, or L3
audio
audio
Audio clip
keyframes
list[image]
Keyframes associated with the case
frame_captions
list[string]… See the full description on the dataset page: https://huggingface.co/datasets/MingweiFu/ScreenASR-Bench.minds14
MInDS-14
MINDS-14 is training and evaluation resource for intent detection task with spoken data. It covers 14
intents extracted from a commercial system in the e-banking domain, associated with spoken examples in 14 diverse language varieties.
Example
MInDS-14 can be downloaded and used as follows:
from datasets import load_dataset
minds_14 = load_dataset("PolyAI/minds14", "fr-FR") # for French
# to download all data for multi-lingual fine-tuning uncomment… See the full description on the dataset page: https://huggingface.co/datasets/abc-123-456/minds14.mindbridge-phq9-hindi-audio-fixtures
MindBridge Hindi PHQ-9/GAD-7 — Audio Fixtures (30 clips)
Hindi audio fixtures for OIWER (Orthographically-Informed Word Error Rate,
AI4Bharat metric) audio-quality benchmarking on Gemma 4 E2B's native USM
conformer audio encoder. Used to verify post-fine-tune audio quality has
NOT regressed vs base E2B (audio modules explicitly frozen via
requires_grad=False during training to preserve the native USM encoder).
Recording setup
30 clips spanning PHQ-9 Sections A-D… See the full description on the dataset page: https://huggingface.co/datasets/Huzayfah-Patel/mindbridge-phq9-hindi-audio-fixtures.minimind-ptbr-kokoro-tts-68k
MiniMind PT-BR Kokoro TTS 68k
Synthetic Brazilian Portuguese speech dataset generated for the MiniMind/Tucano2 -> Mimi Talker fine-tuning experiments.
Contents
audio/: 68,486 WAV files generated with Kokoro TTS.
manifests/kokoro_groq25k_audio_manifest.jsonl: source text manifest.
manifests/kokoro_groq25k_audio_manifest_with_wavs.jsonl: manifest with WAV paths.
manifests/kokoro_groq25k_audio_manifest_with_mimi.jsonl: manifest with Mimi token references.… See the full description on the dataset page: https://huggingface.co/datasets/marcosremar2/minimind-ptbr-kokoro-tts-68k.english-vocal-medical-terminology-mini
FREE PREVIEW: CLINICAL AI VOICE DATASET — MEDICAL TERMINOLOGY SERIES
Format: LJ Speech Standard Compliance | 24-bit Signed Linear PCM Mono WAV | 48kHz
Thank you for downloading this Developer Compatibility Sample Pack. This repository contains enterprise-grade, high-fidelity, ethically sourced human voice data optimized specifically for training, benchmarking, and stress-testing clinical transcription models, medical speech-to-text (STT) pipelines, and health-tech conversational… See the full description on the dataset page: https://huggingface.co/datasets/MarieDeVox/english-vocal-medical-terminology-mini.mia-meeting
MIA Meeting E2E Dataset
Synthetic meeting dataset for end-to-end experiments:
audio to transcript
transcript plus roster to action items
action item extraction benchmark
Splits
train: 200 samples, 0 with linked audio
validation: 5 samples, 5 with linked audio
eval: 205 samples, 5 with linked audio
Structure
data/*.jsonl # split manifests
audio/<split>/* # linked audio files when available
transcripts/<split>/*.json #… See the full description on the dataset page: https://huggingface.co/datasets/minhthien/mia-meeting.fptu-vovinam-dataset
FPTU Vovinam Dataset
Dataset Description
This dataset contains Vietnamese audio recordings with corresponding text transcriptions, specifically focused on Vovinam martial arts terminology and instruction. This dataset is created by FPT University for research and educational purposes in the field of Vietnamese speech recognition and natural language processing.
Dataset Structure
Audio Files: High-quality audio recordings in AAC format (uploaded as raw data)… See the full description on the dataset page: https://huggingface.co/datasets/minhtien2405/fptu-vovinam-dataset.minds14
MInDS-14
MINDS-14 is training and evaluation resource for intent detection task with spoken data. It covers 14
intents extracted from a commercial system in the e-banking domain, associated with spoken examples in 14 diverse language varieties.
Example
MInDS-14 can be downloaded and used as follows:
from datasets import load_dataset
minds_14 = load_dataset("PolyAI/minds14", "fr-FR") # for French
# to download all data for multi-lingual fine-tuning uncomment… See the full description on the dataset page: https://huggingface.co/datasets/yxl10086/minds14.tunartts-mini-hfurdu-tts-mini
Dataset Card for Urdu-TTS-Mini
A curated Urdu speech dataset for Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) research. Audio segments are extracted from publicly available YouTube speech content, processed through a multi-stage quality pipeline, and annotated with Urdu transcriptions. This is a mini release intended to validate the preprocessing pipeline and establish a quality baseline for future large-scale versions.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/salisai/urdu-tts-mini.minspeech
MinSpeech: Cleaned Multi-dialect Min-nan Dataset (Private)
Important Legal Notice & Copyright Status
This repository is a Private Research Fork of the MinSpeech corpus. It is maintained strictly for individual research purposes, specifically for fine-tuning Automatic Speech Recognition (ASR) and Speech-to-Text Translation (S2TT) models.
1. Ownership & Licensing
Annotations & Metadata: The transcriptions and segment metadata are derived from the MinSpeech… See the full description on the dataset page: https://huggingface.co/datasets/scbz/minspeech.stt-mini-bench
am-pranav/stt-mini-bench
Private, curated mini-benchmark assembled on 2025-09-04.
Note: This dataset mirrors small subsets of upstream corpora (LibriSpeech, TED-LIUM 3, VoxPopuli, Common Voice).
Check each upstream license before sharing. This repo is for internal evaluation only.
Schema
audio : Audio(sampling_rate=16000, decode=False) (files stored in repo)
text : reference transcription
lang : short language code (en, de, fr, es, it, pt)
source: upstream… See the full description on the dataset page: https://huggingface.co/datasets/am-pranav/stt-mini-bench.wuw_min
ygyuan/wuw_min
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards.
The input is a Kaldi-style data directory
(wav.scp, text, utt2spk, utt2dur, segments), where each
utterance is packed as a single tar sample.
Layout
data/
<split>/
metadata.csv
audio/
<split>-000.tar
<split>-001.tar
...
Shard counts:
train: 65 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw audio… See the full description on the dataset page: https://huggingface.co/datasets/ygyuan/wuw_min.IndicContextEvalNOTE: This is a duplicate repo of "https://huggingface.co/datasets/ai4bharat/IndicContextEval" - visit the reference dataset - for any new updates made after Jul 30, 2026.
IndicContextEval
A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages
Code and resources: https://github.com/AI4Bharat/IndicContextEval
Dataset at a glance
Languages
Hindi, Bengali, Telugu, Marathi, Gujarati, Malayalam, Odia, Urdu… See the full description on the dataset page: https://huggingface.co/datasets/Minutor/IndicContextEval.vaani-or-dataset
Vaani [Odia (OR) Dataset] Subset
This is a standalone extraction of the Odia language subset from the Project Vaani dataset.
Gated Access & Licensing
To comply with the original data source policies, this repository requires manual approval. By requesting access, you confirm that you have read and accepted the terms of the original ARTPARK-IISc Vaani dataset.
Original Source: ARTPARK-IISc / Project Vaani
License: CC-BY-NC 4.0 (Non-Commercial)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Minutor/vaani-or-dataset.
