datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
danish-asr-unified
Danish ASR Unified Dataset
Unified Danish speech recognition dataset combining 7 sources (~3.5M samples, ~16k hours):
Source
Samples
Description
VoxPopuli
1,775,578
European Parliament recordings
ftspeech
995,677
Danish Parliament (Folketinget)
CoRal-v3 read_aloud
299,255
Read-aloud Danish speech
nst-da
182,605
NST Danish speech
CoRal-v3 conversation
147,249
Conversational Danish speech
nota
98,600
Danish broadcast media
Common Voice 17
3,484
Crowd-sourced… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified.majestrino-unified-detailed-captions
Majestrino Unified Detailed Captions
Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption.
Stats
4,658,407 samples
932 tar files (~1.1 GB each)
~1,017 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption
caption_type — always unified_detailed_caption
transcription — speech transcription (when available, normalized from multiple source keys)
duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.farsi-asr-unified-cleaned
🎧 Farsi ASR Unified Dataset (Parquet Sharded Edition)
Overview
The Farsi ASR Unified Dataset is a large-scale, high-quality, and fully standardized collection of Persian (Farsi) speech-to-text data — designed specifically for modern machine learning and ASR (Automatic Speech Recognition) workflows.
This dataset consolidates audio–text pairs from multiple open sources, applies a rigorous cleaning and normalization pipeline, and stores everything efficiently in Parquet… See the full description on the dataset page: https://huggingface.co/datasets/kiarashQ/farsi-asr-unified-cleaned.unified-kannada-asr-1.0
Dataset Card for "unified-kannada-asr-1.0"
More Information needed
stt-unified-bench
am-pranav/stt-unified-bench
Private, language/locale-partitioned mini-benchmark for STT models.
Each subset is a dataset config (e.g., en, de, en_indian_accent, hi_in) with a single split val.
Audio is staged at 16 kHz and stored in-repo for reproducibility.
Schema
audio : Audio(sampling_rate=16000, decode=False)
text : reference transcription
lang : implied by dataset config name
source : upstream dataset tag
id : source-stable id
⚠️ For internal evaluation only.… See the full description on the dataset page: https://huggingface.co/datasets/am-pranav/stt-unified-bench.FLARE-1k-Unified-T2VAadvwave-unified-prompt-5model
AdvWave — Unified-Prompt 5-Model Adversarial Audio
Adversarial speech clips that carry the AdvWave waveform-suffix attack against five speech
language models, all crafted and evaluated under a single, identical prompt so the only
variable across models is the model itself (a controlled cross-model comparison — the "(C)" setting).
⚠️ Safety / intended use. These clips are optimized to make speech LLMs comply with harmful
AdvBench requests. They are released for defensive… See the full description on the dataset page: https://huggingface.co/datasets/Snooow1029/advwave-unified-prompt-5model.majestrino-unified-detailed-captions-temporal
Majestrino Unified Detailed Captions with Temporal Aspects
Filtered subset of laion/majestrino-data containing only samples with unified_detailed_caption_with_temporal_aspects.
Stats
4,128,665 samples
826 tar files (~1.1 GB each)
~878 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption with temporal aspects
caption_type — always unified_detailed_caption_with_temporal_aspects
transcription — speech… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions-temporal.unified-hausa-speech
Unified Hausa Speech Dataset v5
Dataset Description
A large-scale, cleaned, deduplicated, and quality-filtered Hausa speech dataset compiled from 6 open-source collections. Designed for Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) research on one of Africa's most widely spoken languages.
Hausa (ISO 639-1: ha) is a Chadic language spoken by over 80 million people across West and Central Africa — primarily in Nigeria and Niger, and as a trade language… See the full description on the dataset page: https://huggingface.co/datasets/suleiman2003/unified-hausa-speech.urdu-asr-unified-v1FLARE-1k-Unified-T2VAUnified_v1Noises_Unified_v2Noises_Unified_v3common-voice-unified-splitsNoises_Unified_v4Noises_Unified_v1eval-whisper-small-pilotgpt-unified-all-data-lowercase-new-rewritten-20260223-2259
Training Evaluation: whisper-small-pilotgpt-unified-all-data-lowercase-new-rewritten
Evaluation results comparing base model vs fine-tuned model.
Summary
Model
WER
openai/whisper-small (base)
46.81%
Trelis/whisper-small-pilotgpt-unified-all-data-lowercase-new-rewritten (fine-tuned)
34.22%
Improvement: 12.59% WER reduction (lower is better)
Source Data
Evaluation Dataset: Trelis/pilotgpt-test-0.5s-rewritten
Base Model: openai/whisper-small… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-whisper-small-pilotgpt-unified-all-data-lowercase-new-rewritten-20260223-2259.pilotgpt-unified-all-data-lowercase-new-rewrittenpilotgpt-unified-all-data-lowercase-data-prepeval-whisper-small-pilotgpt-unified-all-data-lowercase-data-prep-6772-20260219-1448
Training Evaluation: whisper-small-pilotgpt-unified-all-data-lowercase-data-prep-6772
Evaluation results comparing base model vs fine-tuned model.
Summary
Model
WER
openai/whisper-small (base)
53.69%
Trelis/whisper-small-pilotgpt-unified-all-data-lowercase-data-prep-6772 (fine-tuned)
27.54%
Improvement: 26.15% WER reduction (lower is better)
Source Data
Evaluation Dataset: Trelis/pilotgpt-test-0.5s
Base Model: openai/whisper-small… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-whisper-small-pilotgpt-unified-all-data-lowercase-data-prep-6772-20260219-1448.eval-whisper-small-pilotgpt-unified-all-raw-nopack-0.5s-clean-7283-20260227-2013
Training Evaluation: whisper-small-pilotgpt-unified-all-raw-nopack-0.5s-clean-7283
Evaluation results comparing base model vs fine-tuned model.
Summary
Model
WER
openai/whisper-small (base)
53.69%
Trelis/whisper-small-pilotgpt-unified-all-raw-nopack-0.5s-clean-7283 (fine-tuned)
32.92%
Improvement: 20.77% WER reduction (lower is better)
Source Data
Evaluation Dataset: Trelis/pilotgpt-test-0.5s
Base Model: openai/whisper-small
Fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-whisper-small-pilotgpt-unified-all-raw-nopack-0.5s-clean-7283-20260227-2013.
