datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons.
License identifiers are normalized to cc-zero, cc-by-4.0,
cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self.
This provides a richer alternative to Common Voice.
Characteristics of the dataset:
One or multiple speakers
Different accents
Different domain texts
761 audio files
We found this dataset useful for audio tasks such as:
Language detection
Evaluation of STT systems
New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.macro_prosody_sample_set
Alexandria Voice Corpus — Multilingual Macro-Prosody Telemetry
Version 1.1 — Replacement release
This pack supersedes the earlier Korean & Hindi two-language release. That release was built on a pipeline with several unresolved quality-gate bugs (documented below). This version corrects all known issues and expands to seven typologically diverse languages.
No audio is included. This is a structured acoustic feature dataset for linguistic research, speech technology, and… See the full description on the dataset page: https://huggingface.co/datasets/moonscape-software/macro_prosody_sample_set.mush_hyMush dataset
An audio–transcription alignment dataset for the Mush dialect of Armenian.
It contains approximately 4.5 hours of speech distributed across three splits:
Train: 4,830 samples
Validation (Dev): 117 samples
Test: 650 samples
Content
Each example includes:
audio: a WAV audio file
transcription: the Armenian transcription text
duration: audio duration in seconds
mwa_hyModern Western Armenian dataset
An audio–transcription alignment dataset for the Modern Western Armenian dialect of Armenian.
It contains approximately 42 hours of speech distributed across three splits:
Train: 13,271 samples / 41 hr.
Validation (Dev): 120 samples / 0.3 hr.
Test: 178 samples / 0.5 hr.
Content
Each example includes:
audio: a WAV audio file
transcription: the Armenian transcription text
duration: audio duration in seconds
Sources
The recordings were sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/Center-of-Advanced-Software-Technologies/mwa_hy.MLAAD_Audit
MLAAD — SSA Acoustic Feature Audit
Moonscape Software | Synthetic Speech Atlas
Research audit contribution to the MLAAD dataset team
Overview
This repository contains acoustic feature measurements extracted from the
MLAAD (Multilingual Audio Anti-Spoofing Dataset) corpus by the Moonscape
Synthetic Speech Atlas (SSA) pipeline.
298,000 rows. 152 columns. One row per MLAAD clip.
No audio files are included. Each row contains classical signal processing
and biomechanical… See the full description on the dataset page: https://huggingface.co/datasets/moonscape-software/MLAAD_Audit.Telecom_Channel_Degredation_Matrix
SSA Codec Degradation Study — Acoustic Feature Exports
Moonscape Software | 2026
A companion to the Synthetic Speech Atlas (SSA)
Overview
This dataset quantifies the effect of 35 codec conditions on 80+ acoustic
features extracted from 7,500 biological speech clips. It answers the question:
"Which acoustic features survive telecommunications codec compression, and which
are destroyed?"
The corpus is the empirical foundation for channel-aware gate calibration in
deepfake… See the full description on the dataset page: https://huggingface.co/datasets/moonscape-software/Telecom_Channel_Degredation_Matrix.lori_hyLori dataset
An audio–transcription alignment dataset for the Lori dialect of Armenian.
It contains approximately 4.5 hours of speech distributed across three splits:
Train: 4,340 samples
Validation (Dev): 90 samples
Test: 580 samples
Content
Each example includes:
audio: a WAV audio file
transcription: the Armenian transcription text
duration: audio duration in seconds
softone
