datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voxbox
VoxBox
This dataset is a curated collection of bilingual speech corpora annotated clean transcriptions and rich metadata incluing age, gender, and emotion.
Dataset Structure
.
├── audios/
│ └── aishell-3/ # Audio files (organised by sub-corpus)
│ └── ...
└── metadata/
├── aishell-3.jsonl
├── casia.jsonl
├── commonvoice_cn.jsonl
├── ...
└── wenetspeech4tts.jsonl # JSONL metadata files
Each JSONL file corresponds to a… See the full description on the dataset page: https://huggingface.co/datasets/SparkAudio/voxbox.CosyVoice2-SparkTTStokenise_spanish_datasetDataset tokenizado para TTS en español. Incluye audio, transcripción, emoción, y códigos SNAC.
spanish_AudiosJ-SPAW_LA
J-SPAW (LA track, eval)
⚠️ NON-COMMERCIAL USE ONLY
The upstream J-SPAW dataset is released "For non-commercial use only"
(see the J-SPAW repository). This
packaging inherits that restriction: do not use it for any commercial
purpose. It is provided solely for non-commercial academic research and
benchmarking. The upstream terms are sparse and do not spell out
redistribution; contact the original authors for any use beyond
non-commercial research.
Benchmark-ready… See the full description on the dataset page: https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/J-SPAW_LA.SpatialAudio
SpatialAudio
This repo hosts the dataset and models of "BAT: Learning to Reason about Spatial Sounds with Large Language Models" [ICML 2024 bib].
Spatial Audio Dataset (Mono/Binaural/Ambisonics)
AudioSet (Anechoic Audio Source)
We provide Balanced train and Evaluation set for your convenience. You can download from SpatialAudio.
For the Unbalanced train set, please refer to Official AudioSet.
Metadata can be downloaded from metadata.
AudioSet
├──… See the full description on the dataset page: https://huggingface.co/datasets/zhisheng01/SpatialAudio.spanish-slang-stt-data
Spanish Regional Speech-to-Text Dataset
A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models.
Dataset Description
This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions:
Region
Samples
Description
Mexico
17,725
Mexican Spanish including CIEMPIESS corpus
Spain
11,360
Castilian Spanish from TEDx and Common Voice
Argentina
5,839
Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.kupe-spark-asr-270m-data
kupe-spark-asr-270m — data
Multilingual ASR corpus for kupe-spark-asr-270m (Gemma-3-270m + Mimi codec).
Languages: en (English), hi (Hindi), gu (Gujarati), bn (Bengali), ur (Urdu), mr (Marathi)
Configs
audio — raw speech resampled to 24 kHz mono (audio/data/shard_*.parquet).
mimi — Mimi codebook-0 tokens (12.5 tok/s) + transcripts (mimi/*.parquet). Used for training.
Shards are uploaded one-by-one as they are fetched. Resume state lives in… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/kupe-spark-asr-270m-data.SPC
Dataset Card for "SPC-v2"
More Information needed
cml_tts_dataset_spanishspassLicense: Creative Commons Attribution 4.0 International
Source
Rhoddy Viveros-Muñoz, Pablo Huijse, Victor Poblete, Victor Vargas, Diego Espejo, Matthieu Vernier, Diego Vergara, Jorge Arenas, & Enrique Suárez. (2022).
SPASS dataset: A synthetic polyphonic dataset with spatiotemporal labels of sound sources (1.0). Zenodo.
https://doi.org/10.5281/zenodo.7484370
benchmark_DEMAND_noise
benchmark_DEMAND_noise
This dataset is a segmented subset derived from DEMAND: Diverse Environments Multichannel Acoustic Noise Database.
It is prepared for the SPARCO noise ablation benchmark. The intended use is to provide fixed 4-second environmental noise segments for:
AUROC-based SAE noise-related feature selection
binary noise-presence scorer training
scorer threshold calibration
final held-out benchmark evaluation
Source
Original source:
DEMAND: Diverse… See the full description on the dataset page: https://huggingface.co/datasets/SPARCO-project/benchmark_DEMAND_noise.voxforge_spanish_enhanced
VoxForge Spanish Enhanced (CleanUNet + FlashSR)
Dataset Summary
This dataset is a processed and enhanced version of the Spanish subset of:
VoxForge.
Furthermore, as this is a personal project, we give no guarantees that the audio is completely clean from any artifacts or noise the CleanUNet model could not remove.
However, we have personally tested the corpus via the fine-tuning of some SOTA speech models and the results have been satisfactory.
It has been created to… See the full description on the dataset page: https://huggingface.co/datasets/ebellob/voxforge_spanish_enhanced.evaluate_spatiallatam-spanish-speech-orpheus-tts-24khz
LATAM Spanish High-Quality Speech Dataset (24kHz - Orpheus TTS Ready)
Dataset Description
This dataset contains approximately 24 hours of high-quality speech audio in Latin American Spanish, specifically prepared for Text-to-Speech (TTS) applications like OrpheusTTS, which require a 24kHz sampling rate.
The audio files are derived from the Crowdsourced high-quality speech datasets made by Google and were obtained via OpenSLR. The original recordings were high-quality… See the full description on the dataset page: https://huggingface.co/datasets/GianDiego/latam-spanish-speech-orpheus-tts-24khz.google-chilean-spanish
Dataset Card for Tamil Speech
Dataset Summary
This dataset consists of 7 hours of transcribed high-quality audio of Chilean Spanish sentences recorded by 31 volunteers. The dataset is intended for speech technologies.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
Supported Tasks
text-to-speech, text-to-audio: The dataset can be used to train a model for Text-To-Speech (TTS).
automatic-speech-recognition… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/google-chilean-spanish.dnr-v3-spagoogle-colombian-spanish
Dataset Card for "google-colombian-spanish"
More Information needed
spanish_voxpopuli_alignedlibrivox_spanish
Dataset Card for librivox_spanish
Dataset Summary
Librivox is a non-commercial, non-profit and ad-free project that is dedicated to make all books in the public domain available, for free, in audio format on the internet. According to this, we downloaded 300 titles in Spanish to create the LIBRIVOX SPANISH CORPUS.
The LIBRIVOX SPANISH CORPUS has a duration of 73 hours and it is constituted by audio files between 3 and 10 seconds long, manually segmented. Transcription are… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/librivox_spanish.google-argentinian-spanish
Dataset Card for "google-argentinian-spanish"
More Information needed
spansynth-edit-gallery
SpanSynth-Edit Gallery
Finished MIDI-guided music edits made with SpanSynth-Edit.
Open a work's shared link to compare the original and edited audio, explore its score, and make an editable copy.
Each work has its own folder:
input.wav and output.wav: the original clip and saved generated audio, at 48 kHz mono.
original.mid and edited.mid: the source and edited score, aligned to the clip.
context.wav: the normalized audio used by the model, including the history before the… See the full description on the dataset page: https://huggingface.co/datasets/mimbres/spansynth-edit-gallery.spacr-tutorials
spaCR tutorial media
Narration and 4K video for the spaCR
interactive tutorial library, served directly to
https://einarolafsson.github.io/spacr/tutorials/.
spaCR is a toolkit for microscopy and single-cell analysis of pooled CRISPR
screens. This repository holds the media its 40-lesson tutorial player streams;
it is not a training dataset.
Why it lives here
GitHub Pages caps a published site at 1 GB. The full narration set is 2,662 MiB
across 54 voices, so the… See the full description on the dataset page: https://huggingface.co/datasets/einarolafsson/spacr-tutorials.CommonVoice-17.0-SpanishBangor-Miami-Spanish-English-Corpus
Bangor Miami Spanish-English Corpus
The Bangor Miami Corpus is a naturalistic Spanish-English code-switching speech dataset collected by Jon Russell Herring at Bangor University. It captures spontaneous bilingual conversations recorded in Miami, Florida, involving proficient Spanish-English bilinguals across multiple speaker groups.
Dataset description
Total recordings
56
Total duration
~32 h
Languages
English (en), Spanish (es)
Format
MP3 audio +… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/Bangor-Miami-Spanish-English-Corpus.spanish-speech-recognition-dataset
Spanish Speech Dataset for recognition task
Dataset comprises 10 hours of telephone dialogues in Spanish, collected from 10 native speakers across various topics and domains. It is a valuable resource for advancing speech recognition technology.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition (ASR) systems, transcribing audio, and natural language processing (NLP). - Get the data
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/spanish-speech-recognition-dataset.spanish-dialects
Dataset Card For "spanish_dialects"
Dataset Summary
This dataset contains 10+ hours of high-quality Spanish audio covering numerous speakers across Spain and Latin America. The following dialects are present in the dataset: Spain, Mexico, Chile, Argentina, Dominican Republic
voxforge_spanish
Dataset Card for voxforge_spanish
Dataset Summary
VoxForge was set up to collect transcribed speech for use with Free and Open Source Speech Recognition Engines (on Linux, Windows and Mac). They promise they will make available all submitted audio files under the GPL license, and then 'compile' them into acoustic models for use with Open Source speech recognition engines such as CMU Sphinx, ISIP, Julius and HTK. According to this, we downloaded the Spanish recordings of… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/voxforge_spanish.spanish_tokenizedmultilingual_librispeech_spanish_phoneme
Multilingual LibriSpeech Spanish Phoneme
Dataset Summary
This dataset is a curated version of the Spanish subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme).
The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into Spanish acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_spanish_phoneme.
