datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-speech-commands-15lang
Multilingual Speech Commands Dataset (15 Languages, Augmented)
This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification.
Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang.multilingual-speech-commands-3lang-raw
Multilingual Speech Commands Dataset (3 Languages, Raw)
This dataset is a curated subset of previously published speech command datasets in Kazakh, Tatar, and Russian. It is intended for use in multilingual speech command recognition and keyword spotting tasks. No data augmentation has been applied.
All files are included in their original form as released in the cited works below. This repository simply reorganizes them for convenience and accessibility.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-3lang-raw.0-9up_google_speech_commands_augmented_raw
Dataset Card for "google_speech_commands_augmented_raw_fixed"
More Information needed
fluent_speech_commands_synth
Dataset Card for "fluent_speech_commands_synth"
More Information needed
speech-commands-v0.02
Speech Commands Dataset v0.02
This is a re-hosted copy of the Google Speech Commands v0.02 dataset in Parquet format for compatibility with the Hugging Face Dataset Viewer.
⚠️ Credits
This dataset was created by Pete Warden / Google. All credit goes to the original authors and the crowdsourcing contributors.
Original source: http://download.tensorflow.org/data/speech_commands_v0.02.tar.gz
Paper: Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/speech-commands-v0.02.speech_commands_enriched_and_annotated
Dataset Summary
📊 Data-centric AI principles have become increasingly important for real-world use cases.At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development.
🔍 This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps the ML community in the following ways:
Enable new researchers to quickly… See the full description on the dataset page: https://huggingface.co/datasets/soerenray/speech_commands_enriched_and_annotated.fluent_speech_commands_femalehey-computer-speech-commands
Hey Computer: Speech Command Recognition
Dataset Summary
A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded.
Splits
Split
Examples
Description
train
13,192
Labeled training data
test
3,295
Public inputs with withheld target labels or annotations
Data Fields
Field
Type
audio
Audio
id
string
label
string (test sentinel: unlabeled)… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/hey-computer-speech-commands.bashkort_commands_omnivoice
Bashkort Commands OmniVoice
Partial eleven-label command snapshot generated with k2-fsa/OmniVoice
using the same cross-lingual voice-cloning recipe as
AigizK/homai_wake_word_omnivoice. Generation was stopped at the user's
request after 41,525 complete reference groups had been committed.
For every included reference row from the train split of:
bond005/sova_rudevices
the dataset contains one recording of every command:
Айвика — Russian
Айвикә — Bashkir
Айһылыу — Bashkir… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_commands_omnivoice.speech_commands_v2fluent_speech_commands_malemultilingual-speech-commands-15lang-zip
Multilingual Speech Commands Dataset (15 Languages, Augmented)
This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification.
Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang-zip.speech_commands_enrichedThis is a set of one-second .wav audio files, each containing a single spoken
English word or background noise. These words are from a small set of commands, and are spoken by a
variety of different speakers. This data set is designed to help train simple
machine learning models. This dataset is covered in more detail at
[https://arxiv.org/abs/1804.03209](https://arxiv.org/abs/1804.03209).
Version 0.01 of the data set (configuration `"v0.01"`) was released on August 3rd 2017 and contains
64,727 audio files.
In version 0.01 thirty different words were recoded: "Yes", "No", "Up", "Down", "Left",
"Right", "On", "Off", "Stop", "Go", "Zero", "One", "Two", "Three", "Four", "Five", "Six", "Seven", "Eight", "Nine",
"Bed", "Bird", "Cat", "Dog", "Happy", "House", "Marvin", "Sheila", "Tree", "Wow".
In version 0.02 more words were added: "Backward", "Forward", "Follow", "Learn", "Visual".
In both versions, ten of them are used as commands by convention: "Yes", "No", "Up", "Down", "Left",
"Right", "On", "Off", "Stop", "Go". Other words are considered to be auxiliary (in current implementation
it is marked by `True` value of `"is_unknown"` feature). Their function is to teach a model to distinguish core words
from unrecognized ones.
This version is not yet supported.
The `_silence_` class contains a set of longer audio clips that are either recordings or
a mathematical simulation of noise.speech_commands_pitchspeech_commands_enrichment_only
Dataset Card for SpeechCommands
Dataset Summary
📊 Data-centric AI principles have become increasingly important for real-world use cases.At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development.
🔍 This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps the ML community in the… See the full description on the dataset page: https://huggingface.co/datasets/renumics/speech_commands_enrichment_only.simplified_google_speech_commands_wav2vec2_960hSyntts-Commands-Media-Dataset
SynTTS-Commands: A Multilingual Synthetic Speech Command Dataset
📖 Introduction
SynTTS-Commands is a large-scale, multilingual synthetic speech command dataset specifically designed for low-power Keyword Spotting (KWS) and speech command recognition tasks. As presented in the paper SynTTS-Commands: A Public Dataset for On-Device KWS via TTS-Synthesized Multilingual Speech, this dataset is generated using advanced Text-to-Speech (TTS) technologies, aiming to… See the full description on the dataset page: https://huggingface.co/datasets/lugan/Syntts-Commands-Media-Dataset.arabic_commands_detection
Dataset Card for "arabic_commands_detection"
More Information needed
synthetic_speech_commands_PA_taggedfluent_speech_commands_test_subset_synth
Dataset Card for "fluent_speech_commands_test_subset_synth"
More Information needed
in_car_commands_26
Dataset Card for "in_car_commands_26"
More Information needed
speech-commands-lt
Speech Commands-LT (Long Tail)
Long-tail variants of Google Speech Commands v0.02
for benchmarking imbalanced audio classification.
Dataset Summary
Speech Commands v0.02 (Warden, 2018) contains 35 spoken word classes with ~1,200-3,200
samples each. This dataset applies exponential decay to the training set to create
long-tail distributions with varying imbalance ratios, simulating real-world class
imbalance in audio classification.
The _silence_ class (label 35) is… See the full description on the dataset page: https://huggingface.co/datasets/tomas-gajarsky/speech-commands-lt.speech_commands--> This is an exact copy of google/speech_commands adapted to be usable with recent datasets 🤗 versions (no remote code). <--
Dataset Card for SpeechCommands
Dataset Summary
This is a set of one-second .wav audio files, each containing a single spoken
English word or background noise. These words are from a small set of commands, and are spoken by a
variety of different speakers. This data set is designed to help train simple
machine learning models. It is covered in… See the full description on the dataset page: https://huggingface.co/datasets/beeneptune/speech_commands.google-speech-commands-wav2vec2-960hcommands_vi_voice_syntheticsimplified-google-speech-commands-wav2vec2-960hspeech_commands_augmented_alpha200speech_commands_10labels_50000synthin_car_commands_60
Dataset Card for "in_car_commands_60"
More Information needed
speech-commands-mini
