speech-commands
ast-finetuned-speech-commands-v2MIT-ast-finetuned-speech-commands-v2-ovwav2vec2-base-finetuned-speech_commands-v0.02ast-finetuned-speech-commands-v2mlcommons-DS-CNN-SpeechCommands-fp32-onnxwav2vec2-conformer-rel-pos-large-finetuned-speech-commandsarm-DS-CNN-Large-SpeechCommands-clustered-fp32-onnxslu-direct-fluent-speech-commands-librispeech-asr
multilingual-speech-commands-15lang
Multilingual Speech Commands Dataset (15 Languages, Augmented)
This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification.
Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang.multilingual-speech-commands-3lang-raw
Multilingual Speech Commands Dataset (3 Languages, Raw)
This dataset is a curated subset of previously published speech command datasets in Kazakh, Tatar, and Russian. It is intended for use in multilingual speech command recognition and keyword spotting tasks. No data augmentation has been applied.
All files are included in their original form as released in the cited works below. This repository simply reorganizes them for convenience and accessibility.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-3lang-raw.0-9up_google_speech_commands_augmented_raw
Dataset Card for "google_speech_commands_augmented_raw_fixed"
More Information needed
speech_commandsThis is a set of one-second .wav audio files, each containing a single spoken
English word or background noise. These words are from a small set of commands, and are spoken by a
variety of different speakers. This data set is designed to help train simple
machine learning models. This dataset is covered in more detail at
[https://arxiv.org/abs/1804.03209](https://arxiv.org/abs/1804.03209).
Version 0.01 of the data set (configuration `"v0.01"`) was released on August 3rd 2017 and contains
64,727 audio files.
In version 0.01 thirty different words were recoded: "Yes", "No", "Up", "Down", "Left",
"Right", "On", "Off", "Stop", "Go", "Zero", "One", "Two", "Three", "Four", "Five", "Six", "Seven", "Eight", "Nine",
"Bed", "Bird", "Cat", "Dog", "Happy", "House", "Marvin", "Sheila", "Tree", "Wow".
In version 0.02 more words were added: "Backward", "Forward", "Follow", "Learn", "Visual".
In both versions, ten of them are used as commands by convention: "Yes", "No", "Up", "Down", "Left",
"Right", "On", "Off", "Stop", "Go". Other words are considered to be auxiliary (in current implementation
it is marked by `True` value of `"is_unknown"` feature). Their function is to teach a model to distinguish core words
from unrecognized ones.
The `_silence_` class contains a set of longer audio clips that are either recordings or
a mathematical simulation of noise.fluent_speech_commands_synth
Dataset Card for "fluent_speech_commands_synth"
More Information needed
speech-commands-v0.02
Speech Commands Dataset v0.02
This is a re-hosted copy of the Google Speech Commands v0.02 dataset in Parquet format for compatibility with the Hugging Face Dataset Viewer.
⚠️ Credits
This dataset was created by Pete Warden / Google. All credit goes to the original authors and the crowdsourcing contributors.
Original source: http://download.tensorflow.org/data/speech_commands_v0.02.tar.gz
Paper: Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/speech-commands-v0.02.
