comma
Datasets
All datasets matching “comma”multilingual-speech-commands-15lang
Multilingual Speech Commands Dataset (15 Languages, Augmented)
This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification.
Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang.comma_v0.1_training_dataset
Comma v0.1 dataset
This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T.
It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data.
If you are looknig for the raw Common Pile v0.1 data, please see this collection.
You can learn more about Common Pile in our paper.
Mixing rates and token counts
The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage.
During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.commaCarSegments
commaCarSegments
commaCarSegments is a dataset of raw CAN bus data recorded from our fleet of openpilot users driving over 300 different production vehicles, around the USA and rest of the world.
Structure
.
├── segments/ # drectory containing all data
│ ├── <device_id>/ # unique device ID directory
│ │ └── <route_id>/ # unique route (aka drive directory)
│ │ └── <segment>/ # unique segment index inside of the route… See the full description on the dataset page: https://huggingface.co/datasets/commaai/commaCarSegments.multilingual-speech-commands-3lang-raw
Multilingual Speech Commands Dataset (3 Languages, Raw)
This dataset is a curated subset of previously published speech command datasets in Kazakh, Tatar, and Russian. It is intended for use in multilingual speech command recognition and keyword spotting tasks. No data augmentation has been applied.
All files are included in their original form as released in the cited works below. This repository simply reorganizes them for convenience and accessibility.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-3lang-raw.comma1M
comma 1M
A large self-driving dataset containing one-minute driving segments with road-camera video, and full localization data.
Dataset collection
Segments were recorded by comma devices installed in real user vehicles. The collection spans several hardware generations (comma two, comma three, comma 3X, and comma four)
Camera systems vary between generations, so native resolution and field of view are not uniform across the dataset.
Each segment includes… See the full description on the dataset page: https://huggingface.co/datasets/commaai/comma1M.comma2k19
comma2k19
comma.ai presents comma2k19, a dataset of over 33 hours of commute in California's 280 highway. This means 2019 segments, 1 minute long each, on a 20km section of highway driving between California's San Jose and San Francisco. comma2k19 is a fully reproducible and scalable dataset. The data was collected using comma EONs that has sensors similar to those of any modern smartphone including a road-facing camera, phone GPS, thermometers and 9-axis IMU. Additionally, the EON… See the full description on the dataset page: https://huggingface.co/datasets/commaai/comma2k19.
