datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common_voice_13_0_dv_preprocessed
Dataset Card for Common Voice Corpus 13.0
Dataset Summary
The Common Voice dataset consists of a unique MP3 and corresponding text file.
Many of the 27141 recorded hours in the dataset also include demographic metadata like age, sex, and accent
that can help improve the accuracy of speech recognition engines.
The dataset currently consists of 17689 validated hours in 108 languages, but more voices and languages are always added.
Take a look at the Languages page to… See the full description on the dataset page: https://huggingface.co/datasets/ssahir/common_voice_13_0_dv_preprocessed.ds007808-sub01-speechopen-pangolin-preprocessed
ds007808 sub-01 / speechopen / pangolin — preprocessed EEG↔speech windows
Ready-to-train EEG↔speech windows for replicating the scaling experiment of
Sato et al. 2024, "Scaling Law in Neural Data: Non-Invasive Speech Decoding with 175 Hours
of EEG Data" (arXiv:2407.07595), built from the public
ds007808 dataset (arXiv:2606.01264).
Slice = subject sub-01, task speechopen (overt speech), device pangolin (128-ch
g.Pangolin) — the rig matching the 175 h paper. Each example is one… See the full description on the dataset page: https://huggingface.co/datasets/tankalapavankalyan/ds007808-sub01-speechopen-pangolin-preprocessed.sada-validation-preprocessed
Details
This is the SADA 2022 dataset with the input_features whish are log mels and the cleaned_labels which is the tokenized version of the cleaned_text. You can directly use this as the validation dataset when training Whisper Tiny, Small, Base & Medium models, as they all use the same tokenizer. Please double check this as well from the original model repo.
In addtition, the following filters were applied to this data:
All audios are less than 30 seconds and greater than 0… See the full description on the dataset page: https://huggingface.co/datasets/mosama/sada-validation-preprocessed.preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2605
5bfe2d098c8486d97fac8be76d86ec9146435245
train
56:46:32
50,557
589,095
11.7
31.9
techiaith/corpws-clllc-wlga
5d00294c31c78b1d7937bb2c2bc6cc70bc18d410
clips
48:20:49
27,579… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607.common_voice_13_0_dv_preprocessed
Dataset Card for Common Voice Corpus 13.0
Dataset Summary
The Common Voice dataset consists of a unique MP3 and corresponding text file.
Many of the 27141 recorded hours in the dataset also include demographic metadata like age, sex, and accent
that can help improve the accuracy of speech recognition engines.
The dataset currently consists of 17689 validated hours in 108 languages, but more voices and languages are always added.
Take a look at the Languages page to… See the full description on the dataset page: https://huggingface.co/datasets/ferno22/common_voice_13_0_dv_preprocessed.common_voice_13_0_dv_preprocessed
Dataset Card for Common Voice Corpus 13.0
Dataset Summary
The Common Voice dataset consists of a unique MP3 and corresponding text file.
Many of the 27141 recorded hours in the dataset also include demographic metadata like age, sex, and accent
that can help improve the accuracy of speech recognition engines.
The dataset currently consists of 17689 validated hours in 108 languages, but more voices and languages are always added.
Take a look at the Languages page to… See the full description on the dataset page: https://huggingface.co/datasets/fmagot01/common_voice_13_0_dv_preprocessed.preprocessed-whisper-btb-cv-cvad-cven-wlga-ca-ec-2601
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-cven-wlga-ca-ec-2601
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2601
main
train
52:38:17
48,608
573,772
11.8
27.7
techiaith/corpws-clllc-wlga
main
clips
20:00:52
18,905
228,962
12.1
10.5
techiaith/commonvoice_23_0_cy
main… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-cven-wlga-ca-ec-2601.preprocessed-whisper-btb-cv-cvad-wlga-ca-2603
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2603
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2602
main
train
52:45:22
48,569
556,542
11.5
28.2
techiaith/corpws-clllc-wlga
2603_rc3
clips
41:54:21
29,446
447,780
15.2
22.4
techiaith/commonvoice_23_0_cy
main… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2603.preprocessed-whisper-btb-cv-cvad-wlga-ca-2606
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2606
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2605
main
train
57:14:31
50,934
594,564
11.7
29.5
techiaith/corpws-clllc-wlga
main
clips
52:27:50
29,794
544,066
18.3
27.1
techiaith/commonvoice_23_0_cy
main… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2606.preprocessed-whisper-btb-cv-cvad-wlga-ca-2602
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2602
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2602
main
train
52:45:22
48,569
556,542
11.5
32.2
techiaith/corpws-clllc-wlga
main
clips
20:02:39
18,851
228,230
12.1
12.2
techiaith/commonvoice_23_0_cy
main
train+dev+other_with_excluded… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2602.
