datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
beat2-additional-annotations
BEAT2 Official Release + Additional Annotations
This is a fork of H-Liu1997/BEAT2
that adds annotations contributed by the
RAG-Gesture (CVPR 2025)
and MIBURI (CVPR 2026) projects.
The base BEAT2-English data (motion, audio, TextGrids, semantic labels,
pretrained motion-autoencoder weights) is inherited verbatim from upstream;
the additional annotations from RAG-Gesture and MIBURI are pushed on top.
Citations
If you use only the original BEAT2 dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/m-hamza-mughal/beat2-additional-annotations.mp3quran-audioRecorrected_Classification_Data_filtered_traindataset_asheeg-motor-imagery-dataRecorrected_Classification_Data_filtered_train_22RepeatAudio
Dataset Card for RepeatAudio
Audio datasets referenced in Class-Agnostic Audio Repetition Counting. Contains two main sections:
RS, RSN and RVN: Synthetic datasets containing varying levels of noise. Uniformly 10 seconds long, with 0-8 repetition events contained in each sample.
Clocks, Heartbeats and Dolphins: Real-world derived samples with variable length across mechanical, ecological and medical domains.
Relevant code can be found in this repo.
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/Hamozwa/RepeatAudio.common-voice-26-en-audio
English — English (en)
This datasheet is for cv-corpus-26.0-2026-06-12 of the Mozilla Common Voice Scripted Speech dataset for English [English - en]. The dataset contains 2583051 clips representing 3780.95 hours of recorded speech (2784.88 hours validated) from 100172 speakers, recorded from a text corpus of 1,721,897 sentences.
Language
English is a West Germanic language with origins in England. There are an estimated 1.5 billion English speakers, making it the… See the full description on the dataset page: https://huggingface.co/datasets/hamayawayuri/common-voice-26-en-audio.machine_noise_dataset
Machine Sound Doctor Dataset
Labeled audio clips of running machinery, used to train the classifier
behind Machine Sound Doctor —
a phone-based predictive maintenance tool: call in, hold your phone near a
running machine, get an SMS back diagnosing the sound.
Classes
Class
Clips
Description
normal
33
Machine running normally
bearing_fault
33
Bearing fault sound signature
belt_slip
33
Belt slip sound signature
99 clips total. Most are 16-bit PCM… See the full description on the dataset page: https://huggingface.co/datasets/hamezksm/machine_noise_dataset.darija-stt-dataset
Dataset Card for "darija-stt-dataset"
More Information needed
persian-tts-dataset-maleRecorrected_Classification_Data_filtered_syr_validatedRecorrected_Classification_Data_filtered_train2hamza-belloumi-tunisian-tts
Hamza Belloumi Tunisian TTS Dataset
A Tunisian Arabic speech dataset for TTS model training.
uclass_clipped_labeled
Dataset Card for "uclass_clipped_labeled"
More Information needed
asr-taskRecorrected_Classification_Data_samplesTest_Data_filtered_samplesStethoBench
StethoBench
StethoBench is a comprehensive benchmark for cardiopulmonary auscultation, comprising 77,027 instruction–response pairs synthesized from 16,125 labeled recordings across 11 public datasets. It is the training and evaluation benchmark for StethoLM, published in the Transactions on Machine Learning Research (TMLR).
Dataset Description
StethoBench was constructed by synthesizing instruction–response pairs from existing labeled cardiopulmonary audio datasets… See the full description on the dataset page: https://huggingface.co/datasets/hamedfrogh/StethoBench.jalandhary_asr_enhancedfb_labeled_v48dretna_daridjamsa_law_asr_testmemoni-audio
Memoni Audio Dataset
Cleaned, voice-activity-detected speech segments for Arabic/regional dialect ASR.
Usage
from datasets import load_dataset
ds = load_dataset("hamzahanif/memoni-audio", split="train", streaming=True)
sample = next(iter(ds))
print(sample["audio"])
print(sample["unique_id"])
print(sample["channel_name"])
asr-task-testdarmyst_single_and_long_utt_balancedfb_labeledfb_labeled_v6_w2v2
