datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
med-mts-audio-kokoro-82m
MTSamples‑Kokoro‑ASR (Synthetic Medical Speech)
Summary: 279 hours of synthetic English medical speech (49,462 clips) created from publicly available transcripts on MTSamples.com using multiple US/UK voices from Kokoro‑82M. Intended for training and evaluating medical ASR.
Dataset
Rows: 49,462
Total audio: ~279 hours (mono)
Source text: Sample medical reports from MTSamples.com (names/dates typically altered or removed)
Audio generation: hexgrad/Kokoro‑82M (various… See the full description on the dataset page: https://huggingface.co/datasets/darknight054/med-mts-audio-kokoro-82m.med-mts-audio-kokoro-82m-noisy16k-v1ds-b9fd6dad5a782cb8
Audio data collection
Audio data distributed as TAR archives. Download access requires manual approval by the repository owner.
Published file paths use opaque identifiers. Archive contents retain their original structure.
This collection contains 1315 source files totaling 963708815360 bytes. All files have been uploaded and checked against source checksums and destination content hashes.
ds-6cc85e8a9cff0c69
Audio data collection
Audio data distributed as TAR archives. Download access requires manual approval by the repository owner.
Published file paths use opaque identifiers. Archive contents retain their original structure.
This collection contains 673 source files totaling 1045990031360 bytes. All files have been uploaded and checked against source checksums and destination content hashes.
conversational-sarcasm-benchmark
Conversational Sarcasm Benchmark — Audio-Grounded, Metadata-Only
A benchmark of 1,168 conversational sarcasm units drawn from 64 English-language
YouTube videos (predominantly stand-up comedy and comedic conversation). Every unit
pairs a short target utterance with the preceding context that makes its
figurative reading available, and carries a categorical label plus a free-text rationale.
This repository contains no audio. It ships annotations, transcriptions, and the
source… See the full description on the dataset page: https://huggingface.co/datasets/darksyntax0/conversational-sarcasm-benchmark.ds-8ecbc2e43a099cf5
Audio data collection
Audio data distributed as TAR archives. Download access requires manual approval by the repository owner.
Published file paths use opaque identifiers. Archive contents retain their original structure.
This collection contains 210 source files totaling 257253861461 bytes. All files have been uploaded and checked against source checksums and destination content hashes.
voxceleb2-40k-part1-preprocess-all-files-separatemichaeljacksonQwen3-TTS-Cloning-Voiceshdtf_generator_preprocess_finalDark_GiovanniBarasuomenemo_datasethdtf_400_audio_60s_chunksTTS-Polish-DarkmanHDTF_audio_videoDitto_videos_hdtf_400_audio_60s_chunks_all_preprocessdarkmelocoton
