datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3-tts-customvoice-ab-clips
qwen3-tts: full 5-way cloning comparison + cross-row diagnostic
Generated 2026-04-14 on RTX 4080 SUPER.
Directories
original/ CustomVoice.generate_custom_voice(speaker=X)
-> the ground truth voice
clone/ Base.generate_voice_clone(ref_audio=original.wav, ref_text=...)
-> full ICL clone via Base's own speaker encoder
transplant/ Base.generate_voice_clone(voice_clone_prompt=[row])
x_vector_only_mode=True… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-tts-customvoice-ab-clips.kk-yt-clipsvideo-clips-and-imgsdata-clipspyken-clipszulu-music-listening-clipswatermarking-clips
Neural Watermarking Clips
Summary
This dataset contains 31 284 short audio clips collected as an unlabeled corpus for neural audio watermarking experiments. The clips cover environmental sounds, bird vocalizations, polyphonic music with predominant instruments and synthetic but realistic jazz drums and ragtime style piano.
Source folders and file counts
ARCA23K.audio 13 470 clips
ff101bird 7 690 clips
IRMAS training 6 706 clips
WaivOps ragtime piano 1 743 clips
WaivOps… See the full description on the dataset page: https://huggingface.co/datasets/benmainbird/watermarking-clips.jdd_topic1_20251224-cliponly_sample100uclass_clipped_labeled
Dataset Card for "uclass_clipped_labeled"
More Information needed
1776-track-clips
1776 Track Clips
Fast, wordless, loop-safe background music built for Shorts, Reels, and edits.
This dataset contains 1776 unique audio clips designed for creators, editors, and developers who need clean, reusable background music without lyrics.
What’s included
1776 unique clips
No lyrics (wordless hooks)
Clean, loop-safe structure
Optimized for short-form video
Editor-first design
Preview
See the preview video(s) in this repository for examples across… See the full description on the dataset page: https://huggingface.co/datasets/BonusLockSMith/1776-track-clips.sep28k-train-4-second-clips
Dataset Card for "sep28k-train-4-second-clips"
More Information needed
OE-DCT-Movie-clipsworst100-testclean-clips
Worst-100 test-clean clips — audio, transcripts, and the vocabulary finding
The 100 LibriSpeech test-clean clips where the block-4 production model (4.60% WER) made
the most word errors — with audio embedded so the failures can be listened to, plus the
model's transcript next to the reference for each clip.
The finding this dataset produced
47% of the word errors in these clips are on words that never appeared in the 30-hour
training vocabulary at all (20,066… See the full description on the dataset page: https://huggingface.co/datasets/Diffusion-ASR/worst100-testclean-clips.music-clips-50There are 50 music clips(of 3~5 seconds).
You can load them by the following code:
from datasets import load_dataset
dataset = load_dataset('yongjian/music-clips-50')
clips = dataset['train'] # all 50 music clips
music_1_np_array = clips[0]['audio']['array'] # numpy array of shape=[N,]
Or you can directly download them from Google Drive: music-clips-50.tar.gz.
sep28k-train-3-second-clips-full-agreement
Dataset Card for "sep28k-train-3-second-clips-full-agreement"
More Information needed
sep28k-train-5-second-clips
Dataset Card for "sep28k-train-5-second-clips"
More Information needed
sep28k-dev-5-second-clips
Dataset Card for "sep28k-dev-5-second-clips"
More Information needed
sep28k-test-4-second-clips
Dataset Card for "sep28k-test-0120-4-second-clips"
More Information needed
sep28k-test-5-second-clips
Dataset Card for "sep28k-test-5-second-clips"
More Information needed
sep28k-dev-4-second-clips
Dataset Card for "sep28k-dev-0120-4-second-clips"
More Information needed
fluencybank-3-second-clips
Dataset Card for "fluencybank-3-second-clips"
More Information needed
sep28k-train-3-second-clips
Dataset Card for "sep28k-train-3-second-clips"
More Information needed
clip13sep28k-dev-3-second-clip-full-agreement
Dataset Card for "sep28k-dev-3-second-clip-full-agreement"
More Information needed
swamiji-tail-artifact-clips
Reported end-of-clip artifact: the raw clips
The exact phrases reported as having "a weird sound at the end", straight out of
the production engine, three independent draws each. No processing at all —
this is what the model produced.
phrase
text
doingwell
I am doing well, thank you.
hereif
also I am here if you need anything.
hereforyou
I am here for you.
welcome
You are very welcome!
glad
Glad you think so!
scope
I can only answer about Sri Ganapati… See the full description on the dataset page: https://huggingface.co/datasets/sw-voice/swamiji-tail-artifact-clips.fluencybank-4-second-clips
Dataset Card for "fluencybank-4-second-clips"
More Information needed
swamiji-voice-clips-clean-transcribedsep28k-test-3-second-clips-full-agreement
Dataset Card for "sep28k-test-3-second-clips-full-agreement"
More Information needed
sep28k-dev-3-second-clips
Dataset Card for "sep28k-dev-3-second-clips"
More Information needed
trance_clips
