datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SleepWalk-DatasetSleep-EDF-V2SleepEDF-V-Datasetmcp-smoke-testAwake_or_sleepWattpad-metadata-hotWattpad-metadata-newRdiffusionRdiffusion
We're releasing the entire corpus of publicly available songs from the Riffusion platform—generated and shared by their user community. Through extensive scraping of their exposed API, we’ve collected over 2,000 artificial songs, including every metadata field and downloadable asset that was accessible at the time.
📦 Included Data
Audio & Visual Assets:audio_url, audio_b64, image_url, image_b64, video_url
Metadata & Structure:id, title, author_id, created_at… See the full description on the dataset page: https://huggingface.co/datasets/sleeping-ai/Rdiffusion.TEKGEN-Wiki
TEKGEN-wiki is derived from the TEKGEN dataset released by Google Research. TEKGEN is a corpus used for fine-tuning the T5-large model to improve Knowledge Graph (KG) generation (NAACL 2021 Paper). This dataset provides the complete collection of original sentences from the TEKGEN dataset.
MemeEffect-382KExcited to release Meme Effect 382K. It is the largest known collection of Meme voice effects to train fundamental text-to-voice models that does not only tackle human emotions rather consider factor like sarcasm and popular meme culture to become more human.
We hope that researchers will consider building human centric TTS models and include our dataset in their training corpus to make text-to-speech/voice models more human.
Data fields
id: Unique identifier for the sound.… See the full description on the dataset page: https://huggingface.co/datasets/sleeping-ai/MemeEffect-382K.FakeProfile600Excited to share 600 state-of-art profiles of fake voice profiles of real people. We are releasing this for training robust and diverse text-to-speech and text-to-music models.
Data fields
displayName:Full name of the speaker (e.g., "Éloïse Gagné").
language:Language code in all caps with underscore (e.g., EN_US).
locale:Regional locale code using ISO format (e.g., fr-CA for French, Canada).
gender:Gender of the speaker (e.g., "female").
imageUrl:URL to the speaker’s image/avatar.… See the full description on the dataset page: https://huggingface.co/datasets/sleeping-ai/FakeProfile600.imnet1k_sleeping_bagfacial_expression_datasetsleep-moviesat-the-back-of-the-north-wind-sleep-series
At the Back of the North Wind — Sleep Series
A complete 3-part sleep audio series from the public domain classic by George MacDonald, paired with 100 therapeutic coloring book images.
Contents
Audio Files (250MB total, ~9 hours)
part1_audio.mp3 (3h 30m) — Chapters 1-4
part2_audio.mp3 (2h 10m) — Chapters 5-9
part3_audio.mp3 (3h 15m) — Chapters 10-38
Coloring Images (100 therapeutic line-art pages)
Sleep-themed illustrations
Wellness & Self-Care collection… See the full description on the dataset page: https://huggingface.co/datasets/ziggylott/at-the-back-of-the-north-wind-sleep-series.Sleeping-GoogleIntroducing Lyria 3, a dataset consructed through 96K artifical intelligence generated songs.
stable_diffusion_generatedSleeping-Eleven-MusicSleeppyLee1__xiaolihaokun__cosplay_Celestine_LucuSleeping-Eleven-Effectsexchangesleep-videos
