CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01its5Q /biggest-ru-bookA bigger version of its5Q/bigger-ru-book, the smaller set being a subset of this one. Almost 1000 hours of high-quality audio. audiotext-to-speech100K<n<1M23 likes1.3k downloads1y agoHugging Face02it-just-works /shot-boundary-detection Shot Boundary Detection Dataset Overview This dataset supports research in shot boundary detection (SBD) by providing over 3.4 million uniformly structured 61-frame video clips. Each clip is centered on a key frame (the 31st), labeled as: C (Cut): a direct shot boundary, T (Transition): a gradual transition (e.g., fade, dissolve), E (Empty): no boundary present. Data is sourced from AutoShot, ClipShots, and crawled Pexels videos, with both real and synthetically… See the full description on the dataset page: https://huggingface.co/datasets/it-just-works/shot-boundary-detection.text1M<n<10M3 likes151 downloads1y agoHugging Face03itzune /antton-dataset Antton Dataset (Synthetic) This is a large-scale synthetic speech corpus designed for training and fine-tuning Basque Text-to-Speech (TTS) models. It consists of 99,996 audio files synthesized from the "Antton" voice model. This dataset was generated by Itzune and serves as the primary source for training the itzune/antton-tts (Piper version) model. Dataset Structure Due to the large volume of data (approx. 100,000 files), the dataset is organized in the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/itzune/antton-dataset.audiotext-to-speech100K<n<1M0 likes149 downloads6mo agoHugging Face04SynthFairCLIP /Synth-So-B-ITimage10M<n<100M0 likes134 downloads10mo agoHugging Face05its5Q /bigger-ru-bookaudio10K<n<100K13 likes58 downloads1y agoHugging Face06Hiswitch /cryo-et-difix-iter0image100K<n<1M0 likes38 downloads1y agoHugging Face07Exgc /iter_1audio100K<n<1M0 likes35 downloads2y agoHugging Face08itzune /maider-dataset Maider Dataset (Synthetic) This is a large-scale synthetic speech corpus designed for training and fine-tuning Basque Text-to-Speech (TTS) models. It consists of 99,996 audio files synthesized from the "Maider" voice model. This dataset was generated by Itzune and serves as the primary source for training the itzune/maider-tts (Piper version) model. Dataset Structure Due to the large volume of data (approx. 100,000 files), the dataset is organized in the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/itzune/maider-dataset.audiotext-to-speech10K<n<100K0 likes29 downloads7mo agoHugging Face09arianhosseini /gemma27b_it_math_500_generationstextn<1K0 likes15 downloads2y agoHugging Face10LeTue09 /traj-itetext10K<n<100K0 likes13 downloads5mo agoHugging Face11arianhosseini /gemma27b_it_math_128_generationstextn<1K0 likes9 downloads2y agoHugging Face12arianhosseini /gemma27b_it_math_train_generationstext1K<n<10K0 likes8 downloads2y agoHugging Face13jiwoohong93 /ita-mdt_sre ITA-MDT Pre-processed Salient Region Images for VITON-HD and DressCode Dataset Salient regions of garments have been pre-extracted and stored for VITON-HD and DressCode datasets. Information about the usage can be found at: https://github.com/jiwoohong93/ita-mdt_code Download 1. Using Python + huggingface_hub from huggingface_hub import snapshot_download snapshot_download( repo_id="jiwoohong93/ita-mdt_sre", repo_type="dataset", local_dir="./ita-mdt_sre" )… See the full description on the dataset page: https://huggingface.co/datasets/jiwoohong93/ita-mdt_sre.image10K<n<100K0 likes6 downloads1y agoHugging Face14tom-jerry-123 /Physical-AI-AV-IT PhysicalAI-AV-SFT Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language waypoint-prediction model. Contains 29,991 samples from 150 000 driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the United States. Format WebDataset — 3 uncompressed .tar shards, each containing pairs of files per sample: Entry Description {key}.jpg Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px) {key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-IT.imagerobotics10K<n<100K0 likes5 downloads5mo agoHugging Face15Krispin /itoddimage10K<n<100K0 likes4 downloads7mo agoHugging Face16Exgc /A_iter1gatedaudio1K<n<10K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.