datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biggest-ru-bookA bigger version of its5Q/bigger-ru-book, the smaller set being a subset of this one. Almost 1000 hours of high-quality audio.
shot-boundary-detection
Shot Boundary Detection
Dataset Overview
This dataset supports research in shot boundary detection (SBD) by providing over 3.4 million uniformly structured 61-frame video clips. Each clip is centered on a key frame (the 31st), labeled as:
C (Cut): a direct shot boundary,
T (Transition): a gradual transition (e.g., fade, dissolve),
E (Empty): no boundary present.
Data is sourced from AutoShot, ClipShots, and crawled Pexels videos, with both real and synthetically… See the full description on the dataset page: https://huggingface.co/datasets/it-just-works/shot-boundary-detection.antton-dataset
Antton Dataset (Synthetic)
This is a large-scale synthetic speech corpus designed for training and fine-tuning Basque Text-to-Speech (TTS) models. It consists of 99,996 audio files synthesized from the "Antton" voice model.
This dataset was generated by Itzune and serves as the primary source for training the itzune/antton-tts (Piper version) model.
Dataset Structure
Due to the large volume of data (approx. 100,000 files), the dataset is organized in the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/itzune/antton-dataset.Synth-So-B-ITbigger-ru-bookcryo-et-difix-iter0iter_1maider-dataset
Maider Dataset (Synthetic)
This is a large-scale synthetic speech corpus designed for training and fine-tuning Basque Text-to-Speech (TTS) models. It consists of 99,996 audio files synthesized from the "Maider" voice model.
This dataset was generated by Itzune and serves as the primary source for training the itzune/maider-tts (Piper version) model.
Dataset Structure
Due to the large volume of data (approx. 100,000 files), the dataset is organized in the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/itzune/maider-dataset.gemma27b_it_math_500_generationstraj-itegemma27b_it_math_128_generationsgemma27b_it_math_train_generationsita-mdt_sre
ITA-MDT Pre-processed Salient Region Images for VITON-HD and DressCode Dataset
Salient regions of garments have been pre-extracted and stored for VITON-HD and DressCode datasets.
Information about the usage can be found at:
https://github.com/jiwoohong93/ita-mdt_code
Download
1. Using Python + huggingface_hub
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="jiwoohong93/ita-mdt_sre",
repo_type="dataset",
local_dir="./ita-mdt_sre"
)… See the full description on the dataset page: https://huggingface.co/datasets/jiwoohong93/ita-mdt_sre.Physical-AI-AV-IT
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 29,991 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 3 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-IT.itoddA_iter1
