datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scientific_papers_DANCERseedance_general_all_dance_scm_latent_lmdb
Seedance General-All + Dance SCM Latent LMDB
This dataset stores precomputed SCM latents used for TurboT2AV training.
Source mapping: seedance_general_all_dance_mapping.csv
Successful latent samples: 44,305
Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007
Video latent shape per sample: (1, 16, 128, 16, 24)
Audio latent shape per sample: (1, 127, 128)
The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.dancegrpo-t2av
MiniMax H3 FL2VA First-Frame Dataset
27,815 FLUX-generated reference images paired with text prompts, built as
first-frame (image) conditions for MiniMax H3 FL2VA (text+image to
audio-video) RL training.
Dataset recipe
Prompts (prompts.txt, 27,815 lines): English video captions from
ConsisID-preview-Data,
as filtered and released by DanceGRPO
(assets/consist-id.txt).
Images (images/{index:06d}.jpg): each prompt rendered offline with
FLUX.1-dev on 8 GPUs — 400x640… See the full description on the dataset page: https://huggingface.co/datasets/zyfenghit/dancegrpo-t2av.taboo-dance
taboo-dance
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-dance")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
dancer-dataclassical-dance-navarasas
