datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seedance_general_all_dance_scm_latent_lmdb
Seedance General-All + Dance SCM Latent LMDB
This dataset stores precomputed SCM latents used for TurboT2AV training.
Source mapping: seedance_general_all_dance_mapping.csv
Successful latent samples: 44,305
Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007
Video latent shape per sample: (1, 16, 128, 16, 24)
Audio latent shape per sample: (1, 127, 128)
The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.scm-mechanism-drift
Structural Causal Model Environment Pairs with Mechanism Drift Labels
Paired-environment structural causal model (SCM) data with ground-truth labels for which
structural mechanism changed between two environments — plus the deterministic generator
that produces it.
Fully synthetic. No external data of any kind: nothing downloaded, scraped, purchased, or
derived from any existing corpus, dataset or benchmark. No large language model output
appears in the data, the labels, the… See the full description on the dataset page: https://huggingface.co/datasets/straxxus/scm-mechanism-drift.scm-regression-mini-trialSCM3K
SCM3K
Benchmark dataset for the paper:
The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction
Shu Wan, Abhinav Gorantla, Huan Liu, K. Selçuk Candan
3,450 tabular prediction tasks sampled from random structural causal
models (SCMs), totalling 3.45M records (1,000 samples per task).
Each task ships with the ground-truth Markov boundary of the target
node, so you can evaluate feature selection and prediction under
known causal structure. Nine feature-count… See the full description on the dataset page: https://huggingface.co/datasets/CSE472-blanket-challenge/SCM3K.fast-autoregressive-inference-scm-train5gblegal-scmlogistics-disruption-archive
Logistics Disruption Archive
Supply chain resilience metrics across 1,000 simulated logistics scenarios, covering five industry sectors under various disruption conditions.
Useful for studying how supplier diversity, delivery reliability, and inventory buffers interact to determine overall chain performance under stress.
Usage
from datasets import load_dataset
dataset = load_dataset("scm-resilience-data/logistics-disruption-archive")
df = dataset["train"].to_pandas()
Or… See the full description on the dataset page: https://huggingface.co/datasets/scm-resilience-data/logistics-disruption-archive.lwm-spectro-scmax-heldoutautophagycode_D_metrics_he_Qwen3-14B_lr0.0001_scm_g1autophagycode_D_metrics_he_Qwen3-0.6B_lr0.0001_scm_g1autophagycode_D_metrics_he_Qwen3-0.6B_lr0.0001_scm_g4github-issues
Dataset Card for "github-issues"
More Information needed
autophagycode_D_metrics_he_Qwen3-8B_lr0.0001_scm_g1autophagycode_D_metrics_he_Qwen3-14B_lr0.0001_scm_g10autophagycode_D_metrics_he_Qwen3-8B_lr0.0001_scm_g10autophagycode_D_metrics_train_Qwen3-8B_lr0.0001_scm_g7autophagycode_D_metrics_he_Qwen3-8B_lr0.0001_scm_g2autophagycode_D_metrics_he_Qwen3-8B_lr0.0001_scm_g3autophagycode_D_metrics_he_Qwen3-8B_lr0.0001_scm_g5autophagycode_D_metrics_he_Qwen3-8B_lr0.0001_scm_g6autophagycode_D_metrics_he_Qwen3-8B_lr0.0001_scm_g7autophagycode_D_metrics_he_Qwen3-8B_lr0.0001_scm_g8autophagycode_D_metrics_he_Qwen3-8B_lr0.0001_scm_g9autophagycode_D_metrics_he_Qwen3-14B_lr0.0001_scm_g7autophagycode_D_metrics_he_Qwen3-0.6B_lr0.0001_scm_g2autophagycode_D_metrics_he_Qwen3-0.6B_lr0.0001_scm_g3lwm-spectro-scmaxautophagycode_D_metrics_he_Qwen3-8B_lr0.0001_scm_g4autophagycode_D_metrics_he_Qwen3-14B_lr0.0001_scm_g3autophagycode_D_metrics_he_Qwen3-14B_lr0.0001_scm_g6
