datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
p2-etf-causal-scm-resultsSCMG_datamachinelearninglm-scm-synthetic-tabularml
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.sc_mscMMA-datasetsseedance_general_all_dance_scm_latent_lmdb
Seedance General-All + Dance SCM Latent LMDB
This dataset stores precomputed SCM latents used for TurboT2AV training.
Source mapping: seedance_general_all_dance_mapping.csv
Successful latent samples: 44,305
Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007
Video latent shape per sample: (1, 16, 128, 16, 24)
Audio latent shape per sample: (1, 127, 128)
The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.hydatascm-mechanism-drift
Structural Causal Model Environment Pairs with Mechanism Drift Labels
Paired-environment structural causal model (SCM) data with ground-truth labels for which
structural mechanism changed between two environments — plus the deterministic generator
that produces it.
Fully synthetic. No external data of any kind: nothing downloaded, scraped, purchased, or
derived from any existing corpus, dataset or benchmark. No large language model output
appears in the data, the labels, the… See the full description on the dataset page: https://huggingface.co/datasets/straxxus/scm-mechanism-drift.scm-sql
SCM-SQL — a supply-chain natural-language-to-SQL evaluation set
500 (question, gold SQL) pairs authored against the live Odoo 17 supply-chain
schema, spanning 6 explicit complexity levels including multi-turn dialogues.
Built for the dissertation Domain-Aware Multi-Agent Natural-Language-to-SQL for
Enterprise Supply Chain Intelligence by Aniruddha Prakash Kawarase (BITS Pilani
WILP, 2026). Released as a public evaluation benchmark so other researchers can
compare domain-aware… See the full description on the dataset page: https://huggingface.co/datasets/AniruddhaAI/scm-sql.ccnews_www.skornorth_scmscm-regression-mini-trialsc_MATLABsc-matlab-validated
SC MATLAB Validated
Validated MATLAB/Octave code–pseudocode pairs for program comprehension and synthesis research.
Each sample was filtered from semran1/yulan-code-MNBVC-matlab,
converted to pseudocode with Gemini, regenerated back to MATLAB, and kept only when Octave execution output matched the original.
Fields
Column
Description
sample_id
Numeric sample index
code
Original MATLAB/Octave source
pseudocode
LLM-generated pseudocode from the… See the full description on the dataset page: https://huggingface.co/datasets/philip120/sc-matlab-validated.SCM3K
SCM3K
Benchmark dataset for the paper:
The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction
Shu Wan, Abhinav Gorantla, Huan Liu, K. Selçuk Candan
3,450 tabular prediction tasks sampled from random structural causal
models (SCMs), totalling 3.45M records (1,000 samples per task).
Each task ships with the ground-truth Markov boundary of the target
node, so you can evaluate feature selection and prediction under
known causal structure. Nine feature-count… See the full description on the dataset page: https://huggingface.co/datasets/CSE472-blanket-challenge/SCM3K.ccnews_www.kalb_scmfast-autoregressive-inference-scm-train5gbViQP
ViQP: Dataset for Vietnamese Question Paraphrasing
Dataset sample
An example of 'viqp_train.json' looks as follows.
{
"source": "Trong thuật toán Caesar Cipher, ký tự K với mã hóa k=4 thì sẽ được chữ mới gì?",
"target": [
"Ký tự K với mã hóa k=4 trong thuật toán Caesar Cipher thì sẽ được chữ gì?",
"Ký tự K với mã hóa k=4 trong thuật toán Caesar Cipher thì sẽ được chữ mới gì?",
"Trong thuật toán Caesar Cipher, ký tự K với mã hóa k=4 thì sẽ được… See the full description on the dataset page: https://huggingface.co/datasets/SCM-LAB/ViQP.Resume_Screening_Data_Classificationccnews_www.fox10phoenix_scmlegal-scmccnews_www.hometownstations_scmsdxl-base-1-scm-corpus
Shamima/sdxl-base-1-scm-corpus
Synthetic image corpus generated with Stable Diffusion XL for studying the
Stereotype Content Model (SCM) structure of text-to-image latent space.
Images: 6,600
Categories: 66 occupation/identity groups
Prompt template: "A portrait of a [group], high quality."
Generator: SDXL base 1.0, DPM++ 2M Karras, 30 steps, CFG 7.0
Resolution: see image features
Fields
field
description
image
RGB JPEG
category
Group/occupation label… See the full description on the dataset page: https://huggingface.co/datasets/Shamima/sdxl-base-1-scm-corpus.ccnews_www.expressnews_scmccnews_bismarcktribune_scmccnews_www.wbay_scmlogistics-disruption-archive
Logistics Disruption Archive
Supply chain resilience metrics across 1,000 simulated logistics scenarios, covering five industry sectors under various disruption conditions.
Useful for studying how supplier diversity, delivery reliability, and inventory buffers interact to determine overall chain performance under stress.
Usage
from datasets import load_dataset
dataset = load_dataset("scm-resilience-data/logistics-disruption-archive")
df = dataset["train"].to_pandas()
Or… See the full description on the dataset page: https://huggingface.co/datasets/scm-resilience-data/logistics-disruption-archive.ccnews_newstalk1290_scmscMPRAforge_models
scMPRA stratified negative binomial fits
Fitted parameters for every stratified negative binomial (NB) and
zero-inflated negative binomial (ZINB) model reported in Modeling,
calibration, and power analysis of single-cell massively parallel reporter
assays. There are fifteen fits over three published scMPRA datasets: three canonical
(one per dataset, the model the paper's analyses use) and twelve
counterfactuals kept so the model-selection comparisons can be reproduced.
The… See the full description on the dataset page: https://huggingface.co/datasets/saarantras1/scMPRAforge_models.ccnews_wcfcourier_scmscm_geosmart_use_case
