datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
splatter-cube-pbmc3k
Splatter cube: controlled scRNA-seq variants with known cluster structure
180 simulated datasets: 60 parameter points x 3 seeds, each 2,000 cells x 16,085 genes with a known number of clusters.
Generated with splatter, with baseline
parameters estimated from real reference data rather than chosen by hand.
Files
data/pP_sS.h5ad -- counts (CSR). obs["Group"] holds the ground-truth cluster label.
metrics.csv -- one row per simulation: parameters, realised sparsity… See the full description on the dataset page: https://huggingface.co/datasets/btraven/splatter-cube-pbmc3k.flow-edit-cube-triple-tkgrid-seeds30003-40004
Edit-placement (t x K x eb) campaign — OGBench cube-triple, task 2 — seeds 30003 & 40004
Seed scope: this repository contains only seeds 30003 and 40004. It is not the
full seed set for this campaign — seeds 10001 and 20002 were trained on separate hardware
and are not included here. Any per-cell mean computed from this repo alone is an n=2
estimate; see Caveats.
290 training runs from the uedit_place agent: a grid over where in the flow a
value-driven edit is applied (t), how… See the full description on the dataset page: https://huggingface.co/datasets/jaehyeokdoo2/flow-edit-cube-triple-tkgrid-seeds30003-40004.cube_text_constraints
Cube Text Constraint Evaluation Dataset (GDPval)
Overview
This dataset provides a focused collection of human-designed evaluation cases for assessing strict text-to-image constraint satisfaction in multimodal generative models.
The tasks are intentionally simple in visual appearance but logically rigid, requiring precise planning, global consistency, and explicit verification.The dataset is designed to expose repeatable failure modes where models appear to understand… See the full description on the dataset page: https://huggingface.co/datasets/8Planetterraforming/cube_text_constraints.
