datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpn-star-p-uniform-v1-background
marin-dna/gpn-star-p-uniform-v1-background
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the background region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-background.vertebrate-v1-background
marin-dna/vertebrate-v1-background
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This
draft covers the background region cohort with all species
scope and preserves source FASTA/2bit letter case.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence case is independent of that filter: lowercase bases preserve source
repeat masking, uppercase bases preserve source… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-background.AI-Background-RemoverImagePulseV2-Edit-Background
ImagePulseV2 Dataset - Background Replacement
The ImagePulseV2 dataset is a collection we constructed for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Background.pi05-libero-plus-background-textures-failures
pi0.5 LIBERO-plus Background Textures Failures
Failure summary artifacts for evaluating TensorAuto/tPi0.5-libero on the LIBERO-plus Background Textures perturbation subset through OpenTau.
Evaluation Setup
Benchmark: LIBERO-plus
Suite: libero_10
Perturbation category: Background Textures
Policy: TensorAuto/tPi0.5-libero
Tasks: 289
Episodes per task: 5
Metric episodes: 1445
Seed schedule: episode seeds 1000 to 1004
Episode length: 520 steps
Source machine: wzxuan… See the full description on the dataset page: https://huggingface.co/datasets/d3d3shan/pi05-libero-plus-background-textures-failures.cc-wet-background-2026-04
Common Crawl background sample — CC-MAIN-2026-04
5 randomly sampled WET files (of 100,000; awk srand(42) selection, list in
sample.paths) from the January 2026 Common Crawl, downloaded from
data.commoncrawl.org on 2026-09-06.
114,234 extracted-text documents, ~0.84 GB plain text (351 MB gzipped).
Purpose: negative-control corpus for phrase-fingerprint false-positive
measurement. The crawl predates the May–July 2026 event under study, so any
phrase hit here is by construction a… See the full description on the dataset page: https://huggingface.co/datasets/thisfffsd/cc-wet-background-2026-04.backgroundeli-why-perceived-background-match
ELI-Why Perceived Background Match
🧠 Dataset Summary
This split contains human judgments on whether an LLM-generated explanation was perceived to match the intended educational background of the audience (e.g., elementary, high school, graduate school).
Each example in this dataset includes:
The original question
The intended education level (based on prompting)
The explanation generated according to the intended education level
The perceived education level (based on… See the full description on the dataset page: https://huggingface.co/datasets/INK-USC/eli-why-perceived-background-match.
