datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
c4-en-html-with-metadatac4-en-html-with-training_metadata_alldesign-bench
SciModelingBench Design-Bench Data
Canonical, provenance-tracked observations for scientific modeling and design Tasks.
GitHub
·
Python Package
·
Documentation
·
Organization
This repository stores the scientific observation layer used by the
SciModelingBench Design-Bench suite. The Python package supplies validators,
Agent-visible Protocols, trusted Objectives, submission contracts, and Task
metrics. Data and evaluation logic… See the full description on the dataset page: https://huggingface.co/datasets/sci-modeling-bench/design-bench.modeling_dataDiffusion-Reward-Modeling-for-Text-Rendering-Dataset
🖼️ Text-to-Image Rendering Dataset
A dataset of 14k text prompts for image generation with text rendering evaluation
📚 Dataset Overview
This dataset contains 14,000 text prompts specifically designed for:
Image generation with text rendering
Evaluating text preservation in generated images
Training diffusion models for better text rendering
Each prompt comes with:
Pre-extracted target text for rendering
5 Stable Diffusion 3 generated latents (70k total)
Dual… See the full description on the dataset page: https://huggingface.co/datasets/leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset.repro-rethinking-genomic-modeling-through-optical-character-recognition-agent-traces
OpticalDNA reproduction — Codex agent trace
This dataset contains the raw Codex JSONL session trace for the ICML 2026
reproduction of Rethinking Genomic Modeling Through Optical Character
Recognition.
Published Trackio logbook
Paper page
Challenge instructions
Agent Trace Viewer announcement
The JSONL is uploaded directly from the matching ~/.codex/sessions entry, as
recommended by the Agent Trace Viewer. It captures the reproduction work,
Hugging Face Jobs audit, poster… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/repro-rethinking-genomic-modeling-through-optical-character-recognition-agent-traces.uplift-modeling-synthetic-benchmark
Synthetic uplift benchmark with known ground-truth treatment effect
100,000 training rows, 20,000 validation rows, generated for
uplift-modeling's Gate 0: checking that
meta-learners (S/T/X-learner) actually recover a real treatment effect before trusting them
on data where no individual ground truth is ever available - which is true of essentially
all real causal-inference data, by the fundamental problem of causal inference (nobody
observes both potential outcomes for the same… See the full description on the dataset page: https://huggingface.co/datasets/Bauxitiego/uplift-modeling-synthetic-benchmark.Churnn_modeling_datasetchurn_modeling_datasetchurn_modelingaviation-flight-control-phase-space-baseline-modeling-v0.1What this dataset tests
Whether a system can model the normal control phase-space attractor
for fly-by-wire surface channels.
The target is phase-space geometry:
dispersion
hysteresis
overshoot
lag
energy efficiency.
Required outputs
phase_space_coherence_index
baseline_dispersion_envelope
command_response_lag_profile
control_energy_efficiency
baseline_confidence
Scoring conventions
indices range 0 to 1
dispersion envelope is a low-high interval
lag profile is p50 and p95 in… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/aviation-flight-control-phase-space-baseline-modeling-v0.1.pharma-toxicity-coherence-baseline-modeling-v0.1What this dataset tests
Whether a system can define the baseline coherence of a multi-tissue biology platform.
This baseline is the reference manifold for toxicity screening.
Decoherence cannot be detected without a stable anchor.
Required outputs
baseline_coherence_score
expected_coupling_map
stress_response_synchrony_index
baseline_variance_envelope
coherence_manifold_anchor
Use case
Pre-clinical safety benchmarking.
Baseline integrity checks for organ-on-chip and multi-omic… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/pharma-toxicity-coherence-baseline-modeling-v0.1.VES-Modeling
VES Modeling Regression Data
公开回归基准:California Housing(sklearn.datasets.fetch_california_housing,seed=42 划分)。
train.csv:16512 行 × 8 特征(MedInc/HouseAge/AveRooms/AveBedrms/Population/AveOccup/Latitude/Longitude)+ target(房价中位数,单位 $100k)
test_features.csv:4128 行 × 8 特征(预测目标)
不含隐藏测试标签:隐藏标签由 Host 持有(不公开),用于 verifier 独立验证。
生成脚本:scripts/generate_regression_data.py --dataset california
项目:https://github.com/Zhuchen00123/VES-Modeling | Space(含交互… See the full description on the dataset page: https://huggingface.co/datasets/235dsds/VES-Modeling.compositional-preference-modeling
Dataset Featurization: Compositional Preference Modeling
This repository contains the datasets used in our case study on compositional preference modeling from Dataset Featurization, demonstrating how our unsupervised featurization pipeline can produce features describing human preferences and match expert-level produced features. This case study is built on top of Compositional Preference Modeling (CPM).
HH-RLHF - Featurization
Utilizing HH-RLHF dataset, we provide… See the full description on the dataset page: https://huggingface.co/datasets/Bravansky/compositional-preference-modeling.topic-modeling-maps
Topic-modelling reference maps
Pre-saved topic-modelling maps for use with the topic-modeling Python package.
Each map is a folder with a saved BERTopic model,
its 2-D UMAP projector, and a refined-labels CSV. You can drop new documents onto a map
(predict topic + 2-D coordinates, no retraining) with load_pretrained_map(...).
Maps
eu_map_60_topics
ESPON / EU map of topics — BERTopic Final_60 (SPECTER embeddings + KMeans, 60 topics),
with refined… See the full description on the dataset page: https://huggingface.co/datasets/SIRIS-Lab/topic-modeling-maps.aviation-vibration-mode-manifold-baseline-modeling-v0.1What this dataset tests
Whether a system can model a healthy airframe vibration spectrum
as a coupled-mode manifold.
It must return:
a coherence index
a coupling graph
a drift envelope
phase-conditioned baselines.
Required outputs
manifold_coherence_index
baseline_mode_coupling_graph
expected_mode_drift_envelope
phase_conditioned_baselines
baseline_confidence
Scoring conventions
indices range 0 to 1
drift envelope uses frequency and coupling tolerances
coupling graph summarizes… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/aviation-vibration-mode-manifold-baseline-modeling-v0.1.reward-modeling-papers
Reward Modeling Papers — FineSet
A research-paper dataset on Reward Modeling Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Reward Modeling Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored: quality_score… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/reward-modeling-papers.amazon-electronics-topic-modelinghn_title_modeling_datasethn_title_modeling_dataset_with_tokensperformance_modeling
