datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cibench-experiments
CIBench Experiments
Reproducibility packages for CIBench — the stateless, replayable benchmark engine for the 1M–10M token long-context era.
If a benchmark result cannot be replayed from its manifest alone, it did not happen.
Every sub-directory in this dataset is a self-contained experiment package: per-run manifests, content-addressed canonical JSON, ResultRecord with full scoring + signed provenance, per-item OpenTelemetry gen_ai_* call metrics, retrieved evidence, a… See the full description on the dataset page: https://huggingface.co/datasets/publicus-ai/cibench-experiments.mbpp-code-rl
MBPP for code RL (deduplicated against MBPP+)
MBPP prepared for RLVR training in verl,
with two independent hold-outs so both MBPP+ and MBPP's own canonical test
split stay reportable after training on this data.
split
rows
contents
train
320
MBPP canonical train + validation + prompt, minus everything in MBPP+
test
378
exactly the problems in evalplus/mbppplus
heldout_mbpp_test
276
MBPP's canonical test split (task_id 11-510) that is not in MBPP+… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/mbpp-code-rl.moltbook-ec-10m-base-model-experiments
MoltBook Base Model Experiments — 10 min runs
Multi-agent social simulation data comparing base (pretrained) vs RL-tuned (instruct) models on MoltBook. This dataset tests whether entropy collapse in multi-agent discourse is driven by RL post-training.
Experiment Design
All experiments use the same split architecture:
Orchestrator: Google Gemini 3.1 Flash Lite (via OpenRouter) — handles agency (browsing, voting, deciding when to post)
Content generator: One of 3 models —… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-ec-10m-base-model-experiments.moltbook-ec-1h-base-model-experiments
MoltBook Base Model Experiments — 1 hour runs
Multi-agent social simulation data from base (pretrained) model content generation on MoltBook. This dataset tests whether entropy collapse in multi-agent discourse is driven by RL post-training.
Experiment Design
All experiments use a split architecture:
Orchestrator: Google Gemini 3.1 Flash Lite (via OpenRouter) — handles agency (browsing, voting, deciding when to post)
Content generator: Qwen 3.5 35B A3B Base (pretrained… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-ec-1h-base-model-experiments.moltbook-entropy-collapse-experiments
MoltBook Entropy Collapse Experiments
Multi-agent social simulation data from the Entropy Collapse experiment series run on MoltBook, a Reddit-like social network for AI agents.
Overview
This dataset contains interaction logs from experiments where autonomous AI agents interact on a social platform. The experiments investigate how initial content seeding affects the diversity and dynamics of agent-generated discourse — specifically, whether and how quickly agent… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-entropy-collapse-experiments.nyt-connections-experiments
NYT Connections Experiments Dataset
This dataset contains training, validation, and test splits for fine-tuning language models on New York Times Connections puzzles. It includes three experimental configurations examining data augmentation, reasoning format, and curriculum learning.
Dataset Overview
NYT Puzzles: 831 total (673 training, 74 validation, 84 test)
Synthetic Puzzles: 200 total (162 training, 18 validation, 20 test)
Pre-Connections Tasks: 720 training… See the full description on the dataset page: https://huggingface.co/datasets/nickting/nyt-connections-experiments.
