datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2026.RA.Frontier-and-Scale-Cells
Rational-Agent Frontier, Scale, and Framing Cells
This public dataset is a sibling of siddharthmb/2026.RA.Negotiation-Campaigns (the frozen P1-P4 experimental record for the ii_mats/experiments/rational_agents negotiation program) and follows the same conventions: raw per-episode JSON, per-turn oracle annotations, Markdown/HTML transcripts, run manifests, analysis tables, and an integrity manifest over every uploaded file. It packages eight later campaigns that were run against… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Frontier-and-Scale-Cells.v11-cells-midtrain-corpus
v11 cells mid-training corpus
The delegating arm of a paired experiment: teach a 115M model to call an external
tool for arithmetic rather than to memorise the answers. Its partner, the maths-only
arm, teaches the same model to absorb the arithmetic into its weights instead.
Pre-tokenized against the v11 tokenizer
(10dd5110…, vocab 71,260), for
chrishayuk/v11-tinystories-115m-base.
Identity: 2115d6aeff3428e217ef2903a8030facd511dcb00183e9fc3faaf49d01038767
(chuk-datasets… See the full description on the dataset page: https://huggingface.co/datasets/chrishayuk/v11-cells-midtrain-corpus.cellsistant_notebook_v7
Cellsistant Notebook v7
A curated dataset for fine-tuning LLMs on JupyterLab cell manipulation and multi-turn tool-calling workflows.
Dataset Structure
Format: ShareGPT / OpenAI-style conversations.
Samples: 1,583 total (319 text-only, 286 single-turn, 978 multi-turn).
Tools Included: create_cell, execute_cell, get_cell_output, update_cell, get_notebook_content.
Use Case
Designed to train models to act as interactive assistants within JupyterLab… See the full description on the dataset page: https://huggingface.co/datasets/p4ulbr4dl3y/cellsistant_notebook_v7.
