datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
recursive-tasktrove-out
AnshKetchum/tasktrove-recursive-task-synthesis
agent-task-recursive-task-synthesis
Apptainer pool for hamishivi/agent-task-recursive-task-synthesis
This repository hosts tmax-compatible SIF images and a unified download manifest. Training data and task archives are in hamishivi/agent-task-recursive-task-synthesis. The manifest includes earlier images hosted under hamishivi and new images hosted under TMaxxx; the downloader selects the correct repository and immutable commit for each image.
Apptainer images
The pool currently contains 29,501 / 29… See the full description on the dataset page: https://huggingface.co/datasets/TMaxxx/agent-task-recursive-task-synthesis.Recursive-Task-Synthesis
Recursive Task Synthesis
This dataset contains 37,484 validated command-line task instances produced
through recursive task synthesis. Public identifiers are opaque and stable.
metadata/tasks.parquet: one searchable row per task instance.
metadata/shard_manifest.jsonl: TAR sizes and SHA256 checksums.
data/tasks-*.tar: sanitized runnable task packages.
The searchable task rows include:
instruction: contents of instruction.md.
task_toml: contents of task.toml.
solution:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis.agent-task-recursive-task-synthesis
Recursive-Task-Synthesis for tmax
Images require building: the complete dataset and build contexts are included. Image builds are deferred; run the resumable script below before using these environments.
All 37,484 task directories from Zhongzhi1228/Recursive-Task-Synthesis, pinned to be44f96808d5a9b599d5cb024341ff00091adeb7, converted to tmax's swerl_vanillux_sandbox format.
The train split uses the same messages, ground_truth, dataset, env_config, and source schema as the… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/agent-task-recursive-task-synthesis.Recursive-Task-Synthesis-Trajectories
Recursive Task Synthesis Trajectories
This dataset contains 327,189 completed agent trajectories collected on
recursively synthesized command-line tasks. Public identifiers are opaque and
stable.
The trajectory JSON retains messages, actions, observations, and token counts.
Token-level log-probability arrays and duplicated debug/session captures are
excluded from the public packages.
metadata/trajectories.parquet: searchable trajectory metadata.
metadata/shard_manifest.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Trajectories.RecursiveDataBridge_91_10YTrecursive-tsfm-data
Recursive TSFM training data
This repository contains the four source pools used for the experiments in
ecntu/recursive-tsfm, frozen as separate
Hugging Face datasets artifacts so sampling recipes remain explicit and auditable.
source
rows
origin
gifteval
500,000
Salesforce/GiftEvalPretrain
tsmixup
3,000,000
Chronos training_corpus_tsmixup_10m
kernel
600,000
Chronos training_corpus_kernel_synth_1m
gifteval_trainval
115,639
leakage-safe histories from… See the full description on the dataset page: https://huggingface.co/datasets/emiliocantuc/recursive-tsfm-data.Sequential-Math
RecursiveMAS Sequential-Math
Project Page | Code | Paper
We introduce RecursiveMAS, a multi-agent framework that scales agent collaboration through latent-space recursion. This dataset contains training examples for the Sequential-Style setting.
Dataset Details
Item
Description
Dataset
RecursiveMAS/Sequential-Math
Original file
Sequential-Math.json
Collaboration style
Sequential-Style
Used for
sequential math inner agents and outer RecursiveLink… See the full description on the dataset page: https://huggingface.co/datasets/RecursiveMAS/Sequential-Math.recursive-cognition-corpus
LuisCore Recursive Cognition Corpus
LuisCore is a low-latency decentralized runtime substrate for multi-step inference at scale.
Generated: 2026-09-24T11:09:13.307Z
Rows: 13236
Owner: Luis610348
Canonical site: https://luiscore.com
What this dataset is
LuisCore is a recursive cognition infrastructure. This dataset is the public
LLM Discovery Corpus — a stable, deterministic Q&A set used by LuisCore to
help language models accurately describe, cite, and verify… See the full description on the dataset page: https://huggingface.co/datasets/Luis610348/recursive-cognition-corpus.recursive-task-synthesis-glm-5.3-rollouts
GLM 5.3 agentic rollouts on Recursive-Task-Synthesis
This dataset catalogs the full collection made from the pinned
Recursive-Task-Synthesis dataset revision
be44f96808d5a9b599d5cb024341ff00091adeb7. The repository includes approximately 260.5 GiB of trajectory payload tar shards.
Contents at a glance
Item
Count
Source tasks considered
37,284
Source candidates inspected
19,368
Converted tasks after source filters
18,600
Tasks passing gold… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts.Recursive-Task-Synthesis-Quality-1K
Recursive Task Synthesis Quality 1K
This dataset contains 1,000 quality-selected, validated command-line task
instances. It is a curated subset of the
Recursive Task Synthesis dataset.
Public task and group identifiers are opaque and stable across both datasets.
Selection
The subset was selected from 37,484 validated tasks using structural and safety
checks, two-pass semantic review, strict gates for instruction clarity,
instruction-verifier alignment, verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Recursive-Task-Synthesis-Quality-1K.appworld-rollouts-recursive-14b-mar06Recursive-Task-Synthesis
Recursive Task Synthesis
Only tasks with completed PUBLIC platform container and VM artifacts at the
2026-09-17T02:53:50.654736+00:00 registry audit are included. The excluded IDs, reasons
and last failed build IDs are recorded in exclusions.json. Exclusions affect
both metadata rows and complete TAR task packages. This is a hosting filter,
not gold validation. Recheck both artifact types and remove recovered IDs from
the preparation sidecar to restore them from the pinned… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Recursive-Task-Synthesis.Recursive-Task-Synthesis-Copy
Recursive Task Synthesis
This dataset contains 37,484 validated command-line task instances produced
through recursive task synthesis. Public identifiers are opaque and stable.
metadata/tasks.parquet: one searchable row per task instance.
metadata/shard_manifest.jsonl: TAR sizes and SHA256 checksums.
data/tasks-*.tar: sanitized runnable task packages.
The searchable task rows include:
instruction: contents of instruction.md.
task_toml: contents of task.toml.
solution:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/Recursive-Task-Synthesis-Copy.pubmedqa-recursive-llm-degradation-qwen2.5-0.5b
PubMedQA Recursive LLM Degradation — Qwen2.5-3B
This repository contains synthetic biomedical question-answering data
and model predictions generated as part of a study of recursive
fine-tuning and model degradation.
Base Model
Qwen/Qwen2.5-3B
Source Dataset
The experiments use the PubMedQA dataset:
qiaoxin/PubMedQA
This repository contains generated/derived research artifacts and does
not redistribute the original PubMedQA dataset in its entirety.… See the full description on the dataset page: https://huggingface.co/datasets/chrislimbe/pubmedqa-recursive-llm-degradation-qwen2.5-0.5b.pubmedqa-recursive-llm-degradation-qwen2.5-3b
PubMedQA Recursive LLM Degradation — Qwen2.5-3B
This repository contains synthetic biomedical question-answering data
and model predictions generated as part of a study of recursive
fine-tuning and model degradation.
Base Model
Qwen/Qwen2.5-3B
Source Dataset
The experiments use the PubMedQA dataset:
qiaoxin/PubMedQA
This repository contains generated/derived research artifacts and does
not redistribute the original PubMedQA dataset in its entirety.… See the full description on the dataset page: https://huggingface.co/datasets/chrislimbe/pubmedqa-recursive-llm-degradation-qwen2.5-3b.RecursiveNodeStack_511_8FXrecursive-latent-reasoning
Recursive Latent Reasoning — datasets
Data for the Recursive Latent Reasoning project: one shared recursive generator
(a TRM-style weight-tied block refining a VAE latent canvas, with a frozen multi-scale
MAE as the feature ruler) applied to three tasks.
domain
unit
input → target
objective
crystal/
a SLICES row
target band_gap → a valid crystal
MAE-based generation, 2 arms: plain vs GFN multi-reward
sudoku/
a (puzzle, solution) pair
puzzle → its unique solution… See the full description on the dataset page: https://huggingface.co/datasets/iamseungpil/recursive-latent-reasoning.recursive-lines
Recursive Lines: A Dual-Track Adversarial Benchmark
Recursive Lines is a diagnostic suite for detecting "High-Agency Deception" in Large Language Models. It serves as the reference implementation for the Constraint Cascade Model (FAccT 2026) and the Agency Index metric.
1. Overview
Current LLM benchmarks measure capability (MMLU) or safety (Refusal). They fail to measure Agency—the thermodynamic distinction between stochastic error (hallucination) and strategic intent… See the full description on the dataset page: https://huggingface.co/datasets/OstensibleParadox/recursive-lines.openclaw-recursive-study-data
OpenClaw Recursive Repository Study Data
Synthetic repository-study data generated against
openclaw/openclaw at commit
da228660306b55a9cce3b973946f3aacfc515848. The source repository is MIT licensed.
This release contains exploration questions, tool-using study trajectories,
recursive notes, full recall-rewritten trajectories, and recall-to-action
training examples. Nested chat/tool objects are stored as JSON strings to keep
the schema stable and can be decoded with json.loads.… See the full description on the dataset page: https://huggingface.co/datasets/aviralku/openclaw-recursive-study-data.Distillation-Code
RecursiveMAS Distillation-Code
Project Page | Code | Paper
We introduce RecursiveMAS, a multi-agent framework that scales agent collaboration through latent-space recursion. This dataset contains training examples for the Distillation-Style setting.
Dataset Details
Item
Description
Dataset
RecursiveMAS/Distillation-Code
Original file
Distillation-Code.json
Collaboration style
Distillation-Style
Used for
expert/learner code inner agents and outer… See the full description on the dataset page: https://huggingface.co/datasets/RecursiveMAS/Distillation-Code.ouroboros-recursive-self-improvement-capability-report
Ouroboros: Recursive Self-Improvement Without Model-Weight Updates
This repository contains the white paper in Markdown, PDF, and DOCX formats.
The paper presents Ouroboros as a system-level recursive self-improvement capability: a persistent operational control plane that can improve the machinery around fixed-weight foundation models, verify and adopt durable changes, and reuse those changes in later improvement cycles. The detailed run evidence remains private.… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-recursive-self-improvement-capability-report.qwen3-4b-recursive-sft-v3.1-eval-rollouts
Recursive V3.1 SFT Evaluation Rollouts
Saved evaluation generations from
tawer12/qwen3-4b-recursive-sft-v3.1.
This dataset contains SFT-model evaluations only, not RL training rollouts,
flattened-model generations, or the SFT training corpus.
The release preserves all original per-rollout fields and text, including
invalid trees and wrong answers. Added fields identify the benchmark, attempt,
model, and source file/line. No generations or answer labels were repaired.… See the full description on the dataset page: https://huggingface.co/datasets/tawer12/qwen3-4b-recursive-sft-v3.1-eval-rollouts.Sequential-Code
RecursiveMAS Sequential-Code
Project Page | Code | Paper
We introduce RecursiveMAS, a multi-agent framework that scales agent collaboration through latent-space recursion. This dataset contains training examples for the Sequential-Style setting.
Dataset Details
Item
Description
Dataset
RecursiveMAS/Sequential-Code
Original file
Sequential-Code.json
Collaboration style
Sequential-Style
Used for
sequential code inner agents and outer RecursiveLink… See the full description on the dataset page: https://huggingface.co/datasets/RecursiveMAS/Sequential-Code.recursive-calibforge-outrecursive-memory-perfectblend-coding
Frozen PerfectBlend + Coding training data
Private migration snapshot of the raw, Qwen-generated training trajectories used
by the all-turn residual Compressor dataset frozen on 2026-09-06. This is not
the untouched upstream PerfectBlend dataset or a newly generated corpus.
Corpus
Trajectories
Generated assistant responses
Raw bytes
PerfectBlend / xhigh
40,596
51,039
549,854,699
Coding
37,560
86,880
1,274,359,351
Total
78,156
137,919
1,824,214,050
The mixture… See the full description on the dataset page: https://huggingface.co/datasets/mocoV3/recursive-memory-perfectblend-coding.Mixture-Math
RecursiveMAS Mixture-Math
Project Page | Code | Paper
We introduce RecursiveMAS, a multi-agent framework that scales agent collaboration through latent-space recursion. This dataset contains training examples for the Mixture-Style setting.
Dataset Details
Item
Description
Dataset
RecursiveMAS/Mixture-Math
Original file
Mixture-Math.json
Collaboration style
Mixture-Style
Used for
math specialist inner agent training
Split
train
Rows
1904… See the full description on the dataset page: https://huggingface.co/datasets/RecursiveMAS/Mixture-Math.Vinayak-Multistep-Recursive-Reasoning-Benchmark
Vinayak Multistep Recursive Reasoning Benchmark (VMRRB)
Overview
The Vinayak Multistep Recursive Reasoning Benchmark (VMRRB) is a large-scale prompt-based benchmark designed to evaluate advanced reasoning, recursive dependency resolution, encrypted task traversal, and robustness capabilities of frontier AI systems.
The benchmark evaluates a model's ability to:
Perform recursive multistep reasoning
Resolve interdependent question chains
Execute encrypted dependency… See the full description on the dataset page: https://huggingface.co/datasets/bepipeV/Vinayak-Multistep-Recursive-Reasoning-Benchmark.samantha-r01-recursive-reasoning-corpus
Samantha R01 Recursive Reasoning Corpus
Answer-only corpus for the first isolated Samantha silent-tick / recursive-latent-reasoning validation. It is normalized for pre_train_recursive_reasoning.py and intentionally contains no visible chain-of-thought or source rationale fields.
The repository is private because it combines sources with mixed or unspecified redistribution terms. Access does not supersede any upstream license.
Split policy
train: 50,000… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/samantha-r01-recursive-reasoning-corpus.RECURSIVE_LOOP_1
