datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rlm-trajectories-seed
HotCopy RLM Trajectories (Seed)
A 12-row seed corpus of synthetic Recursive Language Model trajectories
emitted by the HotCopy two-tier agentic CLI (the orchestrator root, sub-call workers).
Why this dataset exists
The Recursive Language Model paper (Zhang, Kraska, Khattab — MIT CSAIL, 2026,
arxiv.org/abs/2512.24601) reports that
"Fine-tuning Qwen3-8B on 1,000 RLM trajectories improved performance 28.3%"
— a strong signal that the shape of RLM execution can be taught from… See the full description on the dataset page: https://huggingface.co/datasets/HotCopyAI/rlm-trajectories-seed.swe-grep-rlm-reputable-recent-5plus
swe-grep-rlm-reputable-recent-5plus
This dataset is a GitHub-mined collection of issue- or PR-linked retrieval examples for repository-level code search and localization.
Each row is built from a merged pull request in a reputable, actively maintained open-source repository. The target labels are the PR's changed files, with a focus on non-test files.
Summary
Rows: 799
Repositories: 46
Query source:
519 rows use linked issue title/body when GitHub exposed it
280 rows… See the full description on the dataset page: https://huggingface.co/datasets/13point5/swe-grep-rlm-reputable-recent-5plus.rlm-stf-v1rl-mcp
