datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cached-activationsfable-5-coding-and-debugging-traces-synthetic-corrections
Model Synthetic Corrections
1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset.
Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.intern-debug-lerobot
debug: robot demonstrations
Instruction: Nest the three paper cups together into a single stack.
LeRobot v3.0 dataset: 51 episodes, 43257 frames, nominal 30 Hz.
Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner.
Use
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("cloudfan/intern-debug-lerobot", video_backend="torchcodec")
sample = dataset[0]
print(sample["task"]… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-debug-lerobot.chinese-lips-longform-debug
Chinese-LiPS Long-Form (zh long streaming speech)
Reconstructed continuous long-speech streams from
BAAI/Chinese-LiPS, for
slide-aware / streaming speech-translation development and evaluation. Each
source video (one speaker, one scripted lecture with slides) was released as
pre-segmented clips; here they are re-joined into the full talk.
Two variants of the same 3 talks (~97 min speech total):
config
how segments are placed
use
orig_timeline
at their original session… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-longform-debug.fable-5-coding-and-debugging-traces-synthetic
Model Synthetic Corrections
1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset.
Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/11-47/fable-5-coding-and-debugging-traces-synthetic.codeq-debugbench-dpo-pairs
codeq-debugbench-dpo-pairs
Self-generated preference pairs used to train the CodeQ iterative DPO
pipeline on top of Qwen/Qwen2.5-Coder-7B-Instruct. Each pair consists of
a chosen and rejected response to a DebugBench debugging prompt, where
preferences are derived from MCTS rollouts scored by a unit-test verifier.
Files
File
Rows
Description
round1.jsonl
1515
Raw Round 1 preference pairs (reference = base model).
round1_filtered.jsonl
936
Round 1 after… See the full description on the dataset page: https://huggingface.co/datasets/tathadn/codeq-debugbench-dpo-pairs.dpo_debugrequests_debugcodex-debugcodex-whole-debugriceGPN-14-RERUN-DEBUG
