datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nemotron-student-fail-v41-clean-thinking
Nemotron-fail / DeepSeek-V4.1 clean and action-only trajectories
DeepSeek-V4.1 reward-1 trajectories for tasks on which the Nemotron student
did not obtain reward 1. This release was rebuilt from the complete reward-1
audit under v54-high-precision-canonical-reconstruction-relations.
Training paths
Path
Rows
Unique tasks
Thinking
Use
data/strict/train.jsonl.gz
11
11
Preserved and clean
Raw-thinking SFT
data/hybrid/train.jsonl.gz
58
58
Only… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuanhucs/nemotron-student-fail-v41-clean-thinking.student_partial_sciworld
student_partial_sciworld
Partial SciWorld trajectories: Qwen3-1.7B (base) acting for its first 5 turns on
the 2,120-task training split. These are the student prefixes an online-ROSE run sees
before the teacher takes over — captured separately so the take-over point can be
studied, or teacher continuations generated offline.
What is here
rows
2,120 (one per training task)
file
student_partial_sciworld.jsonl (11.4 MB)
model
Qwen/Qwen3-1.7B… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/student_partial_sciworld.
