datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rStar-Coder
rStar-Coder Dataset
Project GitHub | Paper
Dataset Description
rStar-Coder is a large-scale competitive code problem dataset containing 418K programming problems, 580K long-reasoning solutions, and rich test cases of varying difficulty levels. This dataset aims to enhance code reasoning capabilities in large language models, particularly in handling competitive code problems.
Experiments on Qwen models (1.5B-14B) across various code reasoning benchmarks demonstrate… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/rStar-Coder.RSTeller
⚠️ Usage Warning
This is the latest version of RSTeller, updated on 2025-01-28. Users who accessed this dataset before this date can find the legacy version, which is preserved for reference. Additionally, we have released the metadata for this dataset.
For the details and the usage of the dataset, please refer to our github repository page.
Citation
If you find the dataset and our paper useful, please consider citing our paper:
@article{ge2025rsteller… See the full description on the dataset page: https://huggingface.co/datasets/SlytherinGe/RSTeller.R-Star-Distillation-BackupsrStar-Coder-seed-testusaco_2025
USACO 2025 Open Contest Dataset
Dataset Description
The USA Computing Olympiad (USACO) is a prestigious algorithmic programming competition for high school students in the United States, consisting of four difficulty levels: Bronze, Silver, Gold, and Platinum. Each level contains a set of challenging problems that test algorithmic thinking and implementation skills, making USACO a valuable benchmark for evaluating the reasoning and problem-solving capabilities of large… See the full description on the dataset page: https://huggingface.co/datasets/rStar-Reasoning/usaco_2025.RSTeller_legacy
⛔ Usage Warning
This is the legacy version of the RSTeller dataset and is not the latest version referenced in our paper. We are keeping it available here to provide the community with easy access to additional data.
For the details and the usage of the dataset, please refer to our github page.
Citation
If you find the dataset and our paper useful, please consider citing our paper:
@article{ge2025rsteller,
title={RSTeller: Scaling up visual language modeling in… See the full description on the dataset page: https://huggingface.co/datasets/SlytherinGe/RSTeller_legacy.RST-SFT-Qwen3.5-27B
RST SFT trajectories for Qwen3.5-27B
Multi-turn terminal-agent conversations distilled from
Zhongzhi1228/Recursive-Task-Synthesis-Trajectories,
ready for supervised fine-tuning of Qwen/Qwen3.5-27B.
Pipeline, launchers, and the full plan: https://github.com/k1ssloo/RST-Train
cap10 reproduces the paper's SFT example count exactly
The source release has 327,189 trajectories. cap10 ends at 10,778 examples —
the count arXiv:2608.05466v3 states it trained
on. That was… See the full description on the dataset page: https://huggingface.co/datasets/NiuNiu0110/RST-SFT-Qwen3.5-27B.RSTeller_metadata
Metadata for RSTeller
This dataset contains the necessary metadata for the dataset SlytherinGe/RSTeller.
Dataset Details
Description
The metadata table provides detailed information for the RSTeller dataset, with the following columns:
patch_id: The primary key of the table, corresponding to the "__key__" or the "patch_id" field in the JSON of the RSTeller dataset.
patch_lat and patch_lon: The latitude and longitude coordinates of the patch center in WGS84… See the full description on the dataset page: https://huggingface.co/datasets/SlytherinGe/RSTeller_metadata.rStar-Coder-seed-sftrstar-coder-rl-clean
rStar-Coder RL — held-out test split
5,000 held-out problems from the microsoft/rStar-Coder synthetic_rl RL set
(arXiv:2505.21297), for in-domain evaluation.
Test only by design. The training set is built from the original
microsoft/rStar-Coder at load time; only this held-out split is pinned here so every
run evaluates on exactly the same problems. rstar_coder_loader.py removes any training
problem whose question text appears here, so train and test stay disjoint.
question —… See the full description on the dataset page: https://huggingface.co/datasets/Ilia2003Mah/rstar-coder-rl-clean.RS-testrstar_sft
Dataset for rStar-Math
https://github.com/microsoft/rStar
rStar-Coder
rStar-Coder Dataset
Project GitHub | Paper
Dataset Description
rStar-Coder is a large-scale competitive code problem dataset containing 418K programming problems, 580K long-reasoning solutions, and rich test cases of varying difficulty levels. This dataset aims to enhance code reasoning capabilities in large language models, particularly in handling competitive code problems.
Experiments on Qwen models (1.5B-14B) across various code reasoning benchmarks demonstrate… See the full description on the dataset page: https://huggingface.co/datasets/cublya/rStar-Coder.rStarCoder-Human-Test-16krstarCDMAM_RSTtorgo-audio-dataset-with-audioRS-test-fix2rStar-Coder-PyStructure
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): en
License: cc-by-4.0
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More… See the full description on the dataset page: https://huggingface.co/datasets/broadfield-dev/rStar-Coder-PyStructure.cleand_microsoft_rStar-Coder元データ: https://huggingface.co/datasets/microsoft/rStar-Coder
データ件数: 269,863
平均トークン数: 11674
最大トークン数: 31,184
合計トークン数: 3,150,447,484
ファイル形式: JSONL
ファイルサイズ: 不明
加工内容
synthetic_sftを使用
トークン処理が重たいので、文字数でフィルター
seed_question < 6000
generation < 80000
thinkタグ除去 が中途半端なものを除外
トークナイズ処理(速度向上アップデート
繰り返し除去
rStar-Critique-Data
rStar-Critique-Data
This repository contains rStar-Critique-Data, the dataset used for the paper Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning.
This dataset is integral to the Critique Reinforcement Learning (CRL) paradigm, which enhances coder models by explicitly training them to generate critiques for (question, solution) pairs, as described in the accompanying paper.
Data Construction Pipeline is shown:
Paper
Critique-Coder:… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/rStar-Critique-Data.RStestStage1-no-rstellerrStar-Coder-PyStructure-refined
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): en
License: cc-by-4.0
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More… See the full description on the dataset page: https://huggingface.co/datasets/broadfield-dev/rStar-Coder-PyStructure-refined.rstar_ppm
Dataset for rStar-Math
https://github.com/microsoft/rStar
RST-DPO-Qwen3.5-27B
RST DPO preference pairs for Qwen3.5
2,673 preference pairs (2,448 train / 225 holdout) built from
Zhongzhi1228/Recursive-Task-Synthesis-Trajectories.
Each pair is two agent runs on the same task: one whose trajectory the task's own
verifier scored reward 1, one it scored 0.
Builder, trainer, and the numerical gates: https://github.com/k1ssloo/RST-Train
(scripts/17_build_dpo_data.py, scripts/19_train_dpo.py, DPO_PLAN.md).
What the preference actually encodes… See the full description on the dataset page: https://huggingface.co/datasets/NiuNiu0110/RST-DPO-Qwen3.5-27B.RStest2RSTD
RSTD — Russian Semantic Triplets Dataset
Русскоязычный корпус для оценки diversity-ранжирования при семантической
многозначности: проверяет способность поисковой системы покрыть в выдаче
разные смыслы полисемичного слова и не поддаться на отвлекающие документы.
Методология
Построен по методологии RUSSE'2018 (Panchenko et al., 2018) — первого
shared task по разрешению лексической многозначности для русского языка:
значения полисемичных слов — по словарному… See the full description on the dataset page: https://huggingface.co/datasets/DmitriiSablin/RSTD.rs-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "custom_eef",
"total_episodes": 2,
"total_frames": 742,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ywxia/rs-test.rs_tThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch_follower",
"total_episodes": 2,
"total_frames": 589,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ethanCSL/rs_t.
