rst
Datasets
All datasets matching “rst”rStar-Coder
rStar-Coder Dataset
Project GitHub | Paper
Dataset Description
rStar-Coder is a large-scale competitive code problem dataset containing 418K programming problems, 580K long-reasoning solutions, and rich test cases of varying difficulty levels. This dataset aims to enhance code reasoning capabilities in large language models, particularly in handling competitive code problems.
Experiments on Qwen models (1.5B-14B) across various code reasoning benchmarks demonstrate… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/rStar-Coder.RSTeller
⚠️ Usage Warning
This is the latest version of RSTeller, updated on 2025-01-28. Users who accessed this dataset before this date can find the legacy version, which is preserved for reference. Additionally, we have released the metadata for this dataset.
For the details and the usage of the dataset, please refer to our github repository page.
Citation
If you find the dataset and our paper useful, please consider citing our paper:
@article{ge2025rsteller… See the full description on the dataset page: https://huggingface.co/datasets/SlytherinGe/RSTeller.R-Star-Distillation-BackupsrStar-Coder-seed-testusaco_2025
USACO 2025 Open Contest Dataset
Dataset Description
The USA Computing Olympiad (USACO) is a prestigious algorithmic programming competition for high school students in the United States, consisting of four difficulty levels: Bronze, Silver, Gold, and Platinum. Each level contains a set of challenging problems that test algorithmic thinking and implementation skills, making USACO a valuable benchmark for evaluating the reasoning and problem-solving capabilities of large… See the full description on the dataset page: https://huggingface.co/datasets/rStar-Reasoning/usaco_2025.RSTeller_legacy
⛔ Usage Warning
This is the legacy version of the RSTeller dataset and is not the latest version referenced in our paper. We are keeping it available here to provide the community with easy access to additional data.
For the details and the usage of the dataset, please refer to our github page.
Citation
If you find the dataset and our paper useful, please consider citing our paper:
@article{ge2025rsteller,
title={RSTeller: Scaling up visual language modeling in… See the full description on the dataset page: https://huggingface.co/datasets/SlytherinGe/RSTeller_legacy.
