datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts.
Citation
If you find this project or our paper useful, please consider citing our paper:
@misc{wang2026hrmtextefficientpretrainingscaling,
title={HRM-Text: Efficient Pretraining Beyond Scaling},
author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.HRM-Text-Arrowhrm-text-opus46-math-coding
YL95/hrm-text-opus46-math-coding
This dataset keeps only Opus 4.6 math, coding, and nearby technical reasoning tasks from the requested source datasets.
Contents
prompt_completion/train: the main training split for base-model fine-tuning
prompt_completion/over_4096_tokens: rows longer than the token limit
chat/train: a message-form version of the same kept rows
chat/over_4096_tokens: the message-form over-limit subset
Notes
HRM-Text-1B is a base… See the full description on the dataset page: https://huggingface.co/datasets/YL95/hrm-text-opus46-math-coding.HRM-Text-MATH-DAPOHRM-Text-data-io-cleaned-20260515-copyPre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts.
Citation
If you find this project or our paper useful, please consider citing our paper:
@misc{wang2026hrmtextefficientpretrainingscaling,
title={HRM-Text: Efficient Pretraining Beyond Scaling},
author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/huankguan2/HRM-Text-data-io-cleaned-20260515-copy.HRM-Text-MATH-SFTsamvaad-hi-hrm-textsamvaad-hi-hrm-text-only
