hrm_text
Datasets
All datasets matching “hrm_text”HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts.
Citation
If you find this project or our paper useful, please consider citing our paper:
@misc{wang2026hrmtextefficientpretrainingscaling,
title={HRM-Text: Efficient Pretraining Beyond Scaling},
author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.hrm_text_sampledhrm-text-code-tools-sft
HRM-Text Code & Tools SFT
Curated coding and fixed-runtime tool-use supervised fine-tuning data for
HRM-Text-1B. This sealed
research release is not a general chat corpus. It contains only canonical
v2 records—there are no legacy prefix-field files.
Contents
Stage
Rows
Train
Validation
Serialized cap
Train SHA-256
stage-a
1,157,064
1,133,817
23,247
4,096
b6e340ea64568570b24e2d5f08c506a6f3221482d9fca6778a49c5577832171c
stage-b
283,547
278,018
5,529
4… See the full description on the dataset page: https://huggingface.co/datasets/pzarzycki/hrm-text-code-tools-sft.HRM-Text-Arrowhrm-text-opus46-math-coding
YL95/hrm-text-opus46-math-coding
This dataset keeps only Opus 4.6 math, coding, and nearby technical reasoning tasks from the requested source datasets.
Contents
prompt_completion/train: the main training split for base-model fine-tuning
prompt_completion/over_4096_tokens: rows longer than the token limit
chat/train: a message-form version of the same kept rows
chat/over_4096_tokens: the message-form over-limit subset
Notes
HRM-Text-1B is a base… See the full description on the dataset page: https://huggingface.co/datasets/YL95/hrm-text-opus46-math-coding.HRM-Text-MATH-DAPO
