schneiderkamplab/dfm10-sapient-dmmath-filtered-sft
dfm10-sapient-dmmath-filtered-sft The DFM10-safe policy-selected DeepMind Mathematics partition from the Sapient source mirror. Contents Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz Schema: chat messages, optional condition and tools, plus provenance Shards: 448 Rows: 111,999,888 Category: Math reasoning Upstream material sapientinc/HRM-Text-data-io-cleaned-20260515 DeepMind Mathematics Processing Only files… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-sapient-dmmath-filtered-sft.
dfm10-sapient-dmmath-filtered-sft
The DFM10-safe policy-selected DeepMind Mathematics partition from the Sapient source mirror.
Contents
- Format: gzip-compressed JSON Lines under
data/train-*.jsonl.gz - Schema: chat
messages, optionalconditionandtools, plus provenance - Shards: 448
- Rows: 111,999,888
- Category: Math reasoning
Upstream material
sapientinc/HRM-Text-data-io-cleaned-20260515DeepMind Mathematics
Processing
Only files present in the active filtered Sapient symlink tree are packaged. The partition boundary follows the original provenance family and retains the training-visible condition, instruction, and response fields.
Selection policy: config/data/source_filter.yaml.
Every packaged row is taken from the accepted source tree identified in the package manifest. Tokenized arrays and epoch sampling indices are not included; export staging alone does not imply inclusion in a sampled training union.
License and release review
This package does not replace or broaden the licenses of its upstream materials. Review the dataset card, preserve upstream notices and attribution, and record the release decision before upload. Perform a per-family provenance, attribution, licence, privacy, and release review before upload; inclusion in an academic TDM training run does not by itself establish redistribution permission.
Validate
python recreate_dataset.py