openeurollm/dolci-think-sft-tokenized
Dolci-Think-SFT Tokenized Pre-tokenized version of the allenai/Dolci-Think-SFT-7B dataset, ready for training with OLMo-core. This dataset was used to train the openeurollm/OLMo-3-7B-Think-SFT checkpoints. See also: openeurollm/dolci-instruct-sft-tokenized for the instruct (non-thinking) variant. Dataset Details Property Value Source dataset allenai/Dolci-Think-SFT-7B Tokenizer allenai/Olmo-3-7B-Think-SFT Max sequence length 32,768 Total instances… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/dolci-think-sft-tokenized.
Add cross-reference to instruct dataset
Clean up README formatting
Clarify trainable tokens and labels_mask in README
Upload tokenized Dolci-Think-SFT dataset
Update dataset README
Add dataset README
initial commit
