datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dolci-Instruct-SFT-translatedDolci-Think-SFT-32B-Multilingual
Dolci-Think-SFT-32B-Multilingual
Dolci-Think-SFT-32B-Multilingual is a large-scale multilingual long chain-of-thought (CoT) reasoning corpus spanning six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample includes a question, a long-form reasoning trace, and a final answer, all translated into the target language, with sequences up to 32,768 tokens.
It is released alongside the paper Rethinking the Multilingual Reasoning Gap with Layer Swap.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/Dolci-Think-SFT-32B-Multilingual.Dolci-Think-SFT-translated
Dolci-Think-SFT-translated
Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations.
Columns
Each row is a translated conversation plus the result of a post-translation quality filter:
id — source record id.
messages — the translated conversation (list of {content, role}).
filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.Dolci-Instruct-DPO-translatedDolci-Think-SFT-7B-q35instruct
Dolci-Think-SFT-7B-q35instruct
Megatron-format tokenization of
allenai/Dolci-Think-SFT-7B
using the Qwen3.5-0.8B-Instruct tokenizer and its chat template.
Contents
156 shards, each stored as two aligned Megatron indexed datasets (one document
per conversation, one sequence per document):
train-XXXXX-of-00156.bin / .idx — token ids (int32)
train-XXXXX-of-00156_loss_mask.bin / .idx — per-token loss mask (uint8, 0/1)
624 files total (~113 GB).… See the full description on the dataset page: https://huggingface.co/datasets/yangwang92/Dolci-Think-SFT-7B-q35instruct.dolci-distill-packed
Dolci Distill Packed
Pre-packed training data for knowledge distillation from OLMo-3-7B-Instruct to a pruned student model.
Description
This dataset contains tensorized and sequence-packed batches ready for distillation training. The data was preprocessed from hbfreed/Dolci-Instruct-RL-Completions to avoid preprocessing overhead during training.
Format
35 .pt files: 34 training chunks + 1 validation chunk
Pack length: 6144 tokens
~158GB total
Each .pt file… See the full description on the dataset page: https://huggingface.co/datasets/hbfreed/dolci-distill-packed.Dolci-Instruct-RL-Completions
Dolci Instruct RL Completions
Instruction-following completions sampled from OLMo-3-7B-Instruct with teacher logits for knowledge distillation.
Description
This dataset contains instruction-completion pairs with pre-computed teacher logits from OLMo-3-7B-Instruct. Designed for training smaller student models via KL-divergence distillation.
Generation
Completions were sampled from allenai/OLMo-3-7B-Instruct on instruction prompts. For each token position, we… See the full description on the dataset page: https://huggingface.co/datasets/hbfreed/Dolci-Instruct-RL-Completions.Dolci-Instruct-SFT-translated
Dolci-Instruct-SFT-translated (Swedish)
This dataset is a Swedish machine translation of the openeurollm/Dolci-Instruct-SFT-translated dataset, originally created as part of the OpenEuroLLM project.
Dataset details
Examples: 494,841 multi-turn conversations
Language: Swedish (sv-SE)
Format: Chat/messages format (id, messages)
License: Apache 2.0
Translation
All English source texts were machine-translated to Swedish using Google Gemma 3 27B-IT (w8a8_fp8… See the full description on the dataset page: https://huggingface.co/datasets/AI-Sweden-Models/Dolci-Instruct-SFT-translated.Dolci-Instruct-SFT-Tool-Use-Fixed
Dolci-Instruct-SFT-Tool-Use-Fixed
Dataset Description
Dolci-Instruct-SFT-Tool-Use-Fixed is a cleaned and re-formatted version of the allenai/Dolci-Instruct-SFT-Tool-Use tool-use dataset. It is designed as the tool-calling (function-calling) extension of the openbmb/UltraData-SFT-2605 Supervised Fine-Tuning dataset, so that tool-use samples can be mixed into UltraData-SFT-2605 training runs seamlessly.
The raw Dolci-Instruct-SFT-Tool-Use data uses a custom message… See the full description on the dataset page: https://huggingface.co/datasets/nekocyrene/Dolci-Instruct-SFT-Tool-Use-Fixed.Dolci-Instruct-SFT-enPurified-openai-messages
enPurified: Dolci-Instruct-SFT
The original dataset https://huggingface.co/datasets/allenai/Dolci-Instruct-SFT was reduced from ~2,155,000 rows to 38,829 of English only prose.
Project Overview
The enPurified collection is an initiative to curate high-fidelity English prose datasets for language modeling. While the open-source ecosystem is rich with datasets targeting mathematics, code generation, and multilingual capabilities, there is a distinct need for corpora focused… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/Dolci-Instruct-SFT-enPurified-openai-messages.dolci-think-sft-mini
Dolci Think SFT Mini
A compact reasoning dataset derived from allenai/Dolci-Think-SFT-32B. Each row contains id, source, question, steps, and answer; steps is a nonempty list of reasoning-step strings. Instances with fewer than 3 or more than 50 reasoning steps are excluded, and every step is whitespace-stripped. Usage remains subject to the source dataset's licensing terms.
Dataset statistics
Metric
Value
Final records
282,314
File size
1,033,908,072… See the full description on the dataset page: https://huggingface.co/datasets/cs-giung/dolci-think-sft-mini.tis-dolci-random-unbalanced
A Critical Look at Targeted Instruction Selection
This repository contains the pre-computed random unbalanced subsets used as baselines in the paper "A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't)".
Paper: https://huggingface.co/papers/2602.14696
GitHub Repository: https://github.com/dcml-lab/targeted-instruction-selection
Description
Instruction fine-tuning of large language models (LLMs) often involves… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-DCML/tis-dolci-random-unbalanced.dolci-instruct-sft-fi
Dolci Instruct SFT — Finnish (machine-translated)
Finnish machine translation of the Dolci Instruct SFT mixture
(dolci-instruct-sft-filtered-v1, no-math / no-latex subset), produced for
SFT of Finnish LLMs.
Translation model: translategemma-27b (Gemma-based 27B translation model)
Rows: 234,745 (multi-turn chat, mostly single Q→A)
Language: Finnish (fi)
Format: chat messages (role / content)
Provenance & filtering
Translated from the English… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/dolci-instruct-sft-fi.Dolci-Think-DPO-32B-FlatFlat version of AllenAI's Dolci-Think-DPO-32B.
Train set size: 199840
Valid set size: 160
MLX-LM-LoRA
mlx_lm_lora.train \
--model Goekdeniz-Guelmez/Josiefied-Qwen3-0.6B-abliterated-v1 \
--train \
--data mlx-community/Dolci-Think-DPO-32B-Flat \
--epochs 1 \
--batch-size 1 \
--num-layers 1 \
--val-batches 1 \
--steps-per-report 1 \
--adapter-path path/to/adapters \
--max-seq-length 1024 \
--grad-checkpoint \
--train-type lora \
--optimizer adamw \
--train-mode dpo \… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/Dolci-Think-DPO-32B-Flat.qwen3-1p7b-dolci-general-tracesGenerated Qwen3-1.7B responses for Dolci-Instruct-SFT general prompts.
This public export contains gzip-compressed JSONL processed shards only. A lightweight safety pass removed rows with obvious credential-like, RDP/IP/password, exploit-token, and sexual/NSFW indicators before upload. Raw generation logs are not included.
reasoning-sft-dolci-think-sft-32b-1M
Dolci-Think-SFT-32B (converted)
Converted version of allenai/Dolci-Think-SFT-32B, filtered to 1,015,233 rows from 7 selected sources.
Format
Each row has three columns:
input — list of dicts [{"role": "user", "content": "..."}, ...] (conversation turns ending on the last user turn)
response — teacher-generated response string (includes <think> reasoning block)
source — task domain / source dataset name
Filtering
Removed the following sources from the original… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-dolci-think-sft-32b-1M.dfm12-dolci-nl
dfm12-dolci-nl
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target indices are preserved. OPUS… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dolci-nl.dfm12-dolci-pl
dfm12-dolci-pl
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target indices are preserved. OPUS… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dolci-pl.dfm12-dolci-sv
dfm12-dolci-sv
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target indices are preserved. OPUS… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dolci-sv.
