datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
amalia-Dolci-Instruct-SFT
AMALIA Dolci-Instruct-SFT
Version of the allenai/Dolci-Instruct-SFT dataset used in the ramp down phase of AMALIA's Supervised Fine-Tuning stage.
This dataset was developed by sampling the highest quality entries of the selected splits, and translating part of those entries to European Portuguese. Both the quality classification and translation were done using google/gemma-4-31B-it.
This dataset went through a processing pipeline to:
Remove entries that reference… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Dolci-Instruct-SFT.Dolci-Instruct-SFT-enPurified-openai-messages
enPurified: Dolci-Instruct-SFT
The original dataset https://huggingface.co/datasets/allenai/Dolci-Instruct-SFT was reduced from ~2,155,000 rows to 38,829 of English only prose.
Project Overview
The enPurified collection is an initiative to curate high-fidelity English prose datasets for language modeling. While the open-source ecosystem is rich with datasets targeting mathematics, code generation, and multilingual capabilities, there is a distinct need for corpora focused… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/Dolci-Instruct-SFT-enPurified-openai-messages.Dolci-Think-SFT-7B-Propella-AnnotationsDolci-Think-SFT-7B-translationsallenai-Dolci-Thinkamalia-Dolci-Instruct-SFT-No-Tools
AMALIA Dolci-Instruct-SFT-No-Tools
Version of the allenai/Dolci-Instruct-SFT-No-Tools dataset used in the AMALIA's Supervised Fine-Tuning stage.
This dataset went through a processing pipeline to:
Remove entries where the source field was 'allenai/olmo-3-instruct-tagged-wildchat-only-topic-filtered' or 'allenai/hardcoded-olmo';
Remove entries that reference other LLMs or research labs;
Original Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Dolci-Instruct-SFT-No-Tools.dolci-coding-sftDolci-Instruct-SFT-Tool-Use-Codemode
Dolci Instruct SFT Tool Use – Codemode Augmentation
This dataset is a transformed version of allenai/Dolci-Instruct-SFT-Tool-Use.
It preserves the original conversation content and metadata, but rewrites tool calls into executable JavaScript <codemode> blocks plus structured environment outputs.
Only the final codemode-augmented dataset is published here; the original data remains available from the AllenAI dataset above.
Source and Attribution
Source dataset:… See the full description on the dataset page: https://huggingface.co/datasets/deathbyknowledge/Dolci-Instruct-SFT-Tool-Use-Codemode.dolci-think-sft-7b-de27b-part-2-carlesoctav-Dolci-Instruct-SFT-No-Tools-instruct-onlydolci-math-sftdolci-safety-sftdolci-4-lang
Dolci 4-Language (Machine-Translated)
Dataset Summary
This dataset is a 4-language machine-translated derivative of the AllenAI Dolci dataset collection, used in the OLMo 3 post-training pipeline.
It contains translations of reasoning, instruction-following, and alignment data into German, French, Italian, and Spanish for research on multilingual reasoning and alignment.
Translations were generated automatically using Qwen/Qwen3-Next-80B-A3B-Instruct and may… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/dolci-4-lang.dolci-think-dpo-7b-4-langDolci-Think-SFT-32B-q35instructdolci-vi-5kdolci-100k-sharegpt
DOLCI 100K ShareGPT Format
这是 DOLCI 数据集的前 100,000 条样本,已转换为 ShareGPT 格式。
数据集统计
总样本数: 100,000
单轮对话: 70,383 (70.4%)
多轮对话: 29,617 (29.6%)
数据格式
数据采用 ShareGPT 格式,每个样本包含:
{
"conversations": [
{
"from": "human",
"value": "用户消息"
},
{
"from": "gpt",
"value": "助手回复"
}
],
"system": "系统提示词(可选)"
}
使用方法
from datasets import load_dataset
dataset = load_dataset("dongxx1104/dolci-100k-sharegpt")
许可证
MIT… See the full description on the dataset page: https://huggingface.co/datasets/dongxx1104/dolci-100k-sharegpt.dolci-4-lang
Dolci 4-Language (Machine-Translated)
Dataset Summary
This dataset is a 4-language machine-translated derivative of the AllenAI Dolci dataset collection, used in the OLMo 3 post-training pipeline.
It contains translations of reasoning, instruction-following, and alignment data into German, French, Italian, and Spanish for research on multilingual reasoning and alignment.
Translations were generated automatically using Qwen/Qwen3-Next-80B-A3B-Instruct and may contain… See the full description on the dataset page: https://huggingface.co/datasets/ctauchmann/dolci-4-lang.dolci-wildchat-think-singleturndolci-multilingual-sftdolci-wildchat-think-singleturn-filtereddfm12-dolci-nl
dfm12-dolci-nl
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target indices are preserved. OPUS… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dolci-nl.dfm12-dolci-pl
dfm12-dolci-pl
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target indices are preserved. OPUS… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dolci-pl.dfm12-dolci-sv
dfm12-dolci-sv
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target indices are preserved. OPUS… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dolci-sv.
