datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
propella-annotations
This dataset contains document annotations produced with propella-1-4b, a small multilingual LLM that annotates text documents across six categories: core content, classification, quality & value, audience & purpose, safety & compliance, and geographic relevance. The annotations can be used to filter, select, and curate LLM training data at scale.
Properties
Each document is annotated across 18 properties organized into six categories:
Category
Property
Description… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/propella-annotations.Dolci-Instruct-SFT-translatedsmoltalk2-decontaminated
Decontamination
This dataset is a decontaminated version of HuggingFaceTB/smoltalk2.
Benchmarks used
MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test)
AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train)
AIME25: math-ai/aime25 (subset=default, split=test)
AMC23: math-ai/amc23 (subset=default, split=test)
JEEBench: daman1209arora/jeebench (subset=default, split=test)
GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)
LiveCodeBench:… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/smoltalk2-decontaminated.Nemotron-Post-Training-Dataset-v2-decontaminated
Decontamination
This dataset is a decontaminated version of nvidia/Nemotron-Post-Training-Dataset-v2.
Benchmarks used
MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test)
AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train)
AIME25: math-ai/aime25 (subset=default, split=test)
AMC23: math-ai/amc23 (subset=default, split=test)
JEEBench: daman1209arora/jeebench (subset=default, split=test)
GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Nemotron-Post-Training-Dataset-v2-decontaminated.Dolci-Think-SFT-translated
Dolci-Think-SFT-translated
Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations.
Columns
Each row is a translated conversation plus the result of a post-translation quality filter:
id — source record id.
messages — the translated conversation (list of {content, role}).
filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.Dolci-Think-SFT-7B-decontaminated
Decontamination
This dataset is a decontaminated version of allenai/Dolci-Think-SFT-7B.
Benchmarks used
MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test)
AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train)
AIME25: math-ai/aime25 (subset=default, split=test)
AMC23: math-ai/amc23 (subset=default, split=test)
JEEBench: daman1209arora/jeebench (subset=default, split=test)
GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-7B-decontaminated.Dolci-Think-SFT-32B-decontaminated
Decontamination
This dataset is a decontaminated version of allenai/Dolci-Think-SFT-32B.
Benchmarks used
MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test)
AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train)
AIME25: math-ai/aime25 (subset=default, split=test)
AMC23: math-ai/amc23 (subset=default, split=test)
JEEBench: daman1209arora/jeebench (subset=default, split=test)
GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-32B-decontaminated.nemotron-cc-10K-sample-translated
Translated Nemotron-cc-hq samples
This dataset contains translated samples from https://huggingface.co/datasets/spyysalo/nemotron-cc-10K-sample
Currently, the following are available, we will add other models and languages:
Model
Languages
Gemma-3-4b-it
["Bulgarian", "Czech", "Danish", "German", "Estonian", "Finnish", "French", "Croatian", "Dutch"]
EuroLLM-9B-Instruct
["Bulgarian", "Czech", "Danish", "German", "Greek", "Estonia", "Finnish", "French", "Irish"… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/nemotron-cc-10K-sample-translated.Dolci-Instruct-SFT-decontaminated
Decontamination
This dataset is a decontaminated version of allenai/Dolci-Instruct-SFT.
Benchmarks used
MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test)
AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train)
AIME25: math-ai/aime25 (subset=default, split=test)
AMC23: math-ai/amc23 (subset=default, split=test)
JEEBench: daman1209arora/jeebench (subset=default, split=test)
GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Instruct-SFT-decontaminated.prelude-base-eval-scoresorca-agentinstruct-1M-v1-decontaminated
Decontamination
This dataset is a decontaminated version of microsoft/orca-agentinstruct-1M-v1.
Benchmarks used
MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test)
AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train)
AIME25: math-ai/aime25 (subset=default, split=test)
AMC23: math-ai/amc23 (subset=default, split=test)
JEEBench: daman1209arora/jeebench (subset=default, split=test)
GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/orca-agentinstruct-1M-v1-decontaminated.EU-Instruct-Synthetic
EU Instruct Synthetic
Synthetically generated instruction-following SFT data for 11 European
languages. Each example is a single-turn chat (messages: a user instruction
and an assistant response) with a language field.
This is the synthetic counterpart to
openeurollm/Dolci-Instruct-SFT-translated.
Languages and sizes
Code
Language
Examples
cs
Czech
150,129
de
German
135,786
el
Greek
138,048
es
Spanish
132,736
fr
French
119,497
it
Italian
136… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/EU-Instruct-Synthetic.nemotron-cc-10K-sample-translated-judgedArenaHard-EU-v0
ArenaHard-EU Dataset Card
Dataset Description
ArenaHard-EU is a comprehensive multilingual benchmark for evaluating Large Language Models (LLMs) across 35 European and neighboring languages. This dataset extends the original Arena-Hard benchmark through machine translation, enabling robust multilingual LLM evaluation.
Key Features
35 Languages: Covers all official EU languages plus co-official languages, candidate member languages, and Scandinavian languages… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/ArenaHard-EU-v0.reasoning-traces-multilingual
OpenEuroLLM Multilingual Mathematical Reasoning Traces — Two-Stage Pilot
Release status: private v0.2-pilot staging dataset. All published rows passed the
deterministic translation gates described below. This pilot has not yet completed a systematic
native-speaker audit or independent downstream-solver verification and is not a final production
training release.
This dataset contains 3,425 accepted translations sampled from
100 mathematical reasoning traces into 37 non-English… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/reasoning-traces-multilingual.open-perfectblend-decontaminated
Decontamination
This dataset is a decontaminated version of mlabonne/open-perfectblend.
Benchmarks used
MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test)
AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train)
AIME25: math-ai/aime25 (subset=default, split=test)
AMC23: math-ai/amc23 (subset=default, split=test)
JEEBench: daman1209arora/jeebench (subset=default, split=test)
GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/open-perfectblend-decontaminated.contaminated-documentsThis repository will include the contaminated documents from Nemotron and HPLT, extracted using nemo-curator.
The benchmarks are obtained from here, and use the split defined for benchmarking by lm-evaluation-harness
lmsys-chat-1m-decontaminated
Decontamination
This dataset is a decontaminated version of lmsys/lmsys-chat-1m.
Benchmarks used
MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test)
AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train)
AIME25: math-ai/aime25 (subset=default, split=test)
AMC23: math-ai/amc23 (subset=default, split=test)
JEEBench: daman1209arora/jeebench (subset=default, split=test)
GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)
LiveCodeBench:… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/lmsys-chat-1m-decontaminated.jeopardyeval_dashboardcommon-pile-annotatedArenaHard-EU-v0-bisbattle-annotationsopeneurollm-model-identity
OpenEuroLLM Model Identity
A multilingual synthetic conversation dataset for teaching OpenEuroLLM checkpoints accurate,
bounded self-knowledge. It follows Andrej Karpathy's nanochat identity-data pattern—describe the
desired identity, generate varied User/Assistant conversations, mix them into post-training, and
evaluate whether the behavior emerged—but extends the target from a simple persona to a structured
model self-knowledge curriculum.
Version 1.0.0 contains 1,000… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/openeurollm-model-identity.openeurollm-model-identity
OpenEuroLLM Model Identity
A multilingual synthetic conversation dataset for teaching OpenEuroLLM checkpoints accurate,
bounded self-knowledge. It follows Andrej Karpathy's nanochat identity-data pattern—describe the
desired identity, generate varied User/Assistant conversations, mix them into post-training, and
evaluate whether the behavior emerged—but extends the target from a simple persona to a structured
model self-knowledge curriculum.
Version 1.0.0 contains 1,000… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/openeurollm-model-identity.oellm-eu-tooluse-v1
oellm-eu-tooluse-v1
Function-calling / agentic post-training data, normalized to Qwen3.5's native tool-call format
(<tools>…</tools> in the system turn, <tool_call>{json}</tool_call> from the assistant). Built
for the OpenEuroLLM European post-training of Qwen3.5 (folded into the Qwen3.5-4B-EU "v-next"
mobile model as ~10% of the SFT mix, plus a verifiable RL stage).
The value here is format unification: three popular tool-use sources each encode calls
differently (Hermes JSON… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/oellm-eu-tooluse-v1.oellm-math-rlvr
OpenEuroLLM Math RLVR
One million deterministic, verifier-ready mathematical problems for reinforcement learning with
verifiable rewards. The release contains a 760,000-row English depth pool and 10,000 aligned semantic
problems rendered in all 24 official EU languages (240,000 rows).
This is a prompt-and-answer rollout corpus, not a chain-of-thought corpus. Model inputs contain only the
problem and output-format instruction. Reference answers and verifier contracts remain… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/oellm-math-rlvr.oellm-code-rlvr
OpenEuroLLM Code RLVR
oellm-code-rlvr is a deterministic corpus of 100,000 Python programming prompts for reinforcement learning with verifiable rewards. Every task uses standard input/output, includes two model-visible examples, and has 10–13 hidden tests in the Open R1 verification_info format.
The corpus is procedural and Apache-2.0 licensed. It does not copy Codeforces, LeetCode, LiveCodeBench, HumanEval, MBPP, APPS, or other benchmark text.
Design
The release… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/oellm-code-rlvr.
