datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dolci-Instruct-SFT-translatedDolci-Think-SFT-translated
Dolci-Think-SFT-translated
Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations.
Columns
Each row is a translated conversation plus the result of a post-translation quality filter:
id — source record id.
messages — the translated conversation (list of {content, role}).
filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.Dolci-Instruct-DPO-translatednemotron-cc-10K-sample-translated
Translated Nemotron-cc-hq samples
This dataset contains translated samples from https://huggingface.co/datasets/spyysalo/nemotron-cc-10K-sample
Currently, the following are available, we will add other models and languages:
Model
Languages
Gemma-3-4b-it
["Bulgarian", "Czech", "Danish", "German", "Estonian", "Finnish", "French", "Croatian", "Dutch"]
EuroLLM-9B-Instruct
["Bulgarian", "Czech", "Danish", "German", "Greek", "Estonia", "Finnish", "French", "Irish"… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/nemotron-cc-10K-sample-translated.EU-Instruct-Synthetic
EU Instruct Synthetic
Synthetically generated instruction-following SFT data for 11 European
languages. Each example is a single-turn chat (messages: a user instruction
and an assistant response) with a language field.
This is the synthetic counterpart to
openeurollm/Dolci-Instruct-SFT-translated.
Languages and sizes
Code
Language
Examples
cs
Czech
150,129
de
German
135,786
el
Greek
138,048
es
Spanish
132,736
fr
French
119,497
it
Italian
136… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/EU-Instruct-Synthetic.reasoning-traces-multilingual
OpenEuroLLM Multilingual Mathematical Reasoning Traces — Two-Stage Pilot
Release status: private v0.2-pilot staging dataset. All published rows passed the
deterministic translation gates described below. This pilot has not yet completed a systematic
native-speaker audit or independent downstream-solver verification and is not a final production
training release.
This dataset contains 3,425 accepted translations sampled from
100 mathematical reasoning traces into 37 non-English… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/reasoning-traces-multilingual.oellm-code-rlvr
OpenEuroLLM Code RLVR
oellm-code-rlvr is a deterministic corpus of 100,000 Python programming prompts for reinforcement learning with verifiable rewards. Every task uses standard input/output, includes two model-visible examples, and has 10–13 hidden tests in the Open R1 verification_info format.
The corpus is procedural and Apache-2.0 licensed. It does not copy Codeforces, LeetCode, LiveCodeBench, HumanEval, MBPP, APPS, or other benchmark text.
Design
The release… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/oellm-code-rlvr.openeurollm-model-identity
OpenEuroLLM Model Identity
A multilingual synthetic conversation dataset for teaching OpenEuroLLM checkpoints accurate,
bounded self-knowledge. It follows Andrej Karpathy's nanochat identity-data pattern—describe the
desired identity, generate varied User/Assistant conversations, mix them into post-training, and
evaluate whether the behavior emerged—but extends the target from a simple persona to a structured
model self-knowledge curriculum.
Version 1.0.0 contains 1,000… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/openeurollm-model-identity.oellm-eu-tooluse-v1
oellm-eu-tooluse-v1
Function-calling / agentic post-training data, normalized to Qwen3.5's native tool-call format
(<tools>…</tools> in the system turn, <tool_call>{json}</tool_call> from the assistant). Built
for the OpenEuroLLM European post-training of Qwen3.5 (folded into the Qwen3.5-4B-EU "v-next"
mobile model as ~10% of the SFT mix, plus a verifiable RL stage).
The value here is format unification: three popular tool-use sources each encode calls
differently (Hermes JSON… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/oellm-eu-tooluse-v1.openeurollm-model-identity
OpenEuroLLM Model Identity
A multilingual synthetic conversation dataset for teaching OpenEuroLLM checkpoints accurate,
bounded self-knowledge. It follows Andrej Karpathy's nanochat identity-data pattern—describe the
desired identity, generate varied User/Assistant conversations, mix them into post-training, and
evaluate whether the behavior emerged—but extends the target from a simple persona to a structured
model self-knowledge curriculum.
Version 1.0.0 contains 1,000… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/openeurollm-model-identity.
