CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01openeurollm /propella-annotations This dataset contains document annotations produced with propella-1-4b, a small multilingual LLM that annotates text documents across six categories: core content, classification, quality & value, audience & purpose, safety & compliance, and geographic relevance. The annotations can be used to filter, select, and curate LLM training data at scale. Properties Each document is annotated across 18 properties organized into six categories: Category Property Description… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/propella-annotations.text1B<n<10B20 likes7.4k downloads1mo agoHugging Face02openeurollm /Dolci-Instruct-SFT-translatedtexttext-generation1M<n<10M3 likes1.5k downloads3mo agoHugging Face03openeurollm /smoltalk2-decontaminated Decontamination This dataset is a decontaminated version of HuggingFaceTB/smoltalk2. Benchmarks used MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test) AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train) AIME25: math-ai/aime25 (subset=default, split=test) AMC23: math-ai/amc23 (subset=default, split=test) JEEBench: daman1209arora/jeebench (subset=default, split=test) GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train) LiveCodeBench:… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/smoltalk2-decontaminated.text1M<n<10M0 likes1.2k downloads6mo agoHugging Face04openeurollm /Nemotron-Post-Training-Dataset-v2-decontaminated Decontamination This dataset is a decontaminated version of nvidia/Nemotron-Post-Training-Dataset-v2. Benchmarks used MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test) AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train) AIME25: math-ai/aime25 (subset=default, split=test) AMC23: math-ai/amc23 (subset=default, split=test) JEEBench: daman1209arora/jeebench (subset=default, split=test) GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Nemotron-Post-Training-Dataset-v2-decontaminated.text1M<n<10M1 likes1.1k downloads6mo agoHugging Face05openeurollm /Dolci-Think-SFT-translated Dolci-Think-SFT-translated Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations. Columns Each row is a translated conversation plus the result of a post-translation quality filter: id — source record id. messages — the translated conversation (list of {content, role}). filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.tabulartext-generation1M<n<10M0 likes704 downloads10d agoHugging Face06openeurollm /Dolci-Think-SFT-7B-decontaminated Decontamination This dataset is a decontaminated version of allenai/Dolci-Think-SFT-7B. Benchmarks used MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test) AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train) AIME25: math-ai/aime25 (subset=default, split=test) AMC23: math-ai/amc23 (subset=default, split=test) JEEBench: daman1209arora/jeebench (subset=default, split=test) GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-7B-decontaminated.text1M<n<10M0 likes507 downloads6mo agoHugging Face07openeurollm /Dolci-Think-SFT-32B-decontaminated Decontamination This dataset is a decontaminated version of allenai/Dolci-Think-SFT-32B. Benchmarks used MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test) AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train) AIME25: math-ai/aime25 (subset=default, split=test) AMC23: math-ai/amc23 (subset=default, split=test) JEEBench: daman1209arora/jeebench (subset=default, split=test) GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-32B-decontaminated.text1M<n<10M1 likes455 downloads6mo agoHugging Face08openeurollm /nemotron-cc-10K-sample-translated Translated Nemotron-cc-hq samples This dataset contains translated samples from https://huggingface.co/datasets/spyysalo/nemotron-cc-10K-sample Currently, the following are available, we will add other models and languages: Model Languages Gemma-3-4b-it ["Bulgarian", "Czech", "Danish", "German", "Estonian", "Finnish", "French", "Croatian", "Dutch"] EuroLLM-9B-Instruct ["Bulgarian", "Czech", "Danish", "German", "Greek", "Estonia", "Finnish", "French", "Irish"… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/nemotron-cc-10K-sample-translated.texttext-generation100K<n<1M1 likes174 downloads1y agoHugging Face09openeurollm /Dolci-Instruct-SFT-decontaminated Decontamination This dataset is a decontaminated version of allenai/Dolci-Instruct-SFT. Benchmarks used MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test) AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train) AIME25: math-ai/aime25 (subset=default, split=test) AMC23: math-ai/amc23 (subset=default, split=test) JEEBench: daman1209arora/jeebench (subset=default, split=test) GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Instruct-SFT-decontaminated.textother1M<n<10M0 likes174 downloads6mo agoHugging Face10openeurollm /prelude-base-eval-scorestabular10K<n<100K0 likes154 downloads29d agoHugging Face11openeurollm /orca-agentinstruct-1M-v1-decontaminated Decontamination This dataset is a decontaminated version of microsoft/orca-agentinstruct-1M-v1. Benchmarks used MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test) AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train) AIME25: math-ai/aime25 (subset=default, split=test) AMC23: math-ai/amc23 (subset=default, split=test) JEEBench: daman1209arora/jeebench (subset=default, split=test) GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/orca-agentinstruct-1M-v1-decontaminated.textquestion-answering1M<n<10M0 likes136 downloads6mo agoHugging Face12openeurollm /EU-Instruct-Synthetic EU Instruct Synthetic Synthetically generated instruction-following SFT data for 11 European languages. Each example is a single-turn chat (messages: a user instruction and an assistant response) with a language field. This is the synthetic counterpart to openeurollm/Dolci-Instruct-SFT-translated. Languages and sizes Code Language Examples cs Czech 150,129 de German 135,786 el Greek 138,048 es Spanish 132,736 fr French 119,497 it Italian 136… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/EU-Instruct-Synthetic.texttext-generation1M<n<10M3 likes130 downloads3mo agoHugging Face13openeurollm /nemotron-cc-10K-sample-translated-judgedtabular1M<n<10M0 likes112 downloads1y agoHugging Face14openeurollm /ArenaHard-EU-v0 ArenaHard-EU Dataset Card Dataset Description ArenaHard-EU is a comprehensive multilingual benchmark for evaluating Large Language Models (LLMs) across 35 European and neighboring languages. This dataset extends the original Arena-Hard benchmark through machine translation, enabling robust multilingual LLM evaluation. Key Features 35 Languages: Covers all official EU languages plus co-official languages, candidate member languages, and Scandinavian languages… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/ArenaHard-EU-v0.textn<1K0 likes92 downloads11mo agoHugging Face15openeurollm /reasoning-traces-multilingual OpenEuroLLM Multilingual Mathematical Reasoning Traces — Two-Stage Pilot Release status: private v0.2-pilot staging dataset. All published rows passed the deterministic translation gates described below. This pilot has not yet completed a systematic native-speaker audit or independent downstream-solver verification and is not a final production training release. This dataset contains 3,425 accepted translations sampled from 100 mathematical reasoning traces into 37 non-English… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/reasoning-traces-multilingual.texttext-generation1K<n<10K1 likes60 downloads1mo agoHugging Face16openeurollm /open-perfectblend-decontaminated Decontamination This dataset is a decontaminated version of mlabonne/open-perfectblend. Benchmarks used MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test) AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train) AIME25: math-ai/aime25 (subset=default, split=test) AMC23: math-ai/amc23 (subset=default, split=test) JEEBench: daman1209arora/jeebench (subset=default, split=test) GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/open-perfectblend-decontaminated.text1M<n<10M1 likes49 downloads6mo agoHugging Face17openeurollm /contaminated-documentsThis repository will include the contaminated documents from Nemotron and HPLT, extracted using nemo-curator. The benchmarks are obtained from here, and use the split defined for benchmarking by lm-evaluation-harness text10K<n<100K0 likes38 downloads10mo agoHugging Face18openeurollm /lmsys-chat-1m-decontaminated Decontamination This dataset is a decontaminated version of lmsys/lmsys-chat-1m. Benchmarks used MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test) AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train) AIME25: math-ai/aime25 (subset=default, split=test) AMC23: math-ai/amc23 (subset=default, split=test) JEEBench: daman1209arora/jeebench (subset=default, split=test) GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train) LiveCodeBench:… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/lmsys-chat-1m-decontaminated.text100K<n<1M0 likes33 downloads6mo agoHugging Face19openeurollm /jeopardytext1K<n<10K0 likes33 downloads3mo agoHugging Face20openeurollm /eval_dashboardtextn<1K0 likes15 downloads5mo agoHugging Face21openeurollm /common-pile-annotatedtext10K<n<100K0 likes14 downloads5mo agoHugging Face22openeurollm /ArenaHard-EU-v0-bistextn<1K0 likes11 downloads11mo agoHugging Face23openeurollm /battle-annotationstextn<1K0 likes9 downloads8mo agoHugging Face24birgermoell /openeurollm-model-identity OpenEuroLLM Model Identity A multilingual synthetic conversation dataset for teaching OpenEuroLLM checkpoints accurate, bounded self-knowledge. It follows Andrej Karpathy's nanochat identity-data pattern—describe the desired identity, generate varied User/Assistant conversations, mix them into post-training, and evaluate whether the behavior emerged—but extends the target from a simple persona to a structured model self-knowledge curriculum. Version 1.0.0 contains 1,000… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/openeurollm-model-identity.texttext-generation1K<n<10K0 likes9 downloads1mo agoHugging Face25openeurollm /openeurollm-model-identity OpenEuroLLM Model Identity A multilingual synthetic conversation dataset for teaching OpenEuroLLM checkpoints accurate, bounded self-knowledge. It follows Andrej Karpathy's nanochat identity-data pattern—describe the desired identity, generate varied User/Assistant conversations, mix them into post-training, and evaluate whether the behavior emerged—but extends the target from a simple persona to a structured model self-knowledge curriculum. Version 1.0.0 contains 1,000… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/openeurollm-model-identity.texttext-generation1K<n<10K0 likes2h agoHugging Face26openeurollm /oellm-eu-tooluse-v1 oellm-eu-tooluse-v1 Function-calling / agentic post-training data, normalized to Qwen3.5's native tool-call format (<tools>…</tools> in the system turn, <tool_call>{json}</tool_call> from the assistant). Built for the OpenEuroLLM European post-training of Qwen3.5 (folded into the Qwen3.5-4B-EU "v-next" mobile model as ~10% of the SFT mix, plus a verifiable RL stage). The value here is format unification: three popular tool-use sources each encode calls differently (Hermes JSON… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/oellm-eu-tooluse-v1.texttext-generation10K<n<100K0 likes1h agoHugging Face27openeurollm /oellm-math-rlvr OpenEuroLLM Math RLVR One million deterministic, verifier-ready mathematical problems for reinforcement learning with verifiable rewards. The release contains a 760,000-row English depth pool and 10,000 aligned semantic problems rendered in all 24 official EU languages (240,000 rows). This is a prompt-and-answer rollout corpus, not a chain-of-thought corpus. Model inputs contain only the problem and output-format instruction. Reference answers and verifier contracts remain… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/oellm-math-rlvr.tabularquestion-answering1M<n<10M0 likes58m agoHugging Face28openeurollm /oellm-code-rlvr OpenEuroLLM Code RLVR oellm-code-rlvr is a deterministic corpus of 100,000 Python programming prompts for reinforcement learning with verifiable rewards. Every task uses standard input/output, includes two model-visible examples, and has 10–13 hidden tests in the Open R1 verification_info format. The corpus is procedural and Apache-2.0 licensed. It does not copy Codeforces, LeetCode, LiveCodeBench, HumanEval, MBPP, APPS, or other benchmark text. Design The release… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/oellm-code-rlvr.tabulartext-generation100K<n<1M0 likes58m agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.