CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01openeurollm /Dolci-Instruct-SFT-translatedtexttext-generation1M<n<10M3 likes1.5k downloads3mo agoHugging Face02openeurollm /Dolci-Think-SFT-translated Dolci-Think-SFT-translated Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations. Columns Each row is a translated conversation plus the result of a post-translation quality filter: id — source record id. messages — the translated conversation (list of {content, role}). filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.tabulartext-generation1M<n<10M0 likes847 downloads9h agoHugging Face03openeurollm /Dolci-Instruct-DPO-translatedtext-generation100K<n<1M1 likes541 downloads10d agoHugging Face04openeurollm /nemotron-cc-10K-sample-translated Translated Nemotron-cc-hq samples This dataset contains translated samples from https://huggingface.co/datasets/spyysalo/nemotron-cc-10K-sample Currently, the following are available, we will add other models and languages: Model Languages Gemma-3-4b-it ["Bulgarian", "Czech", "Danish", "German", "Estonian", "Finnish", "French", "Croatian", "Dutch"] EuroLLM-9B-Instruct ["Bulgarian", "Czech", "Danish", "German", "Greek", "Estonia", "Finnish", "French", "Irish"… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/nemotron-cc-10K-sample-translated.texttext-generation100K<n<1M1 likes174 downloads1y agoHugging Face05openeurollm /EU-Instruct-Synthetic EU Instruct Synthetic Synthetically generated instruction-following SFT data for 11 European languages. Each example is a single-turn chat (messages: a user instruction and an assistant response) with a language field. This is the synthetic counterpart to openeurollm/Dolci-Instruct-SFT-translated. Languages and sizes Code Language Examples cs Czech 150,129 de German 135,786 el Greek 138,048 es Spanish 132,736 fr French 119,497 it Italian 136… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/EU-Instruct-Synthetic.texttext-generation1M<n<10M3 likes130 downloads3mo agoHugging Face06openeurollm /reasoning-traces-multilingual OpenEuroLLM Multilingual Mathematical Reasoning Traces — Two-Stage Pilot Release status: private v0.2-pilot staging dataset. All published rows passed the deterministic translation gates described below. This pilot has not yet completed a systematic native-speaker audit or independent downstream-solver verification and is not a final production training release. This dataset contains 3,425 accepted translations sampled from 100 mathematical reasoning traces into 37 non-English… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/reasoning-traces-multilingual.texttext-generation1K<n<10K1 likes60 downloads1mo agoHugging Face07openeurollm /oellm-code-rlvr OpenEuroLLM Code RLVR oellm-code-rlvr is a deterministic corpus of 100,000 Python programming prompts for reinforcement learning with verifiable rewards. Every task uses standard input/output, includes two model-visible examples, and has 10–13 hidden tests in the Open R1 verification_info format. The corpus is procedural and Apache-2.0 licensed. It does not copy Codeforces, LeetCode, LiveCodeBench, HumanEval, MBPP, APPS, or other benchmark text. Design The release… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/oellm-code-rlvr.tabulartext-generation100K<n<1M0 likes28 downloads21h agoHugging Face08openeurollm /openeurollm-model-identity OpenEuroLLM Model Identity A multilingual synthetic conversation dataset for teaching OpenEuroLLM checkpoints accurate, bounded self-knowledge. It follows Andrej Karpathy's nanochat identity-data pattern—describe the desired identity, generate varied User/Assistant conversations, mix them into post-training, and evaluate whether the behavior emerged—but extends the target from a simple persona to a structured model self-knowledge curriculum. Version 1.0.0 contains 1,000… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/openeurollm-model-identity.texttext-generation1K<n<10K0 likes23 downloads22h agoHugging Face09openeurollm /oellm-eu-tooluse-v1 oellm-eu-tooluse-v1 Function-calling / agentic post-training data, normalized to Qwen3.5's native tool-call format (<tools>…</tools> in the system turn, <tool_call>{json}</tool_call> from the assistant). Built for the OpenEuroLLM European post-training of Qwen3.5 (folded into the Qwen3.5-4B-EU "v-next" mobile model as ~10% of the SFT mix, plus a verifiable RL stage). The value here is format unification: three popular tool-use sources each encode calls differently (Hermes JSON… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/oellm-eu-tooluse-v1.texttext-generation10K<n<100K0 likes13 downloads21h agoHugging Face10birgermoell /openeurollm-model-identity OpenEuroLLM Model Identity A multilingual synthetic conversation dataset for teaching OpenEuroLLM checkpoints accurate, bounded self-knowledge. It follows Andrej Karpathy's nanochat identity-data pattern—describe the desired identity, generate varied User/Assistant conversations, mix them into post-training, and evaluate whether the behavior emerged—but extends the target from a simple persona to a structured model self-knowledge curriculum. Version 1.0.0 contains 1,000… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/openeurollm-model-identity.texttext-generation1K<n<10K0 likes9 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.