CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CohereLabs /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Data source Prompts from AM-DeepSeek-R1-0528-Distilled Thinking traces and outputs distilled from gpt-oss-120b Translated with command-a-translate and DeepSeek-V3 Languages (44) Language Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M6 likes1.2k downloads13d agoHugging Face02eddie-OB /gsm8k-multilingual-reasoning gsm8k-multilingual-reasoning GSM8K with reasoning translated to multiple languages Schema {"prompt": "...", "answer": "...", "reasoning": "...", "metadata": {...}} Usage from datasets importload_dataset ds = load_dataset("eddie-OB/gsm8k-multilingual-reasoning") print(ds["train"][0]) Source Derived from OpenAI GSM8K. texttext-generationn<1K1 likes683 downloads8mo agoHugging Face03erenyeager-1 /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Languages (44) Language Train Test Total Amharic (am) 3,807 448 4,255 Arabic (ar) 22,968 2,538 25,506 Bulgarian (bg) 4,177 452 4,629 Bengali (bn) 3,803 422 4,225 Catalan (ca) 4,251 512 4,763 Czech… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M0 likes286 downloads14d agoHugging Face04deokhk /multilingual_reasoning_gap_outputs Dataset Card for multilingual_reasoning_gap_outputs Paper | Code Dataset Details Dataset Description This dataset contains experiment outputs for Qwen3-4B used in our study on multilingual reasoning gaps. It includes: Prober checkpoints trained for understanding-failure analysis Intermediate results, such as: Model inference outputs Signals for understanding failure detection Auxiliary artifacts used for probing and analysis The dataset is released to… See the full description on the dataset page: https://huggingface.co/datasets/deokhk/multilingual_reasoning_gap_outputs.text-generation0 likes254 downloads9mo agoHugging Face05MauroPello /multilingual-reasoning-gym-sft Reasoning Gym SFT Dataset This dataset contains Supervised Fine-Tuning (SFT) reasoning data procedurally generated using Reasoning Gym environments. It is designed to train reasoning models (such as DeepSeek-R1-style or Qwen-Coder-style models) to explain their step-by-step reasoning chain before outputting a final answer wrapped inside LaTeX \boxed{...}. Where Does This Dataset Come From? This dataset is procedurally generated from Reasoning Gym, an open-source… See the full description on the dataset page: https://huggingface.co/datasets/MauroPello/multilingual-reasoning-gym-sft.tabulartext-generation100K<n<1M1 likes152 downloads3mo agoHugging Face06openeurollm /reasoning-traces-multilingual OpenEuroLLM Multilingual Mathematical Reasoning Traces — Two-Stage Pilot Release status: private v0.2-pilot staging dataset. All published rows passed the deterministic translation gates described below. This pilot has not yet completed a systematic native-speaker audit or independent downstream-solver verification and is not a final production training release. This dataset contains 3,425 accepted translations sampled from 100 mathematical reasoning traces into 37 non-English… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/reasoning-traces-multilingual.texttext-generation1K<n<10K1 likes60 downloads1mo agoHugging Face07nassimjp /multilingual-proverb-reasoning 📖 Multilingual Proverb Reasoning Dataset This dataset contains 990 unique proverbs primarily in Japanese, with accompanying phonetic readings (Romaji), literal translations, and multilingual reasoning "Chain of Thought" (thought) fields. It is designed for training and evaluating LLMs on cultural nuance, metaphorical reasoning, and multilingual explanation tasks. 📊 Dataset Summary The dataset was curated and cleaned to remove 182 duplicates, resulting in a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/multilingual-proverb-reasoning.texttext-generationn<1K0 likes24 downloads4mo agoHugging Face08birgermoell /reasoning-traces-multilingual OpenEuroLLM Multilingual Mathematical Reasoning Traces — TranslateGemma Pilot Release status: private v0.1-pilot staging dataset. The rows passed the automated translation gates described below, but this pilot has not yet completed a systematic native-speaker audit or independent downstream-solver verification. It should not yet be treated as a final production training release. This dataset contains 2,930 accepted translations of mathematical reasoning traces into 37… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/reasoning-traces-multilingual.texttext-generation1K<n<10K0 likes21 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.