CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01miscovery /Math_CoT_Arabic_English_Reasoning Math CoT Arabic English Dataset A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI. Overview Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.tabularquestion-answering1K<n<10K17 likes73 downloads1y agoHugging Face02erayalp /easy_turkish_math_reasoning Easy Turkish Math Reasoning Dataset Summary The Easy Turkish Math Reasoning dataset is the first phase of a multi-stage curriculum learning pipeline designed to enhance the reasoning abilities of compact language models. This dataset focuses on elementary-level arithmetic and logic problems in Turkish, serving as a warm-up stage for supervised fine-tuning (SFT). Use Case Primarily used for: Bootstrapping reasoning ability in Turkish for compact LLMs. Phase 1… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/easy_turkish_math_reasoning.textquestion-answering1K<n<10K7 likes69 downloads1y agoHugging Face03erayalp /medium_turkish_math_reasoning Dataset Summary The Medium Turkish Math Reasoning dataset is Phase 2 of a curriculum learning pipeline to teach compact models multi-step reasoning in Turkish. It includes moderately difficult math problems involving multiple reasoning steps, such as two-part arithmetic, comparisons, and logical reasoning. Use Case This dataset is ideal for: Continuing SFT after foundational training with simpler problems. Bridging the gap between basic arithmetic and complex GSM8K-style… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/medium_turkish_math_reasoning.textquestion-answering1K<n<10K5 likes66 downloads1y agoHugging Face04veds12 /math-squared Dataset Name MATH2 Dataset Description MATH2 is a mathematical reasoning evaluation dataset curated using a human-in-the-loop approach proposed in the paper AI-Assisted Generation of Difficult Math Questions. The dataset consists of 210 questions formed by combining 2 math domain skills using frontier LLMs. These skills were extracted from the MATH [Hendrycks et al., 2021] dataset. Dataset Sources Paper: AI-Assisted Generation of Difficult Math… See the full description on the dataset page: https://huggingface.co/datasets/veds12/math-squared.textquestion-answeringn<1K6 likes60 downloads2y agoHugging Face05math-across-languages /gsm8k-translated Multilingual GSM8K Translations This dataset contains machine-translated versions of GSM8K in these languages: French (fr) German (de) Hindi (hi) Dataset Structure For each language, we provide the original GSM8K train and test splits: train: 7,473 samples test: 1,319 samples Each sample consists of a question and an answer. The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.textquestion-answering10K<n<100K0 likes60 downloads3mo agoHugging Face06Asib27 /dart_math_banglaThe dataset contains math problems in bangla. hkust-nlp/dart-math-uniform is translated using facebook/nllb-200-3.3B. To achive better performance english sentences are splitted and then fed into the translation model. textquestion-answering1K<n<10K0 likes58 downloads2y agoHugging Face07TaengooTV /math_onettabularquestion-answeringn<1K1 likes52 downloads1y agoHugging Face08prithivMLmods /Math-IIO-68K-Mini Mathematics Dataset for AI Model Training This dataset contains 68,000 rows of mathematical questions and their corresponding solutions. It is designed for training AI models capable of solving mathematical problems or providing step-by-step explanations for a variety of mathematical concepts. The dataset is structured into three columns: input, instruction, and output. Dataset Overview Input: A mathematical question or problem statement (e.g., arithmetic, algebra… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-IIO-68K-Mini.texttext-generation10K<n<100K5 likes48 downloads2y agoHugging Face09flamiinngo /math-code-qa Math & Code QA — Instruction Dataset Worked mathematical solutions and short code answers, built for the Adaption Labs AutoScientist Challenge (Math & Code category). Rows 5,200 Math 3,600 Code 1,600 Distinct answers 5,199 (100%) Duplicate questions none Nulls none Question length median 27 words Answer length median 58 words (max 89) License CC-BY-4.0 What makes the math rows unusual Every math answer is short worked reasoning… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa.textquestion-answering1K<n<10K1 likes35 downloads2mo agoHugging Face10NLPForUA /dumy-zno-ukrainian-math-history-geo-r1-o1 DUMY («Думи»): Ukrainian Multidomain Reasoning Dataset (Part 1: ZNO/NMT tasks with DeepSeek R1 and OpenAI o1 answers) DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks. The word “Dumy” comes from Taras Shevchenko’s famous poem and literally means “thoughts” in Ukrainian: Думи мої, думи мої, Лихо мені з вами! Нащо стали на папері Сумними рядами?.. Work in progress. Stay tuned.… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1.tabulartext-generation1K<n<10K2 likes29 downloads1y agoHugging Face11bilalabic /math-toolcall-tr-benchmark math-toolcall-tr-benchmark bilalabic/gemma_4_math-toolcall-tr_lora LoRA adaptörünü temel Gemma-4 E4B modeliyle karşılaştıran benchmark sonuçları. Bu depo yalnızca değerlendirme çıktılarını içerir. Eğitim veri seti ayrı olarak bilalabic/math-toolcall-tr adresinde yayımlanmaktadır. Benchmark'lar Benchmark Örnek Ölçülen davranış Türkçe MMLU 250 Genel bilgi doğruluğu ve eğitim sonrası bilgi kaybı Matematik Tool-Call 150 Araç seçimi, çekimserlik ve çıktı… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/math-toolcall-tr-benchmark.tabulartext-generation1K<n<10K0 likes29 downloads2mo agoHugging Face12flamiinngo /math-code-qa-v2 Math & Code QA v2 — Instruction Dataset Worked mathematical solutions and short code answers, spanning arithmetic word problems through to algebra, geometry and combinatorics. Built for the Adaption Labs AutoScientist Challenge (Math & Code category). The model trained on this beats Llama-3.3-70B-Instruct 72 to 28 on the held-out Math category evaluation. Rows 5,297 (4,197 math, 1,100 code) Distinct answers 5,297 (100%) Duplicate questions none Nulls none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa-v2.textquestion-answering1K<n<10K0 likes28 downloads2mo agoHugging Face13LTS-VVE /Math-physics-dataset-sq Physics and Math Problems Dataset This repository contains a dataset of 2,600 physics and math problems in Albanian. The dataset is designed to support various NLP tasks and educational applications. Dataset Overview Problems: 2,666 Language: Albanian Fields: algebra lineare: 178 analiza matematike: 176 gjeometri diferenciale: 178 topologji: 179 teoria e numrave: 177 ekuacionet diferenciale: 178 fizika klasike: 177 mekanika kuantike: 179 elektromagnetizmi: 178… See the full description on the dataset page: https://huggingface.co/datasets/LTS-VVE/Math-physics-dataset-sq.textquestion-answering1K<n<10K0 likes27 downloads1y agoHugging Face14prithivMLmods /Math-IIO-Mini Mathematics Dataset for AI Model Training This dataset contains 500 rows of mathematical questions and their corresponding solutions. It is designed for training AI models capable of solving mathematical problems or providing step-by-step explanations for a variety of mathematical concepts. The dataset is structured into three columns: input, instruction, and output. Dataset Overview Input: A mathematical question or problem statement (e.g., arithmetic, algebra… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-IIO-Mini.textquestion-answeringn<1K1 likes25 downloads2y agoHugging Face15prithivMLmods /Math-Solve Overview The Math-Solve dataset is a collection of math problems and their solutions, designed to facilitate training and evaluation of models for tasks such as text generation, question answering, and summarization. The dataset contains nearly 25k rows of math-related problems, each paired with a detailed solution. This dataset is particularly useful for researchers and developers working on AI models that require mathematical reasoning and problem-solving capabilities.… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-Solve.texttext-generation10K<n<100K16 likes25 downloads2y agoHugging Face16Jr23xd23 /Arabic_LLaMA_Math_Dataset Arabic LLaMA Math Dataset Example Entries Dataset Overview Dataset Name: Arabic_LLaMA_Math_Dataset.csv Number of Records: 12,496 Number of Columns: 3 File Format: CSV Dataset Structure Columns: Instruction: The problem statement or question (text, in Arabic) Input: Additional input for model fine-tuning (empty in this dataset) Solution: The solution or answer to the problem (text, in Arabic) Dataset Description The Arabic… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/Arabic_LLaMA_Math_Dataset.textquestion-answering10K<n<100K4 likes24 downloads2y agoHugging Face17Cheukting /mathqaOrigianl dataset from allenai/math_qa textquestion-answering10K<n<100K0 likes23 downloads11mo agoHugging Face18DrUkachi /ktt-math-tutor-data KTT Math Tutor — Data Data artefacts for the AIMS KTT Hackathon Tier-3 submission S2.T3.1 AI Math Tutor for Early Learners. Source code: https://github.com/DrUkachi/ktt-math-tutor. Contents T3.1_Math_Tutor/ Core curriculum + seeds. curriculum.json — 80 items × 5 sub-skills (counting, number sense, addition, subtraction, word problem) with EN / FR / KIN stems, difficulty 1–10, age bands 5–6 / 6–7 / 7–8 / 8–9, visual asset keys, expected integer answer.… See the full description on the dataset page: https://huggingface.co/datasets/DrUkachi/ktt-math-tutor-data.tabularquestion-answeringn<1K0 likes22 downloads5mo agoHugging Face19metr-evals /daft-mathgated DAFT Math: Difficult Automatically-scorable Free-response Tasks for Math Dataset Description ⚠️ Note: The dataset has important limitations and we strongly recommend reading the limitations section below before using it. It is not a formal METR benchmark and was originally designed for a very niche use-case. We present it only as a research artifact. DAFT-Math is a collection of 199 challenging mathematical problems chosen to be at the limit of current LLM abilities… See the full description on the dataset page: https://huggingface.co/datasets/metr-evals/daft-math.tabularquestion-answeringn<1K2 likes19 downloads1y agoHugging Face20prithivMLmods /Math-Forge-Hard Math-Forge-Hard Dataset Overview The Math-Forge-Hard dataset is a collection of challenging math problems designed to test and improve problem-solving skills. This dataset includes a variety of word problems that cover different mathematical concepts, making it a valuable resource for students, educators, and researchers. Dataset Details Modalities Text: The dataset primarily contains text data, including math word problems. Formats… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-Forge-Hard.texttext-generation1K<n<10K6 likes16 downloads2y agoHugging Face21prithivMLmods /Math-Solve-Singleshot Math-Solve-Singleshot Overview This dataset, named Math-Solve-Singleshot, is designed for solving single-shot mathematical problems. It contains a variety of math problems formatted in text, suitable for training and evaluating models on mathematical reasoning tasks. Modalities Text Formats: CSV Size: 1.05M rows Libraries: pandas Croissant License: Apache-2.0 Dataset Details Train Split: 1.05 million rows Problem String Lengths: Length 1: 16… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-Solve-Singleshot.texttext-generation1M<n<10M6 likes15 downloads2y agoHugging Face22Mathivanan2025 /BTP_HFHubspottextquestion-answeringn<1K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.