CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SynthLabsAI /Big-Math-RL-Verifiedgated Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models Big-Math is the largest open-source dataset of high-quality mathematical problems, curated specifically for reinforcement learning (RL) training in language models. With over 250,000 rigorously filtered and verified problems, Big-Math bridges the gap between quality and quantity, establishing a robust foundation for advancing reasoning in LLMs. Request Early Access to Private… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/Big-Math-RL-Verified.textquestion-answering100K<n<1M243 likes5.4k downloads2y agoHugging Face02mkurman /synthlabs-mlabonne-open-perfectblend PerfectBlend Synth Reasoning Synthetic reasoning traces generated for mlabonne/open-perfectblend. Each record contains conversations converted from ShareGPT format (from/value) to standard message format (role/content) with synthetically generated reasoning_content attached to each assistant turn. Dataset Summary 27,265 records across 8 source datasets 37,158 reasoning turns (99.7% format compliance) Average 1,376 chars per reasoning trace Reasoning generated… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-mlabonne-open-perfectblend.tabulartext-generation10K<n<100K0 likes48 downloads2mo agoHugging Face03mkurman /gsm8k-SynthLabs-reasoning GSM8K-SynthLabs This dataset is a refined version of the GSM8K dataset, enriched with complex reasoning traces in the style of Pleias/SYNTH. It is designed for fine-tuning large language models to improve their reasoning capabilities using a structured, step-by-step thinking process. Key Features SYNTH Reasoning: Each problem contains a detailed reasoning trace generated by DeepSeek-V3.2, following the structured format (e.g., Query Parsing, Decomposition… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/gsm8k-SynthLabs-reasoning.tabularquestion-answering1K<n<10K9 likes46 downloads8mo agoHugging Face04mkurman /synthlabs-llm-blender-mix-instruct-19k LLM Blender Synth Reasoning Synthetic reasoning traces for the LLM Blender Mix Instruct dataset, generated with Qwen3.6-27B and Qwen3.6-35B-A3B. Each record contains a general-purpose instruction with SYNTH-style reasoning and a generated answer. Dataset Summary 19,010 records (1,490 dupes + 847 incomplete removed from 21,347 source) 19,010 reasoning turns (99.9% format compliance) Average 1,130 chars per reasoning trace Provider Provider… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-llm-blender-mix-instruct-19k.tabulartext-generation10K<n<100K0 likes42 downloads2mo agoHugging Face05mkurman /synthlabs-openmed-questions-qwen3-235b-a22b-2507 Med Synth Questions (Qwen3-235B questions + DeepSeek V4 Flash and Minimax M2.7 answers) Synthetic reasoning traces for medical questions from openmed-community/med-synth-questions-qwen3-235b-a22b-2507. Each record contains a medical question with SYNTH-style reasoning and a generated answer. Dataset Summary 55,915 records (255 dupes + 3,523 incomplete/truncated removed from 59,693 source) 55,915 reasoning turns (99.9% format compliance) Average 1,881 chars per… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-openmed-questions-qwen3-235b-a22b-2507.tabulartext-generation10K<n<100K0 likes35 downloads2mo agoHugging Face06mkurman /synthlabs-chat-final-cleaned-v3 SynthLabs Chat Final — Cleaned v3 A cleaned instruction-following chat dataset for supervised fine-tuning (SFT) of reasoning-capable language models. Each example is a conversation (user ↔ assistant) with explicit chain-of-thought reasoning (reasoning_content) separated from the final answer (content). Dataset Structure Schema Each record contains: Field Type Description messages list[struct] Conversation turns (2-8+ turns) session_uid… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-chat-final-cleaned-v3.tabulartext-generation10K<n<100K3 likes34 downloads3mo agoHugging Face07mkurman /synthlabs-GLM-5.2-Science GLM-5.2 Science Synth Reasoning Synthetic reasoning traces for science questions from the GLM-5.2 science dataset. Each record contains a complex scientific question with SYNTH-style reasoning and a generated answer. Dataset Summary 33,014 records (605 dupes + 726 incomplete/truncated removed from 34,345 source) 33,014 reasoning turns (99.9% format compliance) Average 3,094 chars per reasoning trace Models Used Model Records… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-GLM-5.2-Science.tabulartext-generation10K<n<100K0 likes28 downloads2mo agoHugging Face08mkurman /medical-reasoning-synthlabs-I Medical Reasoning SynthLabs I Medical Reasoning SynthLabs I is a synthetic medical reasoning dataset generated with SynthLabs. It contains medical question-answer examples paired with structured reasoning traces, intended for research on reasoning-style instruction tuning, medical QA, answer synthesis, and reasoning trace analysis. The dataset is designed for machine learning research and experimentation. It is not intended for clinical decision-making, diagnosis, treatment… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/medical-reasoning-synthlabs-I.texttext-generationn<1K0 likes25 downloads4mo agoHugging Face09mkurman /synthlabs-chat-final-cleaned-v2 SynthLabs Chat Final — Cleaned v2 (Strict) A cleaned instruction-following chat dataset for supervised fine-tuning (SFT) of reasoning-capable language models. Each example is a two-turn conversation (user → assistant) with explicit chain-of-thought reasoning separated from the final answer. Dataset Structure Schema Each record contains: Field Type Description messages list[struct] Conversation turns session_uid string Unique session… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-chat-final-cleaned-v2.texttext-generation10K<n<100K2 likes15 downloads3mo agoHugging Face10mkurman /medical-synthlabs-small Medical SynthLabs Small (MedAlpaca Flashcards Subset) Dataset Summary This dataset is a small synthetic subset (77 examples) derived from medalpaca/medical_meadow_medical_flashcards, expanded with SYNTH-style reasoning traces and packaged as a lightweight Parquet dataset for quick experimentation. Generation was performed using SynthLabs.app, a workflow for creating synthetic datasets and reasoning-augmented conversions. Disclaimer: The content is for research and… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/medical-synthlabs-small.texttext-generationn<1K1 likes14 downloads8mo agoHugging Face11mkurman /math-small-synthlabs Math Small SynthLabs Dataset Summary This dataset is a small synthetic math-focused dataset generated using SynthLabs.app. It is designed for quick experimentation with instruction-following, step-by-step reasoning traces, and short-form question answering in a lightweight format (Parquet). Note: Reasoning traces are model-generated and may contain errors. Use for research/education only. What’s in this Dataset Data Fields Each row typically… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/math-small-synthlabs.texttext-generationn<1K1 likes9 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.