CoolFace
13 results

synthlabs

SynthLabsAI /Big-Math-RL-Verifiedgated Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models Big-Math is the largest open-source dataset of high-quality mathematical problems, curated specifically for reinforcement learning (RL) training in language models. With over 250,000 rigorously filtered and verified problems, Big-Math bridges the gap between quality and quantity, establishing a robust foundation for advancing reasoning in LLMs. Request Early Access to Private… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/Big-Math-RL-Verified.textquestion-answering100K<n<1M243 likes5.4k downloads1y agoHugging FaceSynthLabsAI /PERSONAgated Dataset Card for PERSONAS (Prism Filter) PERSONAS (Prism filter) is one of the largest datasets of synthetic preferences, with over 200k preferences over thousands of questions and 1k personas. Details on the PERSONAS dataset can be found here paper link. Note that you MUST also fill out the form on our site to receive access to the full dataset. The form is available here. Dataset Details Dataset Description The personas dataset is a pluralistic… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/PERSONA.text100K<n<1M23 likes4k downloads2y agoHugging FaceSynthLabsAI /PERSONA_subsetgated Dataset Card for PERSONAS (Prism Filter) PERSONAS (Prism filter) is one of the largest datasets of synthetic preferences, with over 200k preferences over thousands of questions and 1k personas. Details on the PERSONAS dataset can be found here paper link Note that this subset is 5% of the training split of PERSONAS. The full dataset is here, strictly available for academic use. You MUST request access to the full persona dataset here. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/PERSONA_subset.text1K<n<10K3 likes3.3k downloads2y agoHugging Facemkurman /synthlabs-mlabonne-open-perfectblend PerfectBlend Synth Reasoning Synthetic reasoning traces generated for mlabonne/open-perfectblend. Each record contains conversations converted from ShareGPT format (from/value) to standard message format (role/content) with synthetically generated reasoning_content attached to each assistant turn. Dataset Summary 27,265 records across 8 source datasets 37,158 reasoning turns (99.7% format compliance) Average 1,376 chars per reasoning trace Reasoning generated… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-mlabonne-open-perfectblend.tabulartext-generation10K<n<100K0 likes46 downloads2mo agoHugging Facemkurman /gsm8k-SynthLabs-reasoning GSM8K-SynthLabs This dataset is a refined version of the GSM8K dataset, enriched with complex reasoning traces in the style of Pleias/SYNTH. It is designed for fine-tuning large language models to improve their reasoning capabilities using a structured, step-by-step thinking process. Key Features SYNTH Reasoning: Each problem contains a detailed reasoning trace generated by DeepSeek-V3.2, following the structured format (e.g., Query Parsing, Decomposition… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/gsm8k-SynthLabs-reasoning.tabularquestion-answering1K<n<10K9 likes43 downloads8mo agoHugging Facemkurman /synthlabs-llm-blender-mix-instruct-19k LLM Blender Synth Reasoning Synthetic reasoning traces for the LLM Blender Mix Instruct dataset, generated with Qwen3.6-27B and Qwen3.6-35B-A3B. Each record contains a general-purpose instruction with SYNTH-style reasoning and a generated answer. Dataset Summary 19,010 records (1,490 dupes + 847 incomplete removed from 21,347 source) 19,010 reasoning turns (99.9% format compliance) Average 1,130 chars per reasoning trace Provider Provider… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-llm-blender-mix-instruct-19k.tabulartext-generation10K<n<100K0 likes41 downloads2mo agoHugging Face