CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Jackrong /DeepSeek-v3.1-reasoner-Distilled-math-samples DeepSeek-V3.1 Distillation with NVIDIA Nemotron-Post-Training-Dataset-v2 (Math Subset) The release of DeepSeek-V3.1 has attracted wide attention in the AI community. Its significant improvements in reasoning ability provide a new opportunity to explore optimization of domain-specific models. To investigate the potential of this model in complex mathematical reasoning tasks, I selected the math subset from NVIDIA’s newly released Nemotron-Post-Training-Dataset-v2 as seed problems and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-v3.1-reasoner-Distilled-math-samples.tabularquestion-answeringn<1K1 likes28 downloads1y agoHugging Face02vnytht /cancer-screening-evidence-reasoner Cancer Screening Evidence Reasoner (AutoScientist Challenge) Fine-tuning dataset for teaching a language model to answer cancer screening eligibility and evidence questions with exact, verifiable citations — not hedged guesses. Motivation Base models know screening guidelines roughly but invent citations and get exact statistics wrong. Every completion in this dataset is computed by a rule engine from verified USPSTF and SEER ground truth — not LLM-generated.… See the full description on the dataset page: https://huggingface.co/datasets/vnytht/cancer-screening-evidence-reasoner.textquestion-answering1K<n<10K0 likes27 downloads2mo agoHugging Face03alrobles /nemotron-eco-reasoner-v14 Nemotron Eco Reasoner — Dataset v14 Training dataset (9,686 records) for a LoRA adapter on NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, built for the NVIDIA Nemotron Model Reasoning Challenge. Each record is a chat-format example: a system prompt, a user puzzle, and an assistant response ending in </think> then \boxed{<answer>}. What is new in v14 v14 keeps the exact per-category composition of the best prior dataset (v8, the 0.67-Kaggle record) but replaces the… See the full description on the dataset page: https://huggingface.co/datasets/alrobles/nemotron-eco-reasoner-v14.texttext-generation1K<n<10K0 likes25 downloads3mo agoHugging Face04parallel-reasoner /parason-data parason-data Evaluation traces and structural-analysis artifacts for the parallel-reasoning line of work. Training data is not here — it lives in parallel-reasoner/sft-ours (splits 1x, 8x). This repo holds generated traces, so that structural claims about model behaviour can be re-derived rather than taken on trust. Layout aime24/<model-name>/traces.jsonl aime24/Qwen3-8B-sft-ours8x-ar/ Traces from parallel-reasoner/Qwen3-8B-sft-ours8x-ar — the… See the full description on the dataset page: https://huggingface.co/datasets/parallel-reasoner/parason-data.tabulartext-generationn<1K0 likes17 downloads2mo agoHugging Face05seacorn /news-summarizer-reasoner News Summarizer with Reasoning Overview This dataset is designed for news summarization tasks, featuring both original news articles and their corresponding summaries. The dataset includes a 'news' column sourced from three different datasets, along with 'cleaned_summary' and 'cleaned_reasoning' columns generated using large language models. Dataset Structure news: Original news articles. cleaned_summary: Summaries of the news articles in bullet points.… See the full description on the dataset page: https://huggingface.co/datasets/seacorn/news-summarizer-reasoner.textsummarization10K<n<100K0 likes13 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.