CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-SFT-Competitive-Programming-v2 Dataset Description: Nemotron-Competitive-Programming-v2 is a large-scale synthetic coding and reasoning dataset designed to push LLM performance on challenging programming and systems tasks. It combines Python and C++ samples across unique competitive programming questions.In addition, the dataset includes samples targetting a subset of programming language exercises from Exercism. Beyond problem solving, the dataset includes Text-to-SQL, which aims at training LLMs to reason and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Competitive-Programming-v2.text-generation27 likes3.6k downloads7mo agoHugging Face02nvidia /Nemotron-Competitive-Programming-v1 Dataset Description: Nemotron-Competitive-Programming-v1 is a large-scale synthetic coding and reasoning dataset designed to push LLM performance on challenging programming and systems tasks. It combines Python and C++ samples across unique competitive programming questions. Beyond problem solving, the dataset includes InfiniByte, a cross-domain subset with problems derived from scientific fields. This dataset is ready for commercial use. Competitive Coding The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Competitive-Programming-v1.31 likes3.4k downloads9mo agoHugging Face03Jackrong /Competitive-Programming-python-blend Dataset Card for Competitive-Programming-python-blend Summary Competitive-Programming-python-blend is a mixed supervised fine-tuning dataset centered on competitive programming, code reasoning, and instruction-style problem solving. The blend is Python-first, but it also keeps a small amount of C++, agentless SWE, and reasoning-oriented chat supervision to broaden training coverage. The current release is published as a single HF-friendly JSONL file, clean.jsonl.… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Competitive-Programming-python-blend.texttext-generation10K<n<100K21 likes489 downloads6mo agoHugging Face04multimodal-reasoning-lab /Competitive-Programmingimage1K<n<10K1 likes255 downloads1y agoHugging Face05jamesdborin /Nemotron-Competitive-Programming-v1-prompt-only Nemotron-Competitive-Programming-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-Competitive-Programming-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-Competitive-Programming-v1-prompt-only.tabular1M<n<10M0 likes92 downloads3mo agoHugging Face06jamesdborin /Nemotron-SFT-Competitive-Programming-v2-prompt-only Nemotron-SFT-Competitive-Programming-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Competitive-Programming-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Competitive-Programming-v2-prompt-only.tabular100K<n<1M0 likes86 downloads3mo agoHugging Face07amalia-llm /amalia-Nemotron-SFT-Competitive-Programming-v2 AMALIA Nemotron-SFT-Competitive-Programming-v2 Version of the nvidia/Nemotron-SFT-Competitive-Programming-v2 dataset used in the AMALIA's Supervised Fine-Tuning stage. This dataset went through a processing pipeline to remove entries that reference other LLMs or research labs; Original Dataset: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Competitive-Programming-v2 This dataset is provided as part of the AMALIA project and is included in the data mix used to… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SFT-Competitive-Programming-v2.text-generation0 likes54 downloads3mo agoHugging Face08databounty-io /build-competitive-programming-problem-dataset-cmsokff2 Build Competitive Programming Problem Dataset Each item is a build competitive programming problem dataset example providing Problem statement, Constraints, Reference solution, Test cases, Time limit (ms). Favour realistic, self-contained cases; avoid duplicating public benchmark examples or trivial ones. About This dataset was produced by the DataBounty community and published here as part of an open, karma-only program. Accepted items: 1000 Language: Python… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/build-competitive-programming-problem-dataset-cmsokff2.text1K<n<10K0 likes53 downloads12d agoHugging Face09StarsMakeGalaxy /competitive-programming-curated-600 🚀 Competitive Programming & Algorithmic Reasoning (Verbose CoT Reasoning) This dataset contains 600 curated training records with in-depth, verbose 4-phase <Thinking> Chain-of-Thought reasoning, 100 frozen evaluation benchmark samples, and 50 frozen regression verification samples formatted in standard ChatML (messages) and Prompt-Target pairs, strictly following the Pioneer / Prometheus research paper 3-slice curriculum design. 📊 Dataset Composition & 3-Slice… See the full description on the dataset page: https://huggingface.co/datasets/StarsMakeGalaxy/competitive-programming-curated-600.texttext-generationn<1K0 likes51 downloads1mo agoHugging Face10abanm /arjoonn-codechef-competitive-programming-ChatGPT4otext1K<n<10K0 likes44 downloads2y agoHugging Face11Arsh9210 /Nemotron-Competitive-Programming-v1 Dataset Description: Nemotron-Competitive-Programming-v1 is a large-scale synthetic coding and reasoning dataset designed to push LLM performance on challenging programming and systems tasks. It combines Python and C++ samples across unique competitive programming questions. Beyond problem solving, the dataset includes InfiniByte, a cross-domain subset with problems derived from scientific fields. This dataset is ready for commercial use. Competitive Coding The… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-Competitive-Programming-v1.0 likes32 downloads2mo agoHugging Face12nick007x /Nemotron-SFT-Competitive-Programming-v2 Dataset Description: Nemotron-Competitive-Programming-v2 is a large-scale synthetic coding and reasoning dataset designed to push LLM performance on challenging programming and systems tasks. It combines Python and C++ samples across unique competitive programming questions.In addition, the dataset includes samples targetting a subset of programming language exercises from Exercism. Beyond problem solving, the dataset includes Text-to-SQL, which aims at training LLMs to reason… See the full description on the dataset page: https://huggingface.co/datasets/nick007x/Nemotron-SFT-Competitive-Programming-v2.text-generation0 likes21 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.