CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kaysss /leetcode-problem-solutions LeetCode Solution Dataset This dataset contains community-contributed LeetCode solutions scraped from public discussions and solution pages, enriched with metadata such as vote counts, author info, tags, and full code content. The goal is to make high-quality, peer-reviewed coding solutions programmatically accessible for research, analysis, educational use, or developer tooling. Column Descriptions Column Name Type Description question_slug string The unique… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-solutions.tabulartext-classification100K<n<1M9 likes5.3k downloads1y agoHugging Face02greghavens /gpt-5.6-sol-coding-and-debugging-traces GPT-5.6 Sol Coding & Debugging Traces Verified software-engineering, independent model-judging, seed-authoring, defensive-security, and training-harness trajectories from GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an autonomous coding agent. Sessions show the observable development loop: inspecting repositories, reproducing failures, explaining evidence, editing files, running compilers and test suites, correcting mistakes, and verifying the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/gpt-5.6-sol-coding-and-debugging-traces.text-generation10K<n<100K44 likes1.1k downloads2mo agoHugging Face03Crownelius /GPT-5.6-Sol-Luna-Terra-Traces GPT-5.6 — Sol · Terra · Luna Library A maintained mirror of every GPT-5.6 Sol / Terra / Luna dataset on Hugging Face — content-verified, attributed, in one place. Dataset Viewer | Parquet // what this is This is a maintained library — a community mirror of every publicly-available GPT-5.6 Sol / Terra / Luna dataset on Hugging Face, aggregated, validity-filtered, and content-verified with per-row source attribution. It is not Crownelius' own data. Every row… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/GPT-5.6-Sol-Luna-Terra-Traces.tabulartext-generation10K<n<100K18 likes873 downloads2mo agoHugging Face04iiis-lean /NuminaMath-LEAN-Sol NuminaMath-LEAN Cleaned with NL Solutions Dataset Summary This is a cleaned version of the NuminaMath-LEAN dataset, enhanced with natural language (NL) solutions matched from source datasets. The primary goal is to provide paired formal statements/proofs with natural language solutions for proof formalization and theorem proving research. The dataset matches problems from NuminaMath-LEAN with their corresponding natural language solutions from: olympiads-ref: A… See the full description on the dataset page: https://huggingface.co/datasets/iiis-lean/NuminaMath-LEAN-Sol.texttext-generation10K<n<100K0 likes803 downloads8mo agoHugging Face05MCES10-Software /Python-Code-Solutions Python Code Solutions Features 1000k of Python Code Solutions for Text Generation and Question Answering Python Coding Problems labelled by topic and difficulty Recommendations Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models. Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic} textquestion-answering10K<n<100K0 likes626 downloads1y agoHugging Face06olegshulyakov /doocs-leetcode-solutions Doocs LeetCode Solutions A comprehensive dataset of LeetCode problems and solutions created from the Doocs LeetCode repository. This dataset is designed for fine-tuning large language models to understand programming problems and generate code solutions. Description Repository: Doocs LeetCode Solutions Total Problems: 3500+ Total Solutions: 15,000+ (across multiple languages) Size: ~60 MB (Parquet format) Languages: C Cangjie C++ C# Dart Go Java JavaScript Kotlin Nim PHP… See the full description on the dataset page: https://huggingface.co/datasets/olegshulyakov/doocs-leetcode-solutions.texttext-generation10K<n<100K3 likes570 downloads1y agoHugging Face07LLaMAX /BenchMAX_Problem_Solving Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Problem_Solving is a dataset of BenchMAX, sourcing from LiveCodeBench_v4, which evaluates the code generation capability for solving multilingual competitive code problems. We extend the original English dataset by 16 non-English languages. The… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Problem_Solving.texttext-generation10K<n<100K1 likes531 downloads2y agoHugging Face08Solstice-AI /Solace-1.0-Omnigated Project Solace The largest verified frontier-model distillation corpus ever released. 60 datasets · 7 frontier model families · 12,586,893 unique conversations · One file · Zero filler The short version This is synthetic data. The best kind of synthetic data. Every example was generated by a verified 2026 frontier model — GLM-5.2, Claude Fable 5, Mythos 5, GPT-5.6 Sol, GPT-5.5 Codex, DeepSeek V4 Pro 0813, Qwen 3.8-Max, and Kimi K3 — then… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Solace-1.0-Omni.texttext-generation10M<n<100M6 likes530 downloads19d agoHugging Face09solis-soict /xrepotest XRepoTest: Multilingual Repository-Level Unit Test Generation Benchmark Paper | Code The data release for the XRepoTest benchmark (EMNLP 2026 Main). Each task item is a function extracted from a real open-source repository. The goal is to generate a unit test for that function using its surrounding repository context. Generated tests are executed and scored inside Docker on the real codebase. Structure Folder Contents Pipelines data/base per-language… See the full description on the dataset page: https://huggingface.co/datasets/solis-soict/xrepotest.text-generation1 likes420 downloads27d agoHugging Face10wAI-org /swerl-tmax-15k-solvable-gpt-5-6-terra swerl-tmax-15k hardened, post-validation-filter (dataset 3 of 3) Which tasks in hamishivi/swerl-tmax-15k can a strong model actually solve? Every task was attempted twice as a full agentic episode — real sandbox, real bash, real verifier — and a task is verified when at least one attempt earned reward. The last of three artifacts that exist to be compared by task_id: original — hamishivi/swerl-tmax-15k, unchanged — 14,601 tasks hardened, pre-validation-filter —… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-solvable-gpt-5-6-terra.tabulartext-generation10K<n<100K1 likes413 downloads8d agoHugging Face11wAI-org /swerl-tmax-15k-rubric-gpt-5-6-sol swerl-tmax-15k with task-quality rubric labels (gpt-5-6-sol) hamishivi/swerl-tmax-15k, unchanged and unfiltered, with a per-task quality label attached as extra columns. This is not a verified or filtered dataset. Every one of the 14,601 original records is present. Nothing has been dropped, repaired, or reordered. The labels are one model's judgement about whether each task is sound enough to be useful RL training data — an annotation layer, not a correctness guarantee.… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-rubric-gpt-5-6-sol.texttext-generation10K<n<100K0 likes331 downloads11d agoHugging Face12SoLID /shellcode_i_a32Shellcode_IA32 is a dataset for shellcode generation from English intents. The shellcodes are compilable on Intel Architecture 32-bits.texttext-generation1K<n<10K13 likes270 downloads4y agoHugging Face13Solstice-AI /Axiom-1.0-Opus4.7-Kimi2.6-GLM5.2-Deepseek4-Mythos5-Fable5-Qwen3.7 Project Axiom 1.0 (102 GB Reasoning Corpus) 27-Billion Token Pure-Text Chain-of-Thought Corpus Across 7 Frontier Architectures Executive Summary Project Axiom 1.0 is a landmark, high-density, multi-architecture reasoning corpus comprising 102 GB of uncompressed, pure-text JSONL data (axiom.jsonl). Curated by Shreyan Gondaliya and the Solstice-AI research team, the dataset synthesizes ~5.74 million unique samples and ~27.3 billion tokens of… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Axiom-1.0-Opus4.7-Kimi2.6-GLM5.2-Deepseek4-Mythos5-Fable5-Qwen3.7.text-generation13 likes237 downloads19d agoHugging Face14khaimaitien /leetcode_problem_solutionThis dataset contains: problems and solutions in Leetcode, crawled from: https://github.com/AnasImloul/Leetcode-Solutions The format of data: title: title of the problem algo_input: the description of the problem solution_py: the solution in Python solution_js: the solution in Js solution_java: the solution in Java solution_c: the solution in C texttext-generation1K<n<10K2 likes192 downloads3y agoHugging Face15Solstice-AI /Complete-FABLE.5-traces-2M Complete FABLE.5 Traces (2 Million Deduplicated Rows) Comprehensive Agentic Coding & Frontier Reasoning Trajectory Corpus Executive Summary Solstice-AI/Complete-FABLE.5-traces-2M is a clean, fully deduplicated post-training dataset containing 2,006,487 high-entropy agentic coding and multi-step reasoning traces. Originally curated following the closure of Fable and Mythos, this corpus synthesizes frontier agent execution patterns (including Claude… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Complete-FABLE.5-traces-2M.tabulartext-generation1M<n<10M3 likes184 downloads19d agoHugging Face16CollinL /perovskite-solar-cell-efficiency-autoresearch 🔬 Perovskite Solar Cell Text Corpus for Karpathy's autoresearch A 98.9 MB text corpus of perovskite solar cell scientific literature formatted for direct use with karpathy/autoresearch — the autonomous LLM-driven hyperparameter search framework that trains a GPT from scratch and has an AI agent iteratively modify train.py to minimize val_bpb (bits per byte). 📊 Dataset Stats Metric Value Total documents 19,730 Total text 98.9 MB (~103M characters)… See the full description on the dataset page: https://huggingface.co/datasets/CollinL/perovskite-solar-cell-efficiency-autoresearch.texttext-generation10K<n<100K0 likes162 downloads5mo agoHugging Face17DenCT /leetcode-python-solutions-with-exaplanationstabulartext-generation10K<n<100K2 likes153 downloads2y agoHugging Face18Axiom-AI /Small-HLE-Solved Small-HLE-Solved Small-HLE-Solved is a curated dataset consisting of challenging problems selected from the Humanity's Last Exam (HLE) benchmark. Each instance has been processed by an advanced teacher model to generate high-fidelity, multi-step reasoning paths. The dataset is formatted strictly in JSON Lines (jsonl), pairing each complex problem with a structured, step-by-step solution optimized for training next-generation reasoning models. 📂 Data Structure &… See the full description on the dataset page: https://huggingface.co/datasets/Axiom-AI/Small-HLE-Solved.texttext-generationn<1K1 likes153 downloads4mo agoHugging Face19ordlibrary /solana-clawd-model-kit Solana Clawd Model Kit Training data kit for Solana Clawd: SFT / CPT JSONL corpora, manifests, quality reports, and processed shards. Contents (top-level) SFT / CPT JSONL solana_clawd_reasoning_tooling_sft.jsonl (~133 MB) clawd_masterpiece_sft.jsonl (~166 MB) tx_foundation_cpt_clean.jsonl (~21 MB) clawd_future_refinement_sft.jsonl (~1.3 MB) clawd_autoresearch_wiki_sft.jsonl (~1.3 MB) clawd_future_drill_sft.jsonl (~920 KB)… See the full description on the dataset page: https://huggingface.co/datasets/ordlibrary/solana-clawd-model-kit.text-generation100K<n<1M0 likes130 downloads8d agoHugging Face20solanaclawd /solana-clawd-instruct Solana Clawd Instruct A curated instruction-tuning dataset for fine-tuning models into Solana-native Clawd agents with strong Solana, DeFi, ZK, and constitutional-alignment coverage. What it teaches Check every domain your dataset covers: Solana mechanics (PDAs, accounts, instructions, rent, compute budgets, Token-2022) DeFi primitives (AMMs, CLMMs, perpetuals, bonding curves, Jupiter, Phoenix) Memecoin risk analysis (rug detection, holder concentration… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-instruct.texttext-generation10K<n<100K0 likes129 downloads3mo agoHugging Face21samscrack /solidity-audit-cot solidity-audit-cot Long-CoT audit traces for Solidity contracts, generated by Claude Opus 4.7 (adaptive thinking, xhigh effort) over the spec→contract corpus from the Qwopus3.6-27B-solidity training pipeline. This dataset is the Stage 2 training corpus for the multi-stage Qwopus3.6-27B-solidity model — designed to teach long-form security reasoning (8-15 paragraph chain-of-thought) anchored to real Solidity contracts. Why this dataset exists Public Solidity audit… See the full description on the dataset page: https://huggingface.co/datasets/samscrack/solidity-audit-cot.texttext-generation1K<n<10K3 likes110 downloads5mo agoHugging Face22SolusOps /incremental-instruction-creative-writinggated Incremental Instruction Creative Writing Does delivering a writing brief over several conversation turns change what a language model writes? This dataset supports that question with matched creative-writing tasks evaluated under two delivery conditions: FULL: the complete brief is supplied in one turn. SHARDED: the same intended brief is introduced across five to nine turns. The benchmark holds task content fixed while varying how the instructions are delivered. It is… See the full description on the dataset page: https://huggingface.co/datasets/SolusOps/incremental-instruction-creative-writing.tabulartext-generation1K<n<10K0 likes108 downloads22d agoHugging Face23kaushik-harsh-99 /math-sft-solutions-no-cot Math SFT Solutions No CoT A cleaned mathematics supervised fine-tuning dataset containing: instruction → solution pairs mathematical proofs derivations olympiad-style solutions theorem reasoning stepwise mathematical explanations detailed final solutions This dataset was built specifically for mathematical supervised fine-tuning (SFT). Unlike many reasoning datasets, this release removes explicit chain-of-thought tags and hidden thinking traces while preserving high-quality… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/math-sft-solutions-no-cot.texttext-generation100K<n<1M5 likes97 downloads4mo agoHugging Face24Nobody05 /gpt-5.6-sol-coding-and-debugging-traces GPT-5.6 Sol Coding & Debugging Traces Verified software-engineering, independent model-judging, seed-authoring, defensive-security, and training-harness trajectories from GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an autonomous coding agent. Sessions show the observable development loop: inspecting repositories, reproducing failures, explaining evidence, editing files, running compilers and test suites, correcting mistakes, and verifying the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/gpt-5.6-sol-coding-and-debugging-traces.text-generation10K<n<100K1 likes97 downloads2mo agoHugging Face25Januka2009 /GPT5.6_SOL_INVESTIGACION Dataset de Metodología Científica Dataset en español para entrenamiento, validación y evaluación de modelos capaces de razonar sobre metodología de investigación científica. Incluye escenarios de distintas disciplinas y niveles de dificultad, con énfasis en diseño de estudios, inferencia causal, análisis cuantitativo y cualitativo, métodos mixtos, ética, medición, muestreo, interpretación de resultados y revisión crítica de protocolos. 1. Resumen… See the full description on the dataset page: https://huggingface.co/datasets/Januka2009/GPT5.6_SOL_INVESTIGACION.texttext-generation1K<n<10K1 likes95 downloads1d agoHugging Face26CharlieLLL /SWEbench-Verified-eval150-M2.7-solo-selforch-3repeats-w32-20260920 M2.7 Solo and self-orchestration: three independent eval150 runs each All six fresh runs completed the same150 tasks and passed original result/trajectory/task/attempt/fingerprint audits. No previous scores were pooled. Each repeat starts new model processes and cold KV caches after a real telemetry smoke. True failed tasks are retained; infrastructure retries are preserved separately. Mode Repeat1 Repeat2 Repeat3 Mean /150 Sample SD m27-solo 94 98 90 94.00 4.00… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-solo-selforch-3repeats-w32-20260920.text-generation0 likes95 downloads2d agoHugging Face27solanaclawd /solana-clawd-repo-corpus Solana Clawd Core AI Instruct Instruction-tuning dataset derived from the local core-ai source tree and the existing Solana Clawd AI training corpus. Contents Total examples: 441 Existing ai-training SFT examples: 0 Core AI source chunk examples: 0 Core AI knowledge JSONL examples: 0 Format Each row is a chat conversation in OpenAI/Hugging Face messages schema: {"messages": [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-repo-corpus.texttext-generation10K<n<100K1 likes86 downloads16d agoHugging Face28Manusagents /GPT-5.6-Sol-Luna-Terra-Traces GPT-5.6 — Sol · Terra · Luna Library A maintained mirror of every GPT-5.6 Sol / Terra / Luna dataset on Hugging Face — content-verified, attributed, in one place. Dataset Viewer | Parquet // what this is This is a maintained library — a community mirror of every publicly-available GPT-5.6 Sol / Terra / Luna dataset on Hugging Face, aggregated, validity-filtered, and content-verified with per-row source attribution. It is not Crownelius' own data. It exists to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.6-Sol-Luna-Terra-Traces.tabulartext-generation1K<n<10K3 likes83 downloads2mo agoHugging Face29kaushik-harsh-99 /math-sft-solutions-no-cot-v3 Math SFT Solutions No CoT V3 Math SFT Solutions No CoT V3 is a large-scale mathematics supervised fine-tuning (SFT) dataset designed for instruction tuning and mathematical capability adaptation. Version 3 substantially expands mathematical coverage while improving dataset quality through stronger filtering, cleaning, and supervision refinement. Unlike reasoning-heavy datasets, this release focuses on clean instruction → response pairs without hidden chain-of-thought style… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/math-sft-solutions-no-cot-v3.texttext-generation1M<n<10M5 likes81 downloads4mo agoHugging Face30samscrack /solidity-cpt-top10-quality Solidity CPT Top-10% Quality-Filtered Corpus A curated, deduplicated corpus of 23,471 modern Solidity source files (~86M tokens) intended for continued-pretraining (CPT) of code LLMs on smart-contract code. It's the top 10% slice (by composite quality score) of a larger raw corpus that combined: ASSERT-KTH/DISL — 514 k unique deployed Solidity files, deduped at file level 30 hand-picked GitHub blue-chip protocols (OpenZeppelin, Uniswap v2/v3/v4, Aave v3, Compound, Morpho… See the full description on the dataset page: https://huggingface.co/datasets/samscrack/solidity-cpt-top10-quality.texttext-generation10K<n<100K0 likes77 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.