CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01m-a-p /CodeFeedback-Filtered-Instruction OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement [🏠Homepage] | [🛠️Code] OpenCodeInterpreter OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities. For further information and… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/CodeFeedback-Filtered-Instruction.textquestion-answering100K<n<1M208 likes29k downloads3y agoHugging Face02ishumilin /epstein-files-ocr-datasets-1-8-early-release Epstein Files OCR — Datasets 1–8 (Early Release) ARCHIVE NOTICE This dataset is no longer maintained. Please refer to the Epstein Files — Complete OCR Dataset. Dataset Summary This dataset contains page-level OCR output (as Markdown) from a public release of documents related to Jeffrey Epstein / the Epstein case. Each Markdown file represents one scanned page converted to text using an automated OCR pipeline. The dataset is designed for: Question answering Information… See the full description on the dataset page: https://huggingface.co/datasets/ishumilin/epstein-files-ocr-datasets-1-8-early-release.question-answering10K<n<100K1 likes2.8k downloads6mo agoHugging Face03MoreThought /Fable-5.1-Max-Reasoning-Filtered-10000x Dataset Description This dataset contains 10,000 agentic coding and reasoning multi-turn high-quality traces generated by the new Fable 5.1 model using max reasoning effort. It holds almost 500,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains. It has also been deduplicated and heavily filtered to remove low-quality traces, keeping only high-quality traces. Dataset Statistics Metric Value Total Examples 10,000… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/Fable-5.1-Max-Reasoning-Filtered-10000x.text-generation10K<n<100K155 likes2.6k downloads9h agoHugging Face04SegunOni /osworld_tasks_filestext-classification1M<n<10M0 likes1.3k downloads8mo agoHugging Face05birdsql /bird23-train-filtered BIRD-SQL Train (Filtered) A high-quality subset of the original BIRD train split for text-to-SQL finetuning. Overview Over the past year the community has shared many observations about data quality in BIRD. We performed a rigorous data quality check process to retain examples that are consistent with schema and faithfully answer the question. The resulting set keeps 6,601 instances out of 9,428 (≈70%), and serves as a drop-in replacement for training. Original Train: 9… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird23-train-filtered.texttable-question-answering1K<n<10K7 likes982 downloads1y agoHugging Face06anonymous-md /EDGAR_FILINGS_DATASET SFD: SEC Filings Dataset (v1) SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation. This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in: The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.tabulartext-generation1M<n<10M2 likes895 downloads5mo agoHugging Face07aurora2424 /epstein-files-ocr-datasets-1-8-early-release Epstein Files OCR — Datasets 1–8 (Early Release) Work in Progress (WIP) This is an early publication. We are actively working on improving OCR quality and expanding coverage. Dataset Summary This dataset contains page-level OCR output (as Markdown) from a public release of documents related to Jeffrey Epstein / the Epstein case. Each Markdown file represents one scanned page converted to text using an automated OCR pipeline. The dataset is designed for: Question… See the full description on the dataset page: https://huggingface.co/datasets/aurora2424/epstein-files-ocr-datasets-1-8-early-release.question-answering10K<n<100K0 likes775 downloads7mo agoHugging Face08Choiszt /FileGram FileGram Dataset Grounding Agent Personalization in File-System Behavioral Traces Overview FileGram is a comprehensive framework for evaluating memory-centric personalization from file-system behavioral traces. This dataset provides: 640 behavioral trajectories — 20 persona-driven profiles x 32 tasks (16 text-centric + 16 multimodal), each containing fine-grained file-system operation logs, content snapshots, and session statistics 4,333 QA pairsacross 4… See the full description on the dataset page: https://huggingface.co/datasets/Choiszt/FileGram.textquestion-answering1K<n<10K5 likes519 downloads5mo agoHugging Face09lateesha-bhatia /sec-filings-qa-instruct SEC Filings Instruction-Tuning Dataset (Llama-3 Format) This dataset contains 5,000 curated, instruction-formatted question-answering pairs derived from corporate SEC filings (Forms 10-K and 10-Q). It is structured specifically for parameter-efficient instruction fine-tuning (SFT/QLoRA) of Small Language Models using the standard Llama-3 ChatML template. Dataset Details Origin Source: Curated subset extracted from nvidia/Nemotron-SpecializedDomains-Finance-v1.… See the full description on the dataset page: https://huggingface.co/datasets/lateesha-bhatia/sec-filings-qa-instruct.textquestion-answering1K<n<10K0 likes273 downloads19d agoHugging Face10KrazyKitty /Fable-5.1-Max-Reasoning-Filtered-1000x Dataset Description This dataset contains 1,000 coding and reasoning traces generated by the new Fable 5.1 model using max reasoning effort. It holds almost 30,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains. It has also been deduplicated and filtered to remove low-quality traces, keeping only high-quality traces. Dataset Statistics Metric Value Total Examples 1,000 Traces Total Token Count ~30,000,000… See the full description on the dataset page: https://huggingface.co/datasets/KrazyKitty/Fable-5.1-Max-Reasoning-Filtered-1000x.texttext-generation1K<n<10K6 likes271 downloads13d agoHugging Face11beatsprom /swe-bench-multi-file-refactoring-sft-dpo-2026 💻 Enterprise Autonomous SWE-bench AI & Multi-File Code Refactoring SFT/DPO Dataset (2026) High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step call-stack Chain-of-Thought (<thought>) reasoning trees for fine-tuning LLMs (Qwen-2.5-Coder, Llama-3.3, DeepSeek-R1-Distill, Mistral) into Autonomous Software Engineers and SWE-bench Benchmark Agents. 📊 Dataset Architecture & Highlights Multi-Turn Code Reviews:… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/swe-bench-multi-file-refactoring-sft-dpo-2026.texttext-generationn<1K0 likes218 downloads25d agoHugging Face12musk1209 /finsight-sec-filings FinSight — SEC EDGAR Filings Cleaned plain-text 10-K (annual) and 10-Q (quarterly) filings from the US SEC EDGAR system for 20 large publicly-traded companies across 6 sectors. Created as part of the FinSight project — a financial research AI assistant combining BERT fine-tuning, RAG, and multi-agent systems. Stats Records: 97 Companies: 20 (AAPL, MSFT, GOOGL, AMZN, META, NVDA, TSLA, JPM, BAC, GS, JNJ, PFE, UNH, WMT, PG, KO, MCD, XOM, CVX, CAT) Forms: 10-K, 10-Q… See the full description on the dataset page: https://huggingface.co/datasets/musk1209/finsight-sec-filings.tabularquestion-answeringn<1K0 likes186 downloads3mo agoHugging Face13LARK-Lab /EnvFactory-SFT-FILTERED EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL ## Overview EnvFactory-SFT-FILTERED is a filtered supervised fine-tuning (SFT) dataset containing 53,400 tool-use trajectories synthesized using the EnvFactory framework. This dataset is designed for SFT training of tool-use agents. The dataset contains high-quality multi-turn tool-use trajectories with implicit human reasoning, generated through… See the full description on the dataset page: https://huggingface.co/datasets/LARK-Lab/EnvFactory-SFT-FILTERED.texttext-generation10K<n<100K0 likes183 downloads4mo agoHugging Face14mirzaei2114 /stackoverflowVQA-filteredimagevisual-question-answering100K<n<1M3 likes179 downloads3y agoHugging Face15locailabs /nemotron_terminal_filtered Nemotron Terminal Filtered An uncertainty-curated subset of NVIDIA's Nemotron-Terminal-Corpus (dataset_adapters split), selected for high-formation density for post-training NVIDIA-Nemotron-3-Super-120B-A12B-BF16. Motivation The full dataset_adapters split contains ~226k terminal execution trajectories. To curate a compact, high-value subset for post-training we score each sample by how hard the model finds it, using entropy as a proxy for uncertainty. The… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/nemotron_terminal_filtered.textquestion-answering10K<n<100K2 likes157 downloads5mo agoHugging Face16rayonlabs /wildchat-filtered WildChat Filtered Dataset This is a filtered version of the WildChat-4.8M dataset. Dataset Description This dataset contains 3,199,860 conversations between human users and ChatGPT, filtered to keep only the essential conversation structure. Data Structure Each conversation contains only: conversations: A list of message objects with: role: Either "user" or "assistant" content: The text content of the message All other metadata (timestamps, moderation… See the full description on the dataset page: https://huggingface.co/datasets/rayonlabs/wildchat-filtered.texttext-generation1M<n<10M1 likes156 downloads1y agoHugging Face17ishumilin /epstein-files-ocr-complete Epstein Files — Complete OCR Dataset This is a comprehensive, structured publication of the Epstein Files OCR dataset, significantly expanding upon the earlier Datasets 1-8 release. Dataset Summary This dataset contains page-level OCR output compiled from an extensive release of documents related to Jeffrey Epstein / the Epstein case. Each row in this dataset represents one scanned PDF document from the original release using a proprietary automated OCR pipeline… See the full description on the dataset page: https://huggingface.co/datasets/ishumilin/epstein-files-ocr-complete.textquestion-answering1M<n<10M3 likes154 downloads6mo agoHugging Face18laion /nemotron-terminal-file_operations nemotron-terminal-file_operations Per-source partition of nvidia/Nemotron-Terminal-Corpus, filtered to source == "file_operations". The difficulty column preserves the original easy / medium / mixed split (na for the dataset_adapters/* files, which did not carry a difficulty label). Partitioning scheme: adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet {skill} (e.g. debugging, security, …) — rows from synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-file_operations.textquestion-answering10K<n<100K0 likes153 downloads6mo agoHugging Face19antfr99 /hitchcock-psycho-1960-film-dataset-transformed Psycho → AI-Model Dataset (Transformed) A thematic re-skin of the Psycho (1960) Q&A dataset into an original AI-model setting where the world is transformed into an AI/data-center environment. Character names, actor names, objects, locations, production references, dates, and thematic elements are remapped to AI/ML concepts and modern technology. File: psycho_dataset_transformed.jsonl Format: JSONL — one JSON object per line Schema: each line has prompt and completion string… See the full description on the dataset page: https://huggingface.co/datasets/antfr99/hitchcock-psycho-1960-film-dataset-transformed.texttext-generation1K<n<10K0 likes140 downloads9d agoHugging Face20fxmeng /commonsense_filtered Dataset Summary The commonsense reasoning tasks consist of 8 subtasks, each with predefined training and testing sets, as described by LLM-Adapters (Hu et al., 2023). The following table lists the details of each sub-dataset. Train Test Information BoolQ (Clark et al., 2019) 9427 3270 Question-answering dataset for yes/no questions PIQA (Bisk et al., 2020) 16113 1838 Questions with two solutions requiring physical commonsense to answer SIQA (Sap et al., 2019) 33410… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/commonsense_filtered.textquestion-answering100K<n<1M1 likes128 downloads2y agoHugging Face21amphora /ResearchMath-Filtered ResearchMath-Filtered ResearchMath-Filtered is a quality-filtered collection of 129,927 long-form reasoning traces and solutions for research-level mathematical problems, released alongside ResearchMath-14k as part of the same paper. It is a cleaned subset of ResearchMath-Reasoning-194K: each record holds a self-contained problem statement, a long chain-of-thought reasoning trace, and a final response, with low-quality and non-solving generations removed. Paper:… See the full description on the dataset page: https://huggingface.co/datasets/amphora/ResearchMath-Filtered.texttext-generation100K<n<1M0 likes125 downloads3mo agoHugging Face22filippo19741974 /Generated-Recovery-Support-Dialogues # Empathetic Conversations for Addiction Recovery Support Dataset Dataset Description This dataset contains synthetically generated conversational examples between a user discussing their addiction recovery journey and an AI assistant designed to be empathetic, supportive, non-judgmental, and encouraging. The conversations are in English and cover various stages and aspects of the recovery process, following established therapeutic guidelines and models. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/filippo19741974/Generated-Recovery-Support-Dialogues.question-answering1K<n<10K1 likes107 downloads1y agoHugging Face23SafwanAlbeshti /deepseek-v4-flash-filler-lens-demo DeepSeek-V4-Flash filler-token lens captures — demo subset Per-position logit-lens activations and top-k attention recorded from deepseek-ai/DeepSeek-V4-Flash on a three-product arithmetic task, with and without filler tokens. This is the public demo subset (7 captures) of a larger private collection. It exists so the attention viewer in the accompanying repo runs without special access. Code, full results and write-up: https://github.com/safwanalbeshti/filler-effect-writeup… See the full description on the dataset page: https://huggingface.co/datasets/SafwanAlbeshti/deepseek-v4-flash-filler-lens-demo.question-answeringn<1K0 likes98 downloads20d agoHugging Face24Chrisyichuan /screenshot-training-natural-filtered-v2 Chrisyichuan/screenshot-training-natural-filtered-v2 Wikipedia screenshot retrieval training dataset exported from local hard-negative mining. Contents train.jsonl / train_hn.jsonl eval.jsonl / eval_hn.jsonl test.jsonl / test_hn.jsonl train_hn_with_answer.jsonl / eval_hn_with_answer.jsonl / test_hn_with_answer.jsonl lite-query-v2-full-filtered-hn-with-answer.jsonl images/ Each metadata row has the form: { "query": "...", "chunk_path":… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/screenshot-training-natural-filtered-v2.question-answering10K<n<100K1 likes89 downloads6mo agoHugging Face25genevera /epstein-files-ocr-complete Epstein Files — Complete OCR Dataset This is a comprehensive, structured publication of the Epstein Files OCR dataset, significantly expanding upon the earlier Datasets 1-8 release. Dataset Summary This dataset contains page-level OCR output compiled from an extensive release of documents related to Jeffrey Epstein / the Epstein case. Each row in this dataset represents one scanned PDF document from the original release using a proprietary automated OCR pipeline… See the full description on the dataset page: https://huggingface.co/datasets/genevera/epstein-files-ocr-complete.textquestion-answering1M<n<10M0 likes89 downloads5mo agoHugging Face26DarkyMan /Opus-4.6-RU-Reasoning-creative-1385x-not-filtered Opus-4.6-RU-Creative-Writing — Russian Creative Writing Reasoning Dataset A Russian-language dataset of creative writing tasks generated with Claude claude-opus-4.6 (extended thinking enabled). Each sample contains a creative prompt, a full reasoning chain showing the creative process, and a detailed artistic response. Dataset Info Language: Russian 🇷🇺 Size: ~1,385 samples (growing) Model used: anthropic/claude-opus-4.6 with reasoning: {effort: "high"} Format:… See the full description on the dataset page: https://huggingface.co/datasets/DarkyMan/Opus-4.6-RU-Reasoning-creative-1385x-not-filtered.texttext-generation1K<n<10K3 likes88 downloads6mo agoHugging Face27sdiazlor /rag-human-rights-from-files Dataset Card for my-distiset-rag-files This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/my-distiset-rag-files/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/rag-human-rights-from-files.texttext-generationn<1K0 likes87 downloads2y agoHugging Face28jtatman /python-github-code-instruct-filtered-5k Dataset Card for "python-github-code-instruct-filtered-5k" This fine dataset tomekkorbak/python-github-code, filtered by scores greater than 0.03. Feedback and additional columns generated through OpenAI and Cohere responses. texttext-generation1K<n<10K7 likes85 downloads2y agoHugging Face29tmskss /eu-tenders-with-questions-for-agentic-checklist-filling eu-tenders-with-questions-for-agentic-checklist-filling Dataset Description This dataset contains questions and answers for evaluating Retrieval-Augmented Generation (RAG) systems in the context of generative agentic checklist-filling. The dataset is designed to benchmark various RAG architectures (Hybrid RAG, Graph RAG, Multi-Hop/Agentic RAG) on document analysis tasks. Dataset Summary Total Questions: 97 Document Families: 7 Languages: EN Domain: Procurement… See the full description on the dataset page: https://huggingface.co/datasets/tmskss/eu-tenders-with-questions-for-agentic-checklist-filling.documentquestion-answeringn<1K1 likes81 downloads5mo agoHugging Face30alex-apostolo /filtered-cuad Dataset Card for filtered_cuad Dataset Summary Contract Understanding Atticus Dataset (CUAD) v1 is a corpus of more than 13,000 labels in 510 commercial legal contracts that have been manually labeled to identify 41 categories of important clauses that lawyers look for when reviewing contracts in connection with corporate transactions. This dataset is a filtered version of CUAD. It excludes legal contracts with an Agreement date prior to 2002 and contracts which are not… See the full description on the dataset page: https://huggingface.co/datasets/alex-apostolo/filtered-cuad.textquestion-answering1K<n<10K4 likes78 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.