datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CodeFeedback-Filtered-Instruction OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
[🏠Homepage]
|
[🛠️Code]
OpenCodeInterpreter
OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities.
For further information and… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/CodeFeedback-Filtered-Instruction.epstein-files-ocr-datasets-1-8-early-release
Epstein Files OCR — Datasets 1–8 (Early Release)
ARCHIVE NOTICE
This dataset is no longer maintained. Please refer to the Epstein Files — Complete OCR Dataset.
Dataset Summary
This dataset contains page-level OCR output (as Markdown) from a public release of documents related to Jeffrey Epstein / the Epstein case.
Each Markdown file represents one scanned page converted to text using an automated OCR pipeline. The dataset is designed for:
Question answering
Information… See the full description on the dataset page: https://huggingface.co/datasets/ishumilin/epstein-files-ocr-datasets-1-8-early-release.Fable-5.1-Max-Reasoning-Filtered-10000x
Dataset Description
This dataset contains 10,000 agentic coding and reasoning multi-turn high-quality traces generated by the new Fable 5.1 model using max reasoning effort.
It holds almost 500,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains.
It has also been deduplicated and heavily filtered to remove low-quality traces, keeping only high-quality traces.
Dataset Statistics
Metric
Value
Total Examples
10,000… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/Fable-5.1-Max-Reasoning-Filtered-10000x.osworld_tasks_filesbird23-train-filtered
BIRD-SQL Train (Filtered)
A high-quality subset of the original BIRD train split for text-to-SQL finetuning.
Overview
Over the past year the community has shared many observations about data quality in BIRD. We performed a rigorous data quality check process to retain examples that are consistent with schema and faithfully answer the question. The resulting set keeps 6,601 instances out of 9,428 (≈70%), and serves as a drop-in replacement for training.
Original Train: 9… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird23-train-filtered.EDGAR_FILINGS_DATASET
SFD: SEC Filings Dataset (v1)
SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation.
This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in:
The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.epstein-files-ocr-datasets-1-8-early-release
Epstein Files OCR — Datasets 1–8 (Early Release)
Work in Progress (WIP)
This is an early publication. We are actively working on improving OCR quality and expanding coverage.
Dataset Summary
This dataset contains page-level OCR output (as Markdown) from a public release of documents related to Jeffrey Epstein / the Epstein case.
Each Markdown file represents one scanned page converted to text using an automated OCR pipeline. The dataset is designed for:
Question… See the full description on the dataset page: https://huggingface.co/datasets/aurora2424/epstein-files-ocr-datasets-1-8-early-release.FileGram
FileGram Dataset
Grounding Agent Personalization in File-System Behavioral Traces
Overview
FileGram is a comprehensive framework for evaluating memory-centric personalization from file-system behavioral traces. This dataset provides:
640 behavioral trajectories — 20 persona-driven profiles x 32 tasks (16 text-centric + 16 multimodal), each containing fine-grained file-system operation logs, content snapshots, and session statistics
4,333 QA pairsacross 4… See the full description on the dataset page: https://huggingface.co/datasets/Choiszt/FileGram.sec-filings-qa-instruct
SEC Filings Instruction-Tuning Dataset (Llama-3 Format)
This dataset contains 5,000 curated, instruction-formatted question-answering pairs derived from corporate SEC filings (Forms 10-K and 10-Q). It is structured specifically for parameter-efficient instruction fine-tuning (SFT/QLoRA) of Small Language Models using the standard Llama-3 ChatML template.
Dataset Details
Origin Source: Curated subset extracted from nvidia/Nemotron-SpecializedDomains-Finance-v1.… See the full description on the dataset page: https://huggingface.co/datasets/lateesha-bhatia/sec-filings-qa-instruct.Fable-5.1-Max-Reasoning-Filtered-1000x
Dataset Description
This dataset contains 1,000 coding and reasoning traces generated by the new Fable 5.1 model using max reasoning effort.
It holds almost 30,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains.
It has also been deduplicated and filtered to remove low-quality traces, keeping only high-quality traces.
Dataset Statistics
Metric
Value
Total Examples
1,000 Traces
Total Token Count
~30,000,000… See the full description on the dataset page: https://huggingface.co/datasets/KrazyKitty/Fable-5.1-Max-Reasoning-Filtered-1000x.swe-bench-multi-file-refactoring-sft-dpo-2026
💻 Enterprise Autonomous SWE-bench AI & Multi-File Code Refactoring SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step call-stack Chain-of-Thought (<thought>) reasoning trees for fine-tuning LLMs (Qwen-2.5-Coder, Llama-3.3, DeepSeek-R1-Distill, Mistral) into Autonomous Software Engineers and SWE-bench Benchmark Agents.
📊 Dataset Architecture & Highlights
Multi-Turn Code Reviews:… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/swe-bench-multi-file-refactoring-sft-dpo-2026.finsight-sec-filings
FinSight — SEC EDGAR Filings
Cleaned plain-text 10-K (annual) and 10-Q (quarterly) filings from the
US SEC EDGAR system for 20 large publicly-traded companies across 6 sectors.
Created as part of the FinSight project —
a financial research AI assistant combining BERT fine-tuning, RAG, and
multi-agent systems.
Stats
Records: 97
Companies: 20 (AAPL, MSFT, GOOGL, AMZN, META, NVDA, TSLA, JPM, BAC, GS,
JNJ, PFE, UNH, WMT, PG, KO, MCD, XOM, CVX, CAT)
Forms: 10-K, 10-Q… See the full description on the dataset page: https://huggingface.co/datasets/musk1209/finsight-sec-filings.EnvFactory-SFT-FILTERED
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
## Overview
EnvFactory-SFT-FILTERED is a filtered supervised fine-tuning (SFT) dataset containing 53,400 tool-use trajectories synthesized using the EnvFactory framework. This dataset is designed for SFT training of tool-use agents.
The dataset contains high-quality multi-turn tool-use trajectories with implicit human reasoning, generated through… See the full description on the dataset page: https://huggingface.co/datasets/LARK-Lab/EnvFactory-SFT-FILTERED.stackoverflowVQA-filterednemotron_terminal_filtered
Nemotron Terminal Filtered
An uncertainty-curated subset of NVIDIA's Nemotron-Terminal-Corpus (dataset_adapters split), selected for high-formation density for post-training NVIDIA-Nemotron-3-Super-120B-A12B-BF16.
Motivation
The full dataset_adapters split contains ~226k terminal execution trajectories. To curate a compact, high-value subset for post-training we score each sample by how hard the model finds it, using entropy as a proxy for uncertainty. The… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/nemotron_terminal_filtered.wildchat-filtered
WildChat Filtered Dataset
This is a filtered version of the WildChat-4.8M dataset.
Dataset Description
This dataset contains 3,199,860 conversations between human users and ChatGPT, filtered to keep only the essential conversation structure.
Data Structure
Each conversation contains only:
conversations: A list of message objects with:
role: Either "user" or "assistant"
content: The text content of the message
All other metadata (timestamps, moderation… See the full description on the dataset page: https://huggingface.co/datasets/rayonlabs/wildchat-filtered.epstein-files-ocr-complete
Epstein Files — Complete OCR Dataset
This is a comprehensive, structured publication of the Epstein Files OCR dataset, significantly expanding upon the earlier Datasets 1-8 release.
Dataset Summary
This dataset contains page-level OCR output compiled from an extensive release of documents related to Jeffrey Epstein / the Epstein case.
Each row in this dataset represents one scanned PDF document from the original release using a proprietary automated OCR pipeline… See the full description on the dataset page: https://huggingface.co/datasets/ishumilin/epstein-files-ocr-complete.nemotron-terminal-file_operations
nemotron-terminal-file_operations
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "file_operations". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-file_operations.hitchcock-psycho-1960-film-dataset-transformed
Psycho → AI-Model Dataset (Transformed)
A thematic re-skin of the Psycho (1960) Q&A dataset into an original AI-model setting where the world is transformed into an AI/data-center environment.
Character names, actor names, objects, locations, production references, dates, and thematic elements are remapped to AI/ML concepts and modern technology.
File: psycho_dataset_transformed.jsonl
Format: JSONL — one JSON object per line
Schema: each line has prompt and completion string… See the full description on the dataset page: https://huggingface.co/datasets/antfr99/hitchcock-psycho-1960-film-dataset-transformed.commonsense_filtered
Dataset Summary
The commonsense reasoning tasks consist of 8 subtasks, each with predefined training and testing sets, as described by LLM-Adapters (Hu et al., 2023). The following table lists the details of each sub-dataset.
Train
Test
Information
BoolQ (Clark et al., 2019)
9427
3270
Question-answering dataset for yes/no questions
PIQA (Bisk et al., 2020)
16113
1838
Questions with two solutions requiring physical commonsense to answer
SIQA (Sap et al., 2019)
33410… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/commonsense_filtered.ResearchMath-Filtered
ResearchMath-Filtered
ResearchMath-Filtered is a quality-filtered collection of 129,927 long-form reasoning traces and
solutions for research-level mathematical problems, released alongside
ResearchMath-14k as part of the same
paper. It is a cleaned subset of
ResearchMath-Reasoning-194K:
each record holds a self-contained problem statement, a long chain-of-thought reasoning trace, and a
final response, with low-quality and non-solving generations removed.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/amphora/ResearchMath-Filtered.Generated-Recovery-Support-Dialogues
# Empathetic Conversations for Addiction Recovery Support Dataset
Dataset Description
This dataset contains synthetically generated conversational examples between a user discussing their addiction recovery journey and an AI assistant designed to be empathetic, supportive, non-judgmental, and encouraging. The conversations are in English and cover various stages and aspects of the recovery process, following established therapeutic guidelines and models.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/filippo19741974/Generated-Recovery-Support-Dialogues.deepseek-v4-flash-filler-lens-demo
DeepSeek-V4-Flash filler-token lens captures — demo subset
Per-position logit-lens activations and top-k attention recorded from
deepseek-ai/DeepSeek-V4-Flash on a three-product arithmetic task, with and without
filler tokens.
This is the public demo subset (7 captures) of a larger private collection. It exists
so the attention viewer in the accompanying repo runs without special access.
Code, full results and write-up: https://github.com/safwanalbeshti/filler-effect-writeup… See the full description on the dataset page: https://huggingface.co/datasets/SafwanAlbeshti/deepseek-v4-flash-filler-lens-demo.screenshot-training-natural-filtered-v2
Chrisyichuan/screenshot-training-natural-filtered-v2
Wikipedia screenshot retrieval training dataset exported from local hard-negative mining.
Contents
train.jsonl / train_hn.jsonl
eval.jsonl / eval_hn.jsonl
test.jsonl / test_hn.jsonl
train_hn_with_answer.jsonl / eval_hn_with_answer.jsonl / test_hn_with_answer.jsonl
lite-query-v2-full-filtered-hn-with-answer.jsonl
images/
Each metadata row has the form:
{
"query": "...",
"chunk_path":… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/screenshot-training-natural-filtered-v2.epstein-files-ocr-complete
Epstein Files — Complete OCR Dataset
This is a comprehensive, structured publication of the Epstein Files OCR dataset, significantly expanding upon the earlier Datasets 1-8 release.
Dataset Summary
This dataset contains page-level OCR output compiled from an extensive release of documents related to Jeffrey Epstein / the Epstein case.
Each row in this dataset represents one scanned PDF document from the original release using a proprietary automated OCR pipeline… See the full description on the dataset page: https://huggingface.co/datasets/genevera/epstein-files-ocr-complete.Opus-4.6-RU-Reasoning-creative-1385x-not-filtered
Opus-4.6-RU-Creative-Writing — Russian Creative Writing Reasoning Dataset
A Russian-language dataset of creative writing tasks generated with Claude claude-opus-4.6 (extended thinking enabled). Each sample contains a creative prompt, a full reasoning chain showing the creative process, and a detailed artistic response.
Dataset Info
Language: Russian 🇷🇺
Size: ~1,385 samples (growing)
Model used: anthropic/claude-opus-4.6 with reasoning: {effort: "high"}
Format:… See the full description on the dataset page: https://huggingface.co/datasets/DarkyMan/Opus-4.6-RU-Reasoning-creative-1385x-not-filtered.rag-human-rights-from-files
Dataset Card for my-distiset-rag-files
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/my-distiset-rag-files/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/rag-human-rights-from-files.python-github-code-instruct-filtered-5k
Dataset Card for "python-github-code-instruct-filtered-5k"
This fine dataset tomekkorbak/python-github-code, filtered by scores greater than 0.03.
Feedback and additional columns generated through OpenAI and Cohere responses.
eu-tenders-with-questions-for-agentic-checklist-filling
eu-tenders-with-questions-for-agentic-checklist-filling
Dataset Description
This dataset contains questions and answers for evaluating Retrieval-Augmented Generation (RAG) systems in the context of generative agentic checklist-filling. The dataset is designed to benchmark various RAG architectures (Hybrid RAG, Graph RAG, Multi-Hop/Agentic RAG) on document analysis tasks.
Dataset Summary
Total Questions: 97
Document Families: 7
Languages: EN
Domain: Procurement… See the full description on the dataset page: https://huggingface.co/datasets/tmskss/eu-tenders-with-questions-for-agentic-checklist-filling.filtered-cuad
Dataset Card for filtered_cuad
Dataset Summary
Contract Understanding Atticus Dataset (CUAD) v1 is a corpus of more than 13,000 labels in 510 commercial legal contracts that have been manually labeled to identify 41 categories of important clauses that lawyers look for when reviewing contracts in connection with corporate transactions. This dataset is a filtered version of CUAD. It excludes legal contracts with an Agreement date prior to 2002 and contracts which are not… See the full description on the dataset page: https://huggingface.co/datasets/alex-apostolo/filtered-cuad.
