datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-complete
arXiv Complete Corpus
A snapshot of arXiv's metadata, version history, submission files and rendered
documents. It covers 3,148,796 papers and includes file contents, paths, sizes
and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface;
files come from the GCS mirror, S3 source archives and direct PDF fetches.
This release holds a PDF for 99.47% of papers and 99.54% of versions reported
with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.Complete_Data_Source_100K_HOURS
Multi-Language Audio Collection (100K Hours)
This repository is physically reorganized for Absolute 100% Data Visibility.
🏗️ Global Consolidator
Select your language subset to listen to high-quality waveform audio. All shards from legacy and modern pipelines are automatically routed here.
Jailbreak_Complete_DS_labeledpexel-0808-complete-final-testGithub Page: https://github.com/UmiMarch/OpenVideo
license: cc-by-4.0
task_categories:
- video-text-to-text
size_categories:
- 100K<n<1M
Complete-FABLE.5-traces-2M
license: mit
pretty_name: Claude Library — Fable 5 · Opus · Sonnet
annotations_creators:
machine-generated
language:
en
language_creators:
found
machine-generated
multilinguality:
monolingual
size_categories:
10K<n<100K
task_categories:
text-generation
task_ids:
language-modeling
tags:
agent-traces
claude
claude-fable-5
claude-opus
claude-sonnet
chain-of-thought
tool-use
coding-agents
content-verified
maintained-mirror
deduplicated
parquet
configs:
config_name:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Complete-FABLE.5-traces-2M.tvelve_map_complete_datasetzcache-results-completebooksum-complete-cleaned
Description:
This repository contains the Booksum dataset introduced in the paper BookSum: A Collection of Datasets for Long-form Narrative Summarization
.
This dataset includes both book and chapter summaries from the BookSum dataset (unlike the kmfoda/booksum one which only contains the chapter dataset). Some mismatched summaries have been corrected. Uneccessary columns have been discarded. Contains minimal text-to-summary rows. As there are multiple summaries for a given text… See the full description on the dataset page: https://huggingface.co/datasets/ubaada/booksum-complete-cleaned.coco_captioning_complete_formatleetcode-complete
Complete LeetCode Problems Dataset
This dataset contains a comprehensive collection of LeetCode problems (including premium) with AI-generated solutions in JSONL format. It is regularly updated to include new problems as they are added to LeetCode.
Splits
The dataset is divided into the following splits:
train: Contains approximately 80% of the problems for training
validation: Contains approximately 10% of the problems for validation
test: Contains approximately… See the full description on the dataset page: https://huggingface.co/datasets/whiskwhite/leetcode-complete.Benetech_PlotQa_DVQA_combined_matcha_completeISEAR-dataset-completewiki_book_corpus_complete_raw_dataset
Dataset Card for "wiki_book_corpus_complete_raw_dataset"
More Information needed
vie-speech-corpus-completemeta-llama_Llama-3.1-8B-Instruct-jdgfct-Completenesswikipedia-paragraph-embeddings-en-gist-complete
Dataset Summary
Paragraph embeddings for every article in English Wikipedia (not the Simple English version).
Based on wikimedia/wikipedia, 20231101.en.
Embeddings were generated with avsolatorio/GIST-small-Embedding-v0
and are quantized to int8.
You can load the data with the following:
from datasets import load_dataset
ds = load_dataset(path="Abrak/wikipedia-paragraph-embeddings-en-gist-complete", data-dir="20231101.en")
Dataset Structure
The structure of the… See the full description on the dataset page: https://huggingface.co/datasets/Abrak/wikipedia-paragraph-embeddings-en-gist-complete.complete-voiceai-speech-dataset
Silencio Voice AI Sample Dataset
Speaker-attributed spontaneous speech. 363 labelled contributors across 111 self-reported origin varieties, recorded on their own devices in their own environments. Every clip carries origin region / variety, mother tongue, gender, device, OS, recording environment.
Hours
16.05
Clips
1,305
Speakers
363
Origin varieties
111
Languages
21
Configs
44
Speaker metadata
origin region / variety, mother tongue, gender, device, OS… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/complete-voiceai-speech-dataset.thinking-cap-tier-curricula-complete
Thinking Cap Tier Curricula — Complete Reasoning Alignment Suite (TCS v4)
[!IMPORTANT]
Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture:
In v1.2, all 13,477 SFT samples and 3,187 SimPO preference pairs have undergone an automated token purge:
Zero <|pad|> batch residues: 100% eliminated across all files.
Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct, non-repetitive… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-curricula-complete.Graph-R1-dataset-complete
Graph-R1 Complete Dataset
This dataset contains the complete Graph-R1 graph reasoning dataset with all difficulty levels (1-5).
Files
train_graph_all_levels.parquet: Combined training data from all levels (with level column)
test_graph_and_math_all_levels.parquet: Combined test data from all levels (with level column)
test_graph_mixedsize_cleaned.parquet: Mixed size test data (with level='mixed')
Usage
import pandas as pd
# Load combined data… See the full description on the dataset page: https://huggingface.co/datasets/HKUST-DSAIL/Graph-R1-dataset-complete.whittle-teacher32-complete-answers
Whittle teacher32: complete answers with per-token teacher logprobs
Research preview. Part of the Whittle compression campaign, a personal
research project. The compute for this project is self funded and donations
decide whether the next round happens: https://ko-fi.com/davida81328
What this is
Complete answers generated by Qwen3.8-27B (UD-Q5_K_XL via llama.cpp), each
ending on a real end-of-turn token because the answer is finished, with the
teacher's top-32… See the full description on the dataset page: https://huggingface.co/datasets/logic65/whittle-teacher32-complete-answers.nvidia_NVLM-D-72B-jdgfct-CompletenessDeltarune-Complete-Transcript-Cleaned
Deltarune Chapters 1–4 Dataset
Fan-made transcript dataset covering Deltarune Chapters 1 through 4. Processed from video playthroughs and cross-referenced with game data. Intended to provide LLMs with structured narrative context for a game whose content is underrepresented in training corpora.
Why This Exists
As of early 2026, major LLMs (including models with training cutoffs past July 2025) fail to recall basic plot details of Deltarune Chapters 3 and 4 despite their… See the full description on the dataset page: https://huggingface.co/datasets/Deltarunefan/Deltarune-Complete-Transcript-Cleaned.LLaVA_Instruct_150K_complete_formatNexusflow_Athene-70B-jdgfct-Completenessmixed-language-detection-pilot-complete-sentences
Mixed-Language Speech Detection Pilot — Complete Sentences
This is the complete-sentence revision of a 6,000-clip binary
audio-classification pilot. label = 0 denotes one intended language and
label = 1 denotes more than one intended language. The covered languages are
Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani
(ckb), Arabic (ara), Persian (fas), and English (eng).
What changed
Earlier generation forced source transcripts into arbitrary… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-complete-sentences.Complete-FABLE.5-traces-2M
Complete FABLE.5 Traces (2 Million Deduplicated Rows)
Comprehensive Agentic Coding & Frontier Reasoning Trajectory Corpus
Executive Summary
Solstice-AI/Complete-FABLE.5-traces-2M is a clean, fully deduplicated post-training dataset containing 2,006,487 high-entropy agentic coding and multi-step reasoning traces.
Originally curated following the closure of Fable and Mythos, this corpus synthesizes frontier agent execution patterns (including Claude… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Complete-FABLE.5-traces-2M.elliott-wave-market-data-complete
Elliott Wave Market Data - Complete (Quality Validated) ✅
Production-ready, quality-validated OHLCV market data for training Elliott Wave pattern recognition neural networks.
🎯 Key Features
Quality Validated: Rigorous data quality checks applied
Complete Coverage: 1,403 unique instruments
Multi-Timeframe: 1h, 4h, 1d, 1wk data
22,546,189 Total Data Points
Dataset Statistics
Timeframe
Rows
Tickers
1h
9,267,368
1,358
4h
2,849,583
1,358
1d
8,648… See the full description on the dataset page: https://huggingface.co/datasets/usamaahmedsh/elliott-wave-market-data-complete.vipe_fliter_completeepstein-files-ocr-complete
Epstein Files — Complete OCR Dataset
This is a comprehensive, structured publication of the Epstein Files OCR dataset, significantly expanding upon the earlier Datasets 1-8 release.
Dataset Summary
This dataset contains page-level OCR output compiled from an extensive release of documents related to Jeffrey Epstein / the Epstein case.
Each row in this dataset represents one scanned PDF document from the original release using a proprietary automated OCR pipeline… See the full description on the dataset page: https://huggingface.co/datasets/ishumilin/epstein-files-ocr-complete.ChartQA_Benetech_PlotQa_DVQA_combined_matcha_complete
