datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-complete
arXiv Complete Corpus
A snapshot of arXiv's metadata, version history, submission files and rendered
documents. It covers 3,148,796 papers and includes file contents, paths, sizes
and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface;
files come from the GCS mirror, S3 source archives and direct PDF fetches.
This release holds a PDF for 99.47% of papers and 99.54% of versions reported
with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.Jailbreak_Complete_DS_labeledpexel-0808-complete-final-testGithub Page: https://github.com/UmiMarch/OpenVideo
license: cc-by-4.0
task_categories:
- video-text-to-text
size_categories:
- 100K<n<1M
Complete-FABLE.5-traces-2M
license: mit
pretty_name: Claude Library — Fable 5 · Opus · Sonnet
annotations_creators:
machine-generated
language:
en
language_creators:
found
machine-generated
multilinguality:
monolingual
size_categories:
10K<n<100K
task_categories:
text-generation
task_ids:
language-modeling
tags:
agent-traces
claude
claude-fable-5
claude-opus
claude-sonnet
chain-of-thought
tool-use
coding-agents
content-verified
maintained-mirror
deduplicated
parquet
configs:
config_name:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Complete-FABLE.5-traces-2M.wiki_book_corpus_complete_processed_bert_dataset
Dataset Card for "wiki_book_corpus_complete_processed_bert_dataset"
More Information needed
coco_captioning_complete_formatBenetech_PlotQa_DVQA_combined_matcha_completewiki_book_corpus_complete_raw_dataset
Dataset Card for "wiki_book_corpus_complete_raw_dataset"
More Information needed
vie-speech-corpus-completemeta-llama_Llama-3.1-8B-Instruct-jdgfct-Completenesswikipedia-paragraph-embeddings-en-gist-complete
Dataset Summary
Paragraph embeddings for every article in English Wikipedia (not the Simple English version).
Based on wikimedia/wikipedia, 20231101.en.
Embeddings were generated with avsolatorio/GIST-small-Embedding-v0
and are quantized to int8.
You can load the data with the following:
from datasets import load_dataset
ds = load_dataset(path="Abrak/wikipedia-paragraph-embeddings-en-gist-complete", data-dir="20231101.en")
Dataset Structure
The structure of the… See the full description on the dataset page: https://huggingface.co/datasets/Abrak/wikipedia-paragraph-embeddings-en-gist-complete.Graph-R1-dataset-complete
Graph-R1 Complete Dataset
This dataset contains the complete Graph-R1 graph reasoning dataset with all difficulty levels (1-5).
Files
train_graph_all_levels.parquet: Combined training data from all levels (with level column)
test_graph_and_math_all_levels.parquet: Combined test data from all levels (with level column)
test_graph_mixedsize_cleaned.parquet: Mixed size test data (with level='mixed')
Usage
import pandas as pd
# Load combined data… See the full description on the dataset page: https://huggingface.co/datasets/HKUST-DSAIL/Graph-R1-dataset-complete.nvidia_NVLM-D-72B-jdgfct-CompletenessNexusflow_Athene-70B-jdgfct-CompletenessLLaVA_Instruct_150K_complete_formatComplete-FABLE.5-traces-2M
Complete FABLE.5 Traces (2 Million Deduplicated Rows)
Comprehensive Agentic Coding & Frontier Reasoning Trajectory Corpus
Executive Summary
Solstice-AI/Complete-FABLE.5-traces-2M is a clean, fully deduplicated post-training dataset containing 2,006,487 high-entropy agentic coding and multi-step reasoning traces.
Originally curated following the closure of Fable and Mythos, this corpus synthesizes frontier agent execution patterns (including Claude… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Complete-FABLE.5-traces-2M.elliott-wave-market-data-complete
Elliott Wave Market Data - Complete (Quality Validated) ✅
Production-ready, quality-validated OHLCV market data for training Elliott Wave pattern recognition neural networks.
🎯 Key Features
Quality Validated: Rigorous data quality checks applied
Complete Coverage: 1,403 unique instruments
Multi-Timeframe: 1h, 4h, 1d, 1wk data
22,546,189 Total Data Points
Dataset Statistics
Timeframe
Rows
Tickers
1h
9,267,368
1,358
4h
2,849,583
1,358
1d
8,648… See the full description on the dataset page: https://huggingface.co/datasets/usamaahmedsh/elliott-wave-market-data-complete.epstein-files-ocr-complete
Epstein Files — Complete OCR Dataset
This is a comprehensive, structured publication of the Epstein Files OCR dataset, significantly expanding upon the earlier Datasets 1-8 release.
Dataset Summary
This dataset contains page-level OCR output compiled from an extensive release of documents related to Jeffrey Epstein / the Epstein case.
Each row in this dataset represents one scanned PDF document from the original release using a proprietary automated OCR pipeline… See the full description on the dataset page: https://huggingface.co/datasets/ishumilin/epstein-files-ocr-complete.shakespeare-complete-works
Shakespeare Complete Works Dataset
This dataset contains the complete works of William Shakespeare, including:
The Sonnets (154 sonnets)
Plays (Tragedies, Comedies, Histories)
Poems
Dataset Structure
Each entry contains:
work: The title of the work
section: Specific section (e.g., "Sonnet 1") if applicable
text: The actual text content
type: Type of work (sonnet, play, poem)
id: Unique identifier
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/r-three/shakespeare-complete-works.ChartQA_Benetech_PlotQa_DVQA_combined_matcha_completeluanti-complete-dataset
Complete Luanti Package Dataset
Description
Complete Luanti Package Dataset for Luanti (Minetest) expertise fine-tuning.
Dataset Information
Size: 2,592 entries
Format: Harmony format for LLM fine-tuning
Source: Luanti ContentDB package collection
Quality: Filtered and validated Luanti package metadata
Usage
from datasets import load_dataset
dataset = load_dataset("ToddLLM/luanti-complete-dataset")
print(dataset)
Schema
Each entry… See the full description on the dataset page: https://huggingface.co/datasets/ToddLLM/luanti-complete-dataset.dataset_complete_12POPE-completeQwen3-ASR-PostTrain-Complete-Medical-French-Fullosm-completeness-tilesComplete-FABLE.5-traces-2M
Complete FABLE.5 Traces 2M
Full FABLE.5 / Mythos corpus restored, with session-limit answer rows removed.
Dataset Viewer | Parquet | Raw JSONL.gz
This dataset is a post-closure compilation of all available FABLE.5 / Mythos trace datasets found on Hugging Face during the curation pass after the closure of Fable and Mythos. It is deduplicated at the normalized-row level and keeps row-level provenance through first_source_dataset, first_source_config, first_source_split… See the full description on the dataset page: https://huggingface.co/datasets/Quaxicron/Complete-FABLE.5-traces-2M.Methods2Test_CompleteContextcomplete_medical_symptom_datasetso100_rebel_completeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 21666,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/VoicAndrei/so100_rebel_complete.complete_sandwich_in_orderThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 81,
"total_frames": 75333,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:81"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/izuluaga/complete_sandwich_in_order.
