datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MegaMath
MegaMath: Pushing the Limits of Open Math Copora
Megamath is part of TxT360, curated by LLM360 Team.
We introduce MegaMath, an open math pretraining dataset curated from diverse, math-focused sources, with over 300B tokens.
MegaMath is curated via the following three efforts:
Revisiting web data:
We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on… See the full description on the dataset page: https://huggingface.co/datasets/IFM/MegaMath.MegaPairs-Standard
MegaPairs-Standard (Standardized Version)
Dataset Summary
This is a standardized, high-efficiency version of the JUNJIE99/MegaPairs dataset.
Why use this version?
The original dataset is distributed as a massive Tar archive containing millions of images, accompanied by a separate JSONL annotation file.
The Problem: Using the original format requires extracting terabytes of small files (which can exhaust disk inodes) or writing complex logic to read from archives. It… See the full description on the dataset page: https://huggingface.co/datasets/86Cao/MegaPairs-Standard.VDR_MEGA_MultiDomain_DocRetrieval
Visual Document Retrieval Dataset
Overview
This dataset is designed for training visual document retrieval models. It combines multiple datasets from the VDR series, Colpali, and LlamaIndex to create the most comprehensive training resource for visual document retrieval tasks.
Dataset Structure
The dataset contains structured fields including unique identifiers with string lengths ranging from 45 to 50 characters, search query text with variable lengths between… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_MultiDomain_DocRetrieval.megalith-mdqa
Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM.
megaposemegalith-10mVDR_MEGA_2
VDR_MEGA_2
Dataset Summary
VDR_MEGA_2 is a high-quality multimodal dataset created through the merge of multiple domain-specific datasets with enhanced data processing techniques. This dataset represents our most refined approach to multimodal data generation, incorporating filtering algorithms and improved AI-assisted content generation to deliver superior quality for RAG, DSE, question answering, document search, and vision-language model training tasks.… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_2.MegaMath
MegaMath: Pushing the Limits of Open Math Copora
Megamath is part of TxT360, curated by LLM360 Team.
We introduce MegaMath, an open math pretraining dataset curated from diverse, math-focused sources, with over 300B tokens.
MegaMath is curated via the following three efforts:
Revisiting web data:
We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on… See the full description on the dataset page: https://huggingface.co/datasets/NP235/MegaMath.kernelbench-mega-traces
KernelBench-Mega agent traces
Coding agents writing full GPU megakernels across Blackwell / H100 / B200, scored as speedup over reference; contamination-audited (23 verified cells).
Each .jsonl file is one agent run in Claude-Code session format, viewable with the agent trace viewer. Filename = run id; manifest.csv maps each run to model / harness / problem / GPU / score.
23 agent traces · live leaderboard: https://kernelbench.com/mega
Secrets redacted. Full reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-mega-traces.Mega-Brain-Distill
Mega-Brain-Distill
Curated merge of the top 10% highest-scoring examples from
584 community-uploaded LLM distillation/reasoning-trace datasets
on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces,
etc.), deduplicated within and across all of them — many of these source
repos are the same underlying dump re-uploaded by different users.
Auto-generated by run.py — do not hand-edit, it will be overwritten on
the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.MegaMath-Web-Pro-Max
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
The Curation of MegaMath-Web-Pro-Max
Step 1: Uniformly and randomly sample millions of documents from the MegaMath-Web corpus, stratified by publication year;
Step 2: Annotate them using Llama-3.1-70B-instruct with a scoring prompt from FineMath and prepare the seed data;
Step 3: Training a fasttext carefully with proper preprocessing;
Step 4: Filtering documents with a threshold (i.e., 0.4);
Step 5:… See the full description on the dataset page: https://huggingface.co/datasets/OctoThinker/MegaMath-Web-Pro-Max.MegaStyle-1.4MDataset of MegaStyle and MegaStyle++.
MegaStyle-1.4M is a large-scale style dataset built through a scalable pipeline that leverages consistent text-to-image style mapping of Qwen-Image. It combines 170K curated style prompts with 400K content prompts to generate 1.4M high-quality images that share strong intra-style consistency while covering diverse fine-grained styles.
MegaStyle++-8M further scales up the style space through a hierarchical style definition. It covers 150K overall style… See the full description on the dataset page: https://huggingface.co/datasets/tencent/MegaStyle-1.4M.MegaScience
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
Code: https://github.com/GAIR-NLP/MegaScience
Project Page: https://huggingface.co/MegaScience
MegaScience is a large-scale mixture of high-quality open-source datasets consisting of 1.25 million instances. We first collect multiple public datasets, then conduct comprehensive ablation studies across different data selection methods to identify the optimal approach for each dataset, thereby… See the full description on the dataset page: https://huggingface.co/datasets/MegaScience/MegaScience.syntheory
Dataset Card for SynTheory
Dataset Summary
SynTheory is a synthetic dataset of music theory concepts, specifically rhythmic (tempos and time signatures) and tonal (notes, intervals, scales, chords, and chord progressions).
Each of these 7 concepts has its own config.
tempos consist of 161 total integer tempos (bpm) ranging from 50 BPM to 210 BPM (inclusive), 5 percussive instrument types (click_config_name), and 5 random start time offsets (offset_time).
time_signatures… See the full description on the dataset page: https://huggingface.co/datasets/meganwei/syntheory.TextbookReasoning
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
Dataset Description
Scientific reasoning is critical for developing AI scientists and supporting human researchers in advancing the frontiers of natural science discovery. However, the open-source community has primarily focused on mathematics and coding while neglecting the scientific domain, largely due to the absence of open, large-scale, high-quality, verifiable scientific reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MegaScience/TextbookReasoning.Medical-Reasoning-SFT-Mega
Medical-Reasoning-SFT-Mega
The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning.
Dataset Overview
Metric
Value
Total Samples
1,789,998 (after deduplication)
Total Tokens
~3.78 Billion
Content Tokens
~2.22 Billion
Reasoning Tokens
~1.56 Billion
Samples with Reasoning
1,789,764 (100.0%)
Unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Mega.quran-asr-mega-corpusMegaSynth-webdatasetmegalith-10mmegawika-report-generation
Dataset Card for MegaWika for Report Generation
Dataset Summary
MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span
50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a
non-English language, an automated English translation is provided.
This dataset provides the… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/megawika-report-generation.megalith-10m
🗿 Megalith-10m
What is Megalith-10m?
Megalith-10m is a dataset of ~10 million links to Flickr images that were categorized as "photo" with license info of:
No known copyright restrictions (Flickr commons), or
United States Government Work, or
Public Domain Dedication (CC0), or
Public Domain Mark
What's the intended use of Megalith-10m?
Megalith-10m is intended to contain only links to wholesome unedited uncopyrighted photographs - the sort of… See the full description on the dataset page: https://huggingface.co/datasets/madebyollin/megalith-10m.megatron-prof-data-v5cypherbench
CypherBench
CypherBench is a benchmark designed to evaluate text-to-Cypher translation for large language models (LLMs). It includes:
11 large-scale Neo4j property graphs transformed from Wikidata with 7.8 million entities.
Over 10,000 (question, Cypher) pairs for training/evaluating text-to-Cypher translation.
Paper: https://arxiv.org/pdf/2412.18702
Repository & Demo: https://github.com/megagonlabs/cypherbench
Contact: yanlin@megagon.ai
Sample Task
{
"qid":… See the full description on the dataset page: https://huggingface.co/datasets/megagonlabs/cypherbench.megamiryounoryoubokun
Bangumi Image Base of Megami-ryou No Ryoubo-kun.
This is the image base of bangumi Megami-ryou no Ryoubo-kun., we detected 50 characters, 4096 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/megamiryounoryoubokun.MegaTerminal
MegaTerminal
MegaTerminal is a Harbor-style terminal task-folder dataset assembled from the terminal task sources collected for task matching experiments.
The task folders are packed into parquet/ as a custom blob layout (megatask-folder-parquet-v1); each row stores the files of one task folder as binary blobs. Restore the on-disk tasks/<task-id>/ tree with python scripts/deparquetize_megaterminal.py --parquet-dir parquet --out-dir MegaTerminal-restored.
Each task lives under… See the full description on the dataset page: https://huggingface.co/datasets/di-zhang-fdu/MegaTerminal.mega-acceptability-v2MegaDepth-X
MegaDepth-X
MegaDepth-X is a large-scale Internet photo reconstruction dataset for long-tail 3D reconstruction from Internet photos. For more details, please visit the project page.
Paper
Title: Long-tail Internet Photo ReconstructionAuthors: Yuan Li, Yuanbo Xiangli, Hadar Averbuch-Elor, Noah Snavely, Ruojin CaiPaper: https://arxiv.org/abs/2604.22714Project page: https://megadepth-x.github.io/
Abstract
Internet photo collections are highly… See the full description on the dataset page: https://huggingface.co/datasets/y-u-a-n-l-i/MegaDepth-X.megatron-prof-data-v6MegaScale
Mega-scale experimental analysis of protein folding stability in biology and design
The full MegaScale dataset contains 1,841,285 thermodynamic folding stability measurements
using cDNA display proteolysis of natural and designed proteins. From these 776,298 high-quality folding
stabilities (dataset2) cover all single amino acid variants and selected double mutants of 331 natural
and 148 de novo designed protein domains 40–72 amino acids in length. Of these mutations, 607,839 have… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/MegaScale.Rustins_Super_Mega_Awesome_VEDU_Model
Rustin's Super Mega Awesome VEDU Model
A reproducible, heavily-documented pipeline that maps Ventenata dubia ("VEDU", an invasive
winter-annual grass) across Montana from satellite + environmental data.
Science reference: docs/VEDU_48_predictors_detailed.md
Data decisions & gotchas: docs/CONTRADICTIONS.md
Parity with the Earth Engine build: docs/GEE_PARITY.md
Continue-the-build guide: docs/HANDOFF.md
Label inventory: docs/DATA_SOURCES.md
What it produces
57… See the full description on the dataset page: https://huggingface.co/datasets/UniversityOfMontanaSAL/Rustins_Super_Mega_Awesome_VEDU_Model.
