datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MegaMath
MegaMath: Pushing the Limits of Open Math Copora
Megamath is part of TxT360, curated by LLM360 Team.
We introduce MegaMath, an open math pretraining dataset curated from diverse, math-focused sources, with over 300B tokens.
MegaMath is curated via the following three efforts:
Revisiting web data:
We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on… See the full description on the dataset page: https://huggingface.co/datasets/IFM/MegaMath.VDR_MEGA_MultiDomain_DocRetrieval
Visual Document Retrieval Dataset
Overview
This dataset is designed for training visual document retrieval models. It combines multiple datasets from the VDR series, Colpali, and LlamaIndex to create the most comprehensive training resource for visual document retrieval tasks.
Dataset Structure
The dataset contains structured fields including unique identifiers with string lengths ranging from 45 to 50 characters, search query text with variable lengths between… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_MultiDomain_DocRetrieval.megalith-mdqa
Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM.
megalith-10mVDR_MEGA_2
VDR_MEGA_2
Dataset Summary
VDR_MEGA_2 is a high-quality multimodal dataset created through the merge of multiple domain-specific datasets with enhanced data processing techniques. This dataset represents our most refined approach to multimodal data generation, incorporating filtering algorithms and improved AI-assisted content generation to deliver superior quality for RAG, DSE, question answering, document search, and vision-language model training tasks.… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_2.MegaMath
MegaMath: Pushing the Limits of Open Math Copora
Megamath is part of TxT360, curated by LLM360 Team.
We introduce MegaMath, an open math pretraining dataset curated from diverse, math-focused sources, with over 300B tokens.
MegaMath is curated via the following three efforts:
Revisiting web data:
We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on… See the full description on the dataset page: https://huggingface.co/datasets/NP235/MegaMath.Mega-Brain-Distill
Mega-Brain-Distill
Curated merge of the top 10% highest-scoring examples from
584 community-uploaded LLM distillation/reasoning-trace datasets
on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces,
etc.), deduplicated within and across all of them — many of these source
repos are the same underlying dump re-uploaded by different users.
Auto-generated by run.py — do not hand-edit, it will be overwritten on
the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.MegaStyle-1.4MDataset of MegaStyle and MegaStyle++.
MegaStyle-1.4M is a large-scale style dataset built through a scalable pipeline that leverages consistent text-to-image style mapping of Qwen-Image. It combines 170K curated style prompts with 400K content prompts to generate 1.4M high-quality images that share strong intra-style consistency while covering diverse fine-grained styles.
MegaStyle++-8M further scales up the style space through a hierarchical style definition. It covers 150K overall style… See the full description on the dataset page: https://huggingface.co/datasets/tencent/MegaStyle-1.4M.MegaScience
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
Code: https://github.com/GAIR-NLP/MegaScience
Project Page: https://huggingface.co/MegaScience
MegaScience is a large-scale mixture of high-quality open-source datasets consisting of 1.25 million instances. We first collect multiple public datasets, then conduct comprehensive ablation studies across different data selection methods to identify the optimal approach for each dataset, thereby… See the full description on the dataset page: https://huggingface.co/datasets/MegaScience/MegaScience.syntheory
Dataset Card for SynTheory
Dataset Summary
SynTheory is a synthetic dataset of music theory concepts, specifically rhythmic (tempos and time signatures) and tonal (notes, intervals, scales, chords, and chord progressions).
Each of these 7 concepts has its own config.
tempos consist of 161 total integer tempos (bpm) ranging from 50 BPM to 210 BPM (inclusive), 5 percussive instrument types (click_config_name), and 5 random start time offsets (offset_time).
time_signatures… See the full description on the dataset page: https://huggingface.co/datasets/meganwei/syntheory.MegaMath-Web-Pro-Max
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
The Curation of MegaMath-Web-Pro-Max
Step 1: Uniformly and randomly sample millions of documents from the MegaMath-Web corpus, stratified by publication year;
Step 2: Annotate them using Llama-3.1-70B-instruct with a scoring prompt from FineMath and prepare the seed data;
Step 3: Training a fasttext carefully with proper preprocessing;
Step 4: Filtering documents with a threshold (i.e., 0.4);
Step 5:… See the full description on the dataset page: https://huggingface.co/datasets/OctoThinker/MegaMath-Web-Pro-Max.TextbookReasoning
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
Dataset Description
Scientific reasoning is critical for developing AI scientists and supporting human researchers in advancing the frontiers of natural science discovery. However, the open-source community has primarily focused on mathematics and coding while neglecting the scientific domain, largely due to the absence of open, large-scale, high-quality, verifiable scientific reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MegaScience/TextbookReasoning.Medical-Reasoning-SFT-Mega
Medical-Reasoning-SFT-Mega
The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning.
Dataset Overview
Metric
Value
Total Samples
1,789,998 (after deduplication)
Total Tokens
~3.78 Billion
Content Tokens
~2.22 Billion
Reasoning Tokens
~1.56 Billion
Samples with Reasoning
1,789,764 (100.0%)
Unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Mega.megalith-10mmegalith-10m
🗿 Megalith-10m
What is Megalith-10m?
Megalith-10m is a dataset of ~10 million links to Flickr images that were categorized as "photo" with license info of:
No known copyright restrictions (Flickr commons), or
United States Government Work, or
Public Domain Dedication (CC0), or
Public Domain Mark
What's the intended use of Megalith-10m?
Megalith-10m is intended to contain only links to wholesome unedited uncopyrighted photographs - the sort of… See the full description on the dataset page: https://huggingface.co/datasets/madebyollin/megalith-10m.MegaTerminal
MegaTerminal
MegaTerminal is a Harbor-style terminal task-folder dataset assembled from the terminal task sources collected for task matching experiments.
The task folders are packed into parquet/ as a custom blob layout (megatask-folder-parquet-v1); each row stores the files of one task folder as binary blobs. Restore the on-disk tasks/<task-id>/ tree with python scripts/deparquetize_megaterminal.py --parquet-dir parquet --out-dir MegaTerminal-restored.
Each task lives under… See the full description on the dataset page: https://huggingface.co/datasets/di-zhang-fdu/MegaTerminal.MegaScale
Mega-scale experimental analysis of protein folding stability in biology and design
The full MegaScale dataset contains 1,841,285 thermodynamic folding stability measurements
using cDNA display proteolysis of natural and designed proteins. From these 776,298 high-quality folding
stabilities (dataset2) cover all single amino acid variants and selected double mutants of 331 natural
and 148 de novo designed protein domains 40–72 amino acids in length. Of these mutations, 607,839 have… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/MegaScale.MEGA-Bench
MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks [ICLR 2025]
🌐 Homepage | 🏆 Leaderboard | 🤗 Dataset | 🤗 Paper | 🔎 Visualiaztion | 📖 arXiv | GitHub
🔔 News
[2025-01]: Paper accepted by ICLR 2025.
[2024-10-18]: Initial release of the evaluation code on our Github repo.
[2024-10-14]: Paper released on arXiv.
❗❗ Data Information
We put the file path of images/videos in HF datasets. Please download the zipped data here.
We chose… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MEGA-Bench.Twin-2K-500-Mega-Study
Twin-2K-500-Mega-Study Dataset
GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study
To see more details for how to process these data, please refer to this GitHub repository.
This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants).
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.bn-asr-mega-open-datafiltered_MegaMathmegamath-synthmegadiff
Megadiff, a dataset of source code changes
If you use Megadiff, please cite the following technical report:
"Megadiff: A Dataset of 600k Java Source Code Changes Categorized by Diff Size". Technical Report 2108.04631, Arxiv; 2021.
@techreport{megadiff,
TITLE = {{Megadiff: A Dataset of 600k Java Source Code Changes Categorized by Diff Size}},
AUTHOR = {Martin Monperrus and Matias Martinez and He Ye and Fernanda Madeiral and Thomas Durieux and Zhongxing Yu},
URL =… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/megadiff.MasriAudio-Mega-v0megamath-web-prodaily-aqi
Daily AQI — US Air Quality
Sharing datasets helps agents analyze them — giving everyone the ability to make sense of complex data.
Official US EPA air quality data updated daily. All 6 NAAQS criteria pollutants, hourly granularity, ~1,500 monitoring stations across the United States.
Explore it interactively in the Daily AQI Space.
Dataset Structure
readings.parquet — hourly readings
One row per monitoring station per hour per pollutant.
Column… See the full description on the dataset page: https://huggingface.co/datasets/meganariley/daily-aqi.mega-ssum
Mega-SSum
A large-scale English sentence-wise speech summarization (Sen-SSum) dataset
Consists of 3.8M+ synthesized speech, transcription, summary triplets
Derived from the Gigaword dataset Rush+2015
Overview
The dataset is divided into five splits: train/core/dev/eval/duc2003. (See below table)
We added a new evaluation split "test" for in-domain evaluation.
The train split is here: MegaSSum(train).
orig. data
split
#samples
#speakers
total dur. (hrs)
ave.… See the full description on the dataset page: https://huggingface.co/datasets/komats/mega-ssum.MegaTerminal-Hard-1K
MegaTerminal-Hard-1K
The 1,000 highest-quality, hardest terminal-agent tasks selected from the 11,599 tasks in
di-zhang-fdu/MegaTerminal.
Every task is a complete, self-contained Harbor / Terminal-Bench task folder — instruction,
container definition, verifier, and reference solution — restorable byte-for-byte from the
Parquet blobs in parquet/.
Tasks
1,000 (top 9.8% of the deduplicated upstream pool)
Restored size
156 MB across 10,119 files
With reference… See the full description on the dataset page: https://huggingface.co/datasets/jwu323/MegaTerminal-Hard-1K.mega-moledit-large
MEGA: A Large-Scale Molecular Editing Dataset for Guided-Action Optimization
Large-scale annotated molecular editing dataset with 57M examplesfor training models to modify molecular structures based on natural language instructions.
Paper: MEGA: A Large-Scale Molecular Editing Dataset for Guided-Action Optimization
Official Repository: https://github.com/nfsrules/MEGA-moledit
Dataset Structure
Each example will contain:
task_id: Task identifier
prompt: Natural… See the full description on the dataset page: https://huggingface.co/datasets/nfsrulesFR/mega-moledit-large.megalith-cc0
Megalith-CC0
A CC0-filtered version of the Megalith-10m dataset. The images have also been persisted to an independent public S3 bucket, supported by the AWS Open Data Registry program, for durability.
Why filter by CC0?
The images in Megalith-10m, having been gathered from Flickr, have attached licenses of CC0 and public domain. However, it is not clear if users assigning the public domain license to their works understand the implications of the public domain… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/megalith-cc0.
