datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MegaMath
MegaMath: Pushing the Limits of Open Math Copora
Megamath is part of TxT360, curated by LLM360 Team.
We introduce MegaMath, an open math pretraining dataset curated from diverse, math-focused sources, with over 300B tokens.
MegaMath is curated via the following three efforts:
Revisiting web data:
We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on… See the full description on the dataset page: https://huggingface.co/datasets/IFM/MegaMath.megawikaMegaWika is a multi- and crosslingual text dataset containing 30 million
Wikipedia passages with their scraped and cleaned web citations. The
passages span 50 Wikipedias in 50 languages, and the articles in which
the passages were originally embedded are included for convenience. Where
a Wikipedia passage is in a non-English language, an automated English
translation is provided. Furthermore, nearly 130 million English
question/answer pairs were extracted from the passages, and FrameNet events
occurring in the passages are detected using the LOME FrameNet parser.MegaPairs-Standard
MegaPairs-Standard (Standardized Version)
Dataset Summary
This is a standardized, high-efficiency version of the JUNJIE99/MegaPairs dataset.
Why use this version?
The original dataset is distributed as a massive Tar archive containing millions of images, accompanied by a separate JSONL annotation file.
The Problem: Using the original format requires extracting terabytes of small files (which can exhaust disk inodes) or writing complex logic to read from archives. It… See the full description on the dataset page: https://huggingface.co/datasets/86Cao/MegaPairs-Standard.VDR_MEGA_MultiDomain_DocRetrieval
Visual Document Retrieval Dataset
Overview
This dataset is designed for training visual document retrieval models. It combines multiple datasets from the VDR series, Colpali, and LlamaIndex to create the most comprehensive training resource for visual document retrieval tasks.
Dataset Structure
The dataset contains structured fields including unique identifiers with string lengths ranging from 45 to 50 characters, search query text with variable lengths between… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_MultiDomain_DocRetrieval.MegaSynthmegalith-mdqa
Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM.
megaposeMegadepthmegalith-10mVDR_MEGA_2
VDR_MEGA_2
Dataset Summary
VDR_MEGA_2 is a high-quality multimodal dataset created through the merge of multiple domain-specific datasets with enhanced data processing techniques. This dataset represents our most refined approach to multimodal data generation, incorporating filtering algorithms and improved AI-assisted content generation to deliver superior quality for RAG, DSE, question answering, document search, and vision-language model training tasks.… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_2.MegaMath
MegaMath: Pushing the Limits of Open Math Copora
Megamath is part of TxT360, curated by LLM360 Team.
We introduce MegaMath, an open math pretraining dataset curated from diverse, math-focused sources, with over 300B tokens.
MegaMath is curated via the following three efforts:
Revisiting web data:
We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on… See the full description on the dataset page: https://huggingface.co/datasets/NP235/MegaMath.megawika-2
MegaWika 2
MegaWika 2 is an improved multilingual text dataset containing a structured view of Wikipedia articles, the web sources they cite, source text quality estimates, article text translations, and additional article enrichments.
Note: Web citations (sources) in the HuggingFace dataset do not include scraped source text; use rehydrate-citations.py to rehydrate them.
The initial data release is based on Wikipedia dumps from May 1, 2024.
In total, the data contains about 77… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/megawika-2.kernelbench-mega-traces
KernelBench-Mega agent traces
Coding agents writing full GPU megakernels across Blackwell / H100 / B200, scored as speedup over reference; contamination-audited (23 verified cells).
Each .jsonl file is one agent run in Claude-Code session format, viewable with the agent trace viewer. Filename = run id; manifest.csv maps each run to model / harness / problem / GPU / score.
23 agent traces · live leaderboard: https://kernelbench.com/mega
Secrets redacted. Full reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-mega-traces.Mega60k
Mega60k: Chart Question Answering Dataset
Dataset Overview
A multimodal chart question answering dataset featuring charts in multiple formats (CSV, PNG, SVG) and degraded PNG images with components omission, occlusion, blurring, and rotation to enhance robustness evaluation.
Languages: English
Chart Type Distribution
Chart Type
Count
Chart Type
Count
Chart Type
Count
Area
200
Bar
200
Box
200
Bubble
200
Chord
200
Fill-bubble
200
Funnel
200… See the full description on the dataset page: https://huggingface.co/datasets/guodaosun/Mega60k.Mega-Brain-Distill
Mega-Brain-Distill
Curated merge of the top 10% highest-scoring examples from
584 community-uploaded LLM distillation/reasoning-trace datasets
on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces,
etc.), deduplicated within and across all of them — many of these source
repos are the same underlying dump re-uploaded by different users.
Auto-generated by run.py — do not hand-edit, it will be overwritten on
the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.MegaMath-Web-Pro-Max
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
The Curation of MegaMath-Web-Pro-Max
Step 1: Uniformly and randomly sample millions of documents from the MegaMath-Web corpus, stratified by publication year;
Step 2: Annotate them using Llama-3.1-70B-instruct with a scoring prompt from FineMath and prepare the seed data;
Step 3: Training a fasttext carefully with proper preprocessing;
Step 4: Filtering documents with a threshold (i.e., 0.4);
Step 5:… See the full description on the dataset page: https://huggingface.co/datasets/OctoThinker/MegaMath-Web-Pro-Max.MegaStyle-1.4MDataset of MegaStyle and MegaStyle++.
MegaStyle-1.4M is a large-scale style dataset built through a scalable pipeline that leverages consistent text-to-image style mapping of Qwen-Image. It combines 170K curated style prompts with 400K content prompts to generate 1.4M high-quality images that share strong intra-style consistency while covering diverse fine-grained styles.
MegaStyle++-8M further scales up the style space through a hierarchical style definition. It covers 150K overall style… See the full description on the dataset page: https://huggingface.co/datasets/tencent/MegaStyle-1.4M.MegaScience
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
Code: https://github.com/GAIR-NLP/MegaScience
Project Page: https://huggingface.co/MegaScience
MegaScience is a large-scale mixture of high-quality open-source datasets consisting of 1.25 million instances. We first collect multiple public datasets, then conduct comprehensive ablation studies across different data selection methods to identify the optimal approach for each dataset, thereby… See the full description on the dataset page: https://huggingface.co/datasets/MegaScience/MegaScience.highresolution-laioncoco-aesthetic-MEGThis dataset is filtered from laioncoco-aesthetic, which is used for academic research on mobile edge generation (MEG).
It includes high-resolution 1024-by-1024 text-to-image samples generated by a distilled SDXL with 4-12 denoising steps.
The dataset mainly involves the following fields:
caption: The text prompt of the image.
image: The target image corresponding to the prompt.
diffusion: The generative results of the distilled SDXL.
latents: The latent features of the distilled SDXL.
3D-data-megatronsyntheory
Dataset Card for SynTheory
Dataset Summary
SynTheory is a synthetic dataset of music theory concepts, specifically rhythmic (tempos and time signatures) and tonal (notes, intervals, scales, chords, and chord progressions).
Each of these 7 concepts has its own config.
tempos consist of 161 total integer tempos (bpm) ranging from 50 BPM to 210 BPM (inclusive), 5 percussive instrument types (click_config_name), and 5 random start time offsets (offset_time).
time_signatures… See the full description on the dataset page: https://huggingface.co/datasets/meganwei/syntheory.python-megatronhephaestus-ccx-runs-megarepomiomio_cp1_cachethings-meg
THINGS-MEG
This dataset is a processed version of THINGS-MEG, derived from the paper Bridging the Vision-Brain Gap with an Uncertainty-Aware Blur Prior (CVPR 2025). In this version, the MEG data is stored in float16 format, reducing the storage size by half. The original official dataset can be accessed from the OSF repository.
Original official dataset:
THINGS-data, a multimodal collection of large-scale datasets for investigating object representations in human brain and… See the full description on the dataset page: https://huggingface.co/datasets/Haitao999/things-meg.TextbookReasoning
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
Dataset Description
Scientific reasoning is critical for developing AI scientists and supporting human researchers in advancing the frontiers of natural science discovery. However, the open-source community has primarily focused on mathematics and coding while neglecting the scientific domain, largely due to the absence of open, large-scale, high-quality, verifiable scientific reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MegaScience/TextbookReasoning.Medical-Reasoning-SFT-Mega
Medical-Reasoning-SFT-Mega
The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning.
Dataset Overview
Metric
Value
Total Samples
1,789,998 (after deduplication)
Total Tokens
~3.78 Billion
Content Tokens
~2.22 Billion
Reasoning Tokens
~1.56 Billion
Samples with Reasoning
1,789,764 (100.0%)
Unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Mega.MegaDepth-Syn
MegaDepth-Syn Dataset
The MegaDepth-Syn Dataset is generated from the MegaDepth dataset
using our MINIMA data engine, which contains for extra 6 modalities: infrared, depth, event, normal, sketch, and paint.
Abstract
Image matching for both cross-view and cross-modality plays a critical role in multimodal perception. In practice, the
modality gap caused by different imaging systems/styles poses great challenges to the matching task. Existing works try
to extract… See the full description on the dataset page: https://huggingface.co/datasets/lsxi77777/MegaDepth-Syn.megaminocafeterrace
Bangumi Image Base of Megami No Café Terrace
This is the image base of bangumi Megami no Café Terrace, we detected 80 characters, 8688 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/megaminocafeterrace.megatonkyuumusashi
Bangumi Image Base of Megaton-kyuu Musashi
This is the image base of bangumi Megaton-kyuu Musashi, we detected 81 characters, 5660 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/megatonkyuumusashi.
