datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MegaPairs-Standard
MegaPairs-Standard (Standardized Version)
Dataset Summary
This is a standardized, high-efficiency version of the JUNJIE99/MegaPairs dataset.
Why use this version?
The original dataset is distributed as a massive Tar archive containing millions of images, accompanied by a separate JSONL annotation file.
The Problem: Using the original format requires extracting terabytes of small files (which can exhaust disk inodes) or writing complex logic to read from archives. It… See the full description on the dataset page: https://huggingface.co/datasets/86Cao/MegaPairs-Standard.VDR_MEGA_MultiDomain_DocRetrieval
Visual Document Retrieval Dataset
Overview
This dataset is designed for training visual document retrieval models. It combines multiple datasets from the VDR series, Colpali, and LlamaIndex to create the most comprehensive training resource for visual document retrieval tasks.
Dataset Structure
The dataset contains structured fields including unique identifiers with string lengths ranging from 45 to 50 characters, search query text with variable lengths between… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_MultiDomain_DocRetrieval.megalith-mdqa
Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM.
Megadepthmegalith-10mVDR_MEGA_2
VDR_MEGA_2
Dataset Summary
VDR_MEGA_2 is a high-quality multimodal dataset created through the merge of multiple domain-specific datasets with enhanced data processing techniques. This dataset represents our most refined approach to multimodal data generation, incorporating filtering algorithms and improved AI-assisted content generation to deliver superior quality for RAG, DSE, question answering, document search, and vision-language model training tasks.… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_2.Mega60k
Mega60k: Chart Question Answering Dataset
Dataset Overview
A multimodal chart question answering dataset featuring charts in multiple formats (CSV, PNG, SVG) and degraded PNG images with components omission, occlusion, blurring, and rotation to enhance robustness evaluation.
Languages: English
Chart Type Distribution
Chart Type
Count
Chart Type
Count
Chart Type
Count
Area
200
Bar
200
Box
200
Bubble
200
Chord
200
Fill-bubble
200
Funnel
200… See the full description on the dataset page: https://huggingface.co/datasets/guodaosun/Mega60k.MegaStyle-1.4MDataset of MegaStyle and MegaStyle++.
MegaStyle-1.4M is a large-scale style dataset built through a scalable pipeline that leverages consistent text-to-image style mapping of Qwen-Image. It combines 170K curated style prompts with 400K content prompts to generate 1.4M high-quality images that share strong intra-style consistency while covering diverse fine-grained styles.
MegaStyle++-8M further scales up the style space through a hierarchical style definition. It covers 150K overall style… See the full description on the dataset page: https://huggingface.co/datasets/tencent/MegaStyle-1.4M.MegaDepth-Syn
MegaDepth-Syn Dataset
The MegaDepth-Syn Dataset is generated from the MegaDepth dataset
using our MINIMA data engine, which contains for extra 6 modalities: infrared, depth, event, normal, sketch, and paint.
Abstract
Image matching for both cross-view and cross-modality plays a critical role in multimodal perception. In practice, the
modality gap caused by different imaging systems/styles poses great challenges to the matching task. Existing works try
to extract… See the full description on the dataset page: https://huggingface.co/datasets/lsxi77777/MegaDepth-Syn.megatonkyuumusashi
Bangumi Image Base of Megaton-kyuu Musashi
This is the image base of bangumi Megaton-kyuu Musashi, we detected 81 characters, 5660 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/megatonkyuumusashi.MegaSynth-webdatasetmegalith-10mmegalith-10m
🗿 Megalith-10m
What is Megalith-10m?
Megalith-10m is a dataset of ~10 million links to Flickr images that were categorized as "photo" with license info of:
No known copyright restrictions (Flickr commons), or
United States Government Work, or
Public Domain Dedication (CC0), or
Public Domain Mark
What's the intended use of Megalith-10m?
Megalith-10m is intended to contain only links to wholesome unedited uncopyrighted photographs - the sort of… See the full description on the dataset page: https://huggingface.co/datasets/madebyollin/megalith-10m.megamiryounoryoubokun
Bangumi Image Base of Megami-ryou No Ryoubo-kun.
This is the image base of bangumi Megami-ryou no Ryoubo-kun., we detected 50 characters, 4096 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/megamiryounoryoubokun.Rustins_Super_Mega_Awesome_VEDU_Model
Rustin's Super Mega Awesome VEDU Model
A reproducible, heavily-documented pipeline that maps Ventenata dubia ("VEDU", an invasive
winter-annual grass) across Montana from satellite + environmental data.
Science reference: docs/VEDU_48_predictors_detailed.md
Data decisions & gotchas: docs/CONTRADICTIONS.md
Parity with the Earth Engine build: docs/GEE_PARITY.md
Continue-the-build guide: docs/HANDOFF.md
Label inventory: docs/DATA_SOURCES.md
What it produces
57… See the full description on the dataset page: https://huggingface.co/datasets/UniversityOfMontanaSAL/Rustins_Super_Mega_Awesome_VEDU_Model.MEGA-Bench
MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks [ICLR 2025]
🌐 Homepage | 🏆 Leaderboard | 🤗 Dataset | 🤗 Paper | 🔎 Visualiaztion | 📖 arXiv | GitHub
🔔 News
[2025-01]: Paper accepted by ICLR 2025.
[2024-10-18]: Initial release of the evaluation code on our Github repo.
[2024-10-14]: Paper released on arXiv.
❗❗ Data Information
We put the file path of images/videos in HF datasets. Please download the zipped data here.
We chose… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MEGA-Bench.MegaStyle-1.4MDataset of MegaStyle. MegaStyle-1.4M is a large-scale style dataset built through a scalable pipeline that leverages consistent text-to-image style mapping of Qwen-Image. It combines 170K curated style prompts with 400K content prompts to generate 1.4M high-quality images that share strong intra-style consistency while covering diverse fine-grained styles.
All assets and code are under the license unless specified otherwise.
If this work is helpful for your research, please consider citing the… See the full description on the dataset page: https://huggingface.co/datasets/Adelacici/MegaStyle-1.4M.Megalith-10M-Camera
Megalith-10M-Camera
Per-image camera parameter annotations for the Megalith-10M dataset
(the image-bearing build drawthingsai/megalith-10m; ~9.58M Flickr photos across
959 shards), captioned by the Puffin-World model. More captioned datasets are provided in our Puffin-16M website.
The collage above visualizes the camera maps on sample images — each
pair shows the up field (green arrows: the projected gravity-up direction)
and the latitude field (colored contours: angle… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/Megalith-10M-Camera.megalith-cc0
Megalith-CC0
A CC0-filtered version of the Megalith-10m dataset. The images have also been persisted to an independent public S3 bucket, supported by the AWS Open Data Registry program, for durability.
Why filter by CC0?
The images in Megalith-10m, having been gathered from Flickr, have attached licenses of CC0 and public domain. However, it is not clear if users assigning the public domain license to their works understand the implications of the public domain… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/megalith-cc0.MegaDepth-v1mega_1mThis dataset MegaSG contains 1M images annotated with scene graphs , which was introduced in the paper What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation
nighthawk-mega
Nighthawk Mega — 1.72M Captioned UAV Aerial Images Across 6 Conditions
Every drone image, in every condition, fully captioned. The first dataset of its kind.
1,718,541 captioned images · 5 synthesized adverse conditions · 1.4M YOLO labels · 5 trained translation models · One reproducible pipeline.
Same drone, six modalities: original RGB, night, dusk, fog, rain, thermal — every image captioned.
Why this exists
UAV computer vision has a deployment problem.
Models… See the full description on the dataset page: https://huggingface.co/datasets/robotflowlabs/nighthawk-mega.megalith-qa-resizedMEGA-Benchflickr-megalith-10m-internvl2-multi-caption
Dataset Card for flickr-megalith-10m-internvl2-multi-caption
Dataset Summary
This is approximately 57.3 million synthetic captions for the images found in madebyollin/megalith-10m.
It includes the following captions:
InternVL2 8B long captions (by CaptionEmporium)
InternVL2 8B short captions (by CaptionEmporium)
Florence2 long captions (by aipicasso)
Florence2 short captions (by CaptionEmporium)
ShareCaptioner long captions (by drawthingsai)
ShareCaptioner short… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/flickr-megalith-10m-internvl2-multi-caption.SpanishGovernmentMultimodalQA
Spanish Government Multimodal QA
This dataset consists of official Curriculim Vitaes and images of 683 members of the Spanish Government obtained from https://transparencia.gob.es. Additionally, 323 questions regarding properties of the CVs and images of each Government member are included. For example, one question is:
{
"question_es": "¿Quién tiene una chaqueta de color verde en su foto de perfil y reporta que trabaja en el Ministerio de Economía, Comercio y Empresa?"… See the full description on the dataset page: https://huggingface.co/datasets/megaelius/SpanishGovernmentMultimodalQA.MegaKPT
MegaKPT: A Large-Scale High-Quality GKD Dataset
Project
GitHub: https://github.com/AlanLuSun/General-Keypoint-Detection
1. Introduction
MegaKPT dataset unifies 29 public keypoint datasets into the same annotation format, namely, COCO format, resulting over 1.3 million object instances. Moreover, we correct noisy annotations, supplement accurate keypoint texts, and give clear super-categories and indexes, rendering a high-quality and… See the full description on the dataset page: https://huggingface.co/datasets/changshenglu/MegaKPT.Continual-MEGA-Benchmark
Continual-MEGA: A Large-scale Benchmark for Generalizable Continual Anomaly Detection
This repository contains the dataset for Continual-MEGA, a new benchmark for continual learning in anomaly detection, introduced in the paper Continual-MEGA: A Large-scale Benchmark for Generalizable Continual Anomaly Detection.
Continual-MEGA aims to better reflect real-world deployment scenarios. It features a large and diverse dataset that significantly expands existing evaluation settings by… See the full description on the dataset page: https://huggingface.co/datasets/Continual-Mega/Continual-MEGA-Benchmark.mega_nerf_rubble_colmapThis repository redistributes the Rubble dataset released by Mega-NeRF. Their original release uses a different data format, so we reran COLMAP and provide a COLMAP-compatible version of Rubble here.
Please cite the original Mega-NeRF paper if you use this dataset:
@InProceedings{Turki_2022_CVPR,
author = {Turki, Haithem and Ramanan, Deva and Satyanarayanan, Mahadev},
title = {Mega-NERF: Scalable Construction of Large-Scale NeRFs for Virtual Fly-Throughs}… See the full description on the dataset page: https://huggingface.co/datasets/HexuZhao/mega_nerf_rubble_colmap.us-ghost-towns-photos
