datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
banned-historical-archives
和谐历史档案馆数据集 - Banned Historical Archives Datasets
和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。
目录结构
banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步
raw # 原始文件
config # 配置文件
todo # 存放暂未录入网站的文件
部分报纸和图片资料存放在单独的仓库:
名称
地址
状态
参考消息
https://huggingface.co/datasets/banned-historical-archives/ckxx
未录入
人民日报
https://huggingface.co/datasets/banned-historical-archives/rmrb
已精选重要的文章录入
文汇报… See the full description on the dataset page: https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives.picbreeder-vlm-archive
Picbreeder-VLM Archive
Every image evolved by the swarm of vision-language-model "breeders" in
In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models
(GECCO 2026), together with the CPPN genomes that produced them, the agents' reasoning transcripts, the
lineage graphs, and the analysis artifacts behind the paper and blog.
The original Picbreeder (Secretan et al., 2008) let crowds of
humans collaboratively evolve images from
CPPN… See the full description on the dataset page: https://huggingface.co/datasets/picbreeder-vlm/picbreeder-vlm-archive.symile-m3
Dataset Card for Symile-M3
Symile-M3 is a multilingual dataset of (audio, image, text) samples. The dataset is specifically designed to test a model's ability to capture higher-order information between three distinct high-dimensional data types: by incorporating multiple languages, we construct a task where text and audio are both needed to predict the image, and where, importantly, neither text nor audio alone would suffice.
Paper: https://arxiv.org/abs/2411.01053
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/arsaporta/symile-m3.VaaniVAANI is an India-representative multi-modal multi-lingual dataset.
The current version (phase 1- 80 districts, phase 2- 85 districts) contains ~31278 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts, talking about 288K images covering 105 languages.
From this audio data, 2,122 hours of transcribed data(text) is available, spanning almost evenly across the 165 districts.
Project Vaani, by IISc, Bangalore and ARTPARK, is capturing the true diversity of India’s… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani.3d-arenaFor more information, visit the 3D Arena Space.
Inputs are sourced from iso3D.
To assist with easily running inputs, are input image URLs are provided in inputs.txt.
ArxivCap
Dataset Card for ArxivCap
Data Instances
Example-1 of single (image, caption) pairs
"......" stands for omitted parts.
{
'src': 'arXiv_src_2112_060/2112.08947',
'meta':
{
'meta_from_kaggle':
{
'journey': '',
'license': 'http://arxiv.org/licenses/nonexclusive-distrib/1.0/',
'categories': 'cs.ET'
},
'meta_from_s2':
{
'citationCount': 8… See the full description on the dataset page: https://huggingface.co/datasets/MMInstruction/ArxivCap.3d-front-ararankpartyworidatsushitaorewamotooshiegotachitomeikyuushinbuwomezasu
Bangumi Image Base of A-rank Party Wo Ridatsu Shita Ore Wa, Moto Oshiego-tachi To Meikyuu Shinbu Wo Mezasu.
This is the image base of bangumi A-Rank Party wo Ridatsu shita Ore wa, Moto Oshiego-tachi to Meikyuu Shinbu wo Mezasu., we detected 168 characters, 15579 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/arankpartyworidatsushitaorewamotooshiegotachitomeikyuushinbuwomezasu.banned-historical-archives
和谐历史档案馆数据集 - Banned Historical Archives Datasets
和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。
目录结构
banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步
raw # 原始文件
config # 配置文件
todo # 存放暂未录入网站的文件
部分报纸和图片资料存放在单独的仓库:
名称
地址
状态
参考消息
https://huggingface.co/datasets/banned-historical-archives/ckxx
未录入
人民日报
https://huggingface.co/datasets/banned-historical-archives/rmrb
已精选重要的文章录入
文汇报… See the full description on the dataset page: https://huggingface.co/datasets/Dragonegg2026/banned-historical-archives.CircuitSense
CircuitSense
This dataset is a comprehensive multimodal circuit question-answering benchmark designed to evaluate visual reasoning and problem-solving capabilities across three main domains: Perception, Analysis, and Design. The dataset contains structured question-answer pairs with accompanying visual content, targeting different engineering cognitive levels and reasoning tasks.
Dataset Structure
The dataset is organized into three primary folders, each containing… See the full description on the dataset page: https://huggingface.co/datasets/armanakbari4/CircuitSense.cola
COLA: Compose Objects Localized with Attributes
Self-contained Hugging Face port of the COLA benchmark from the paper
"How to adapt vision-language models to Compose Objects Localized with Attributes?".
📄 Paper: https://arxiv.org/abs/2305.03689
🌐 Project page: https://cs-people.bu.edu/array/research/cola/
💻 Original code & data: https://github.com/ArijitRay1993/COLA
This repository bundles the benchmark annotations as Parquet files and the referenced
images as regular files… See the full description on the dataset page: https://huggingface.co/datasets/array/cola.ASMR-Archive-Processed
ASMR-Archive-Processed (WIP)
Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended.
Work in Progress — expect breaking changes while the pipeline and data layout stabilize.
This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.MINT-1T-ArXiv
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.banana-vidorev3-synthetic-arms
Banana ViDoRe v3 Synthetic Arms
Domain-separated ViDoRe v3 synthetic training arms for finance and industrial adaptation.
The Hub dataset uses finance and industrial as dataset configs/subsets. Within each config, splits separate
vlm_in_batch, vlm_ocr_bm25, banana_fullpipe, and hybrid_vlm_ocr_bm25_banana_fullpipe.
Generated at: 2026-06-29T11:49:45.670912+00:00
Total JSONL rows across configs/splits: 151691.
Images are stored once per subset under… See the full description on the dataset page: https://huggingface.co/datasets/vkehfdl1/banana-vidorev3-synthetic-arms.celebA_spoof
Dataset Card for "celebA_spoof"
More Information needed
pega-hardarxiv-latex-5TThe dataset used for https://github.com/TIGER-AI-Lab/ScholarCopilot.
veo3-video-prompts
Veo 3 Video Generation Dataset
English | Português do Brasil
English
Summary
A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant.
Videos: 5,811
Input images: 1,354
Configurations: 6
Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.arifuretashokugyoudesekaisaikyouseason3
Bangumi Image Base of Arifureta Shokugyou De Sekai Saikyou Season 3
This is the image base of bangumi Arifureta Shokugyou de Sekai Saikyou Season 3, we detected 64 characters, 6786 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/arifuretashokugyoudesekaisaikyouseason3.ultimateids-project-artifactsMOVA_benchmark_for_arena
MOVA Benchmark for Arena
This is the benchmark used for the subjective arena experiments of MOVA (MOVA: Towards Scalable and Synchronized Video–Audio Generation). All prompts are rewritten by the workflow introduced in the paper.
Paper: MOVA: Towards Scalable and Synchronized Video–Audio Generation
Code: https://github.com/OpenMOVA/MOVA
Overview
The benchmark contains 732 samples in total, organized into two subsets:
Subset
Samples
MOVA-Bench
132… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuzhang-0212/MOVA_benchmark_for_arena.artist-styles
artist-styles
Static gallery of artist styles. Plain HTML/JS (index.html) reading from
data/artists.json and images/ — no build step, no dependencies.
Running with Docker
docker.sh runs the site in a python:3.12-slim container serving the project
directory with python3 scripts/server.py (a stdlib-only server: static files
plus a small favorites API). The container is named artist-styles, restarts
automatically (--restart=always), and serves on port 7803 by… See the full description on the dataset page: https://huggingface.co/datasets/jtreminio/artist-styles.laion-art-en-colorcanny
Dataset Card for "laion-art-en-colorcanny"
More Information needed
ArchWorldsmart-bin-detect
arudaev/smart-bin-detect
Training data for Smart Bin Recognition – a validator ("is there a bin?")
and an identifier ("which bin?"). The design lives in docs/04-ml-pipeline.md
in the project repo, which is private; the manifests here carry per-image
provenance and are the authoritative record of what this dataset contains.
Every image carries provenance: source, source URL, licence, region,
capture date, annotator where known, label origin (human / machine /
legacy /… See the full description on the dataset page: https://huggingface.co/datasets/arudaev/smart-bin-detect.stocksrelaion-art
Relaion Art - LLM-Annotated
Original Source
📌 Introduction
This dataset comprises images and annotations from the original Relaion Art Dataset.
Out of the 8M images, a subset of 3.66M images has been annotated with automatic methods (Image-text-to-text models).
Captions
The annotations include four annotation columns:
dense_caption: A dense annotation about the image
vqa: Visual Question-Answers related to the image. JSON dictionary embedded as a… See the full description on the dataset page: https://huggingface.co/datasets/Fhrozen/relaion-art.ledger-long-context-multi-kpi
the LEDGER Long-Context Multi-KPI extraction datasets and benchmarks.
OCR'd annual reports with ground-truth KPI values for financial information extraction benchmarking.
Dataset Description
This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values. It is designed for evaluating LLM-based financial information extraction, retrieval, and needle-in-a-haystack tasks.
Configs
Config
Reports… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-multi-kpi.family-archival-scans
Family archival scans
Page scans of primary-source archival records used in the bobpanil/family
genealogy project. Public and world-readable, so tools and agents without
credentials can fetch pages directly over plain HTTPS.
6,475 files across 50 archival units from five holding institutions, plus a
database-index folder and a screenshot folder. All records are pre-1943 and
concern people long deceased; no information about living individuals is
included.
Structure… See the full description on the dataset page: https://huggingface.co/datasets/bobpanil/family-archival-scans.
