datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagesMMStar
MMStar (Are We on the Right Way for Evaluating Large Vision-Language Models?)
🌐 Homepage | 🤗 Dataset | 🤗 Paper | 📖 arXiv | GitHub
Dataset Details
As shown in the figure below, existing benchmarks lack consideration of the vision dependency of evaluation samples and potential data leakage from LLMs' and LVLMs' training data.
Therefore, we introduce MMStar: an elite vision-indispensible multi-modal benchmark, aiming to ensure each curated sample exhibits… See the full description on the dataset page: https://huggingface.co/datasets/Lin-Chen/MMStar.D3HROwnerAnime-LineArt-Dataset
Anime Lineart Sketch Dataset
Dataset Summary
This dataset contains high-quality lineart sketch images automatically extracted from raw anime images using the LineartAnimeDetector model from ControlNet Annotators (lllyasviel/Annotators).
It is designed to support research and development in:
Anime-style sketch generation
Text-to-sketch pipelines
ControlNet conditioning
Sketch-to-image and image-to-sketch translation
The raw source images are sourced from the Anime Images… See the full description on the dataset page: https://huggingface.co/datasets/ityizNola/Anime-LineArt-Dataset.UniML3D
UniML3D
UniML3D is the text-paired, topology-annotated motion dataset behind UniMate (SIGGRAPH Asia 2026): motion clips from three sources with very different skeletons — Mixamo humanoids, Truebones ZOO animals and rigged Objaverse-XL objects — brought into one canonical layout, captioned, and annotated with cleaned joint names, a body-plan category and a facing-direction joint pair per skeleton. Every annotation in it was generated by this project's own data… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/UniML3D.IAM-line
IAM - line level
Dataset Summary
The IAM Handwriting Database contains forms of handwritten English text which can be used to train and test handwritten text recognizers and to perform writer identification and verification experiments.
Note that all images are resized to a fixed height of 128 pixels.
Languages
All the documents in the dataset are written in English.
Dataset Structure
Data Instances
{
'image':… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/IAM-line.UAV3DCrop
UAV3DCrop
UAV3DCrop is a multi-year UAV crop dataset containing field imagery and
associated reconstruction metadata for agricultural research.
Project page: UAV3DCrop
Paper: UAV3DCrop: Benchmarking 3D Reconstruction in Repeated Multi-Angle UAV Crop Surveys
Code: GitHub
Companion depth dataset: UAV3DCrop Depth
Repository Contents
The dataset is organized by year and acquisition day. A scene may include:
RGB images in an images/ directory;
camera parameters and… See the full description on the dataset page: https://huggingface.co/datasets/Link-Dev/UAV3DCrop.UNS-SSDS-2025Dataset from UNS SSDS 2025 Competition
spada-dataset
SPADA Dataset
This dataset contains images and sparse labels used in the paper Land Cover Segmentation with Sparse Annotations from Sentinel-2 Imagery
, published at IGARSS 2023.
Repository: https://github.com/links-ads/igarss-spada
Paper: https://paperswithcode.com/paper/land-cover-segmentation-with-sparse
Dataset Preparation
The dataset has been compressed into segmented tarballs for ease of use within Git LFS (that is, tar > gzip > split).
To revert the process… See the full description on the dataset page: https://huggingface.co/datasets/links-ads/spada-dataset.Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset
Description
This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations.
It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.lingbot-depth-subset
Dataset Card for lingbot-depth-subset
This is a FiftyOne dataset with 13,149 samples
(10,207 groups) spanning 3 sub-collections (RobbyReal, RobbyVla, RobbySim).
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/lingbot-depth-subset.OCRBench_v2UAV3DCrop_depth
UAV3DCrop Depth
UAV3DCrop Depth contains per-image depth maps associated with the UAV3DCrop
multi-year agricultural UAV dataset.
Project page: UAV3DCrop
Paper: UAV3DCrop: Benchmarking 3D Reconstruction in Repeated Multi-Angle UAV Crop Surveys
Code: UAV3DCrop GitHub
RGB dataset: UAV3DCrop
Repository Contents
The depth data are organized by year and acquisition scene. Scene directories
contain TIFF depth maps named to correspond to images in the main UAV3DCrop… See the full description on the dataset page: https://huggingface.co/datasets/Link-Dev/UAV3DCrop_depth.ocr-pl-lines
ocr-pl-lines
Syntetyczny zbiór linii tekstu po polsku do fine-tuningu OCR (TrOCR).
Pary NNNNN.png (obraz linii) + NNNNN.txt (transkrypcja).
Struktura
train/ — 2000 par (seed 42)
val/ — 200 par (seed 123)
Generowanie
OCR_engine —
python -m training.generate_synthetic
Korpus: zdania potoczne i urzędowe, domeny (faktury, umowy, medyczne,
prawnicze), losowe daty/kwoty/adresy/NIP/PESEL, zdania z pl.wikipedia.org.
Augmentacje: pochylenie, blur, szum… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/ocr-pl-lines.MDPBench
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
We introduce Multilingual Document Parsing Benchmark, the first benchmark for multilingual digital and photographed document parsing. Document parsing has made remarkable strides, yet almost exclusively on clean, digital, well-formatted pages in a handful of dominant languages. No systematic benchmark exists to evaluate how models perform on digital and photographed documents across diverse scripts and… See the full description on the dataset page: https://huggingface.co/datasets/Delores-Lin/MDPBench.WebCompass
WebCompass
A unified multimodal benchmark for evaluating LLMs' ability to generate, edit, and repair functional web pages. WebCompass spans three input modalities — text design documents, reference screenshots, and video demonstrations — and three task families — generation, editing, and repair.
GitHub: NJU-LINK/WebCompass
Project Page: nju-link.github.io/WebCompass
Quick Start
from datasets import load_dataset
# Generation tasks (existing)
ds_text =… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/WebCompass.3DHarnessBench
Probing Agentic 3D-to-Code Capabilities of Frontier Vision-Language Models
Project Page
·
GitHub
·
Arxiv
Abstract
3DHarnessBench evaluates the agentic capacity of frontier vision-language models (VLMs) to recover 3D geometry as executable Blender Python code from multiple forms of target evidence. Rather than restricting every system to a single fixed input, the benchmark compares four progressively richer harnesses: Single-view, Multi-view, Active… See the full description on the dataset page: https://huggingface.co/datasets/lingada/3DHarnessBench.genhome3d-1280
GenHome3D-1280
1,280 validated household and spatial-design assets in USDZ format, organized
across 64 categories.
Explore the visual catalog ·
Browse the GitHub repository ·
Download the versioned release ·
Read the generation method
Dataset summary
Assets
1,280
Categories
64
Assets per category
20
Runtime format
USDZ
Units
Meters
Asset license
CC BY 4.0
Technical validation
1,280/1,280 pass
Package validation
1… See the full description on the dataset page: https://huggingface.co/datasets/linxy97/genhome3d-1280.tifa_train_v3telugu-synthetic-line-imagesPELLET-Casimir-Marius-line
PELLET Casimir Marius - Line level
Dataset Summary
The PELLET Casimir Marius dataset includes 100 annotated French letters written between 1914 and 1918.
Annotations were done at line-level and all images do not have any text.
Note that all images are resized to a fixed height of 128 pixels.
Languages
All the documents in the dataset are written in French.
Dataset Structure
Data Instances
{
'image': <PIL.JpegImagePlugin.JpegImageFile… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/PELLET-Casimir-Marius-line.wealth-of-nationsai2thor-sideview-onlyLaTeX_OCR
LaTeX OCR 的数据仓库
本数据仓库是专为 LaTeX_OCR 及 LaTeX_OCR_PRO 制作的数据,来源于 https://zenodo.org/record/56198#.V2p0KTXT6eA 以及 https://www.isical.ac.in/~crohme/ 以及我们自己构建。
如果这个数据仓库有帮助到你的话,请点亮 ❤️like ++
后续追加新的数据也会放在这个仓库 ~~
原始数据仓库在github LinXueyuanStdio/Data-for-LaTeX_OCR.
数据集
本仓库有 5 个数据集
small 是小数据集,样本数 110 条,用于测试
full 是印刷体约 100k 的完整数据集。实际上样本数略小于 100k,因为用 LaTeX 的抽象语法树剔除了很多不能渲染的 LaTeX。
synthetic_handwrite 是手写体 100k 的完整数据集,基于 full 的公式,使用手写字体合成而来,可以视为人类在纸上的手写体。样本数实际上略小于 100k,理由同上。… See the full description on the dataset page: https://huggingface.co/datasets/linxy/LaTeX_OCR.lingyuCASIA-HWDB2-line
CASIA-HWDB2 - line level
Dataset Summary
The offline Chinese handwriting database (CASIA-HWDB2) was built by the National Laboratory of Pattern Recognition (NLPR), Institute of Automation of Chinese Academy of Sciences (CASIA).
The handwritten samples were produced by 1,020 writers using Anoto pen on papers, such that both online and offline data were obtained.
Note that all images are resized to a fixed height of 128 pixels.
Languages
All the documents in the… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/CASIA-HWDB2-line.mnist-cleaned-full
Dataset Card for 2025.11.21.16.40.44.970939
This is a FiftyOne dataset with 69807 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Linus-L/mnist-cleaned-full")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Linus-L/mnist-cleaned-full.lingbot-map-demoYCB_Video_Dataset
