datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
food101
Dataset Card for Food-101
Dataset Summary
This dataset consists of 101 food categories, with 101'000 images. For each class, 250 manually reviewed test images are provided as well as 750 training images. On purpose, the training images were not cleaned, and thus still contain some amount of noise. This comes mostly in the form of intense colors and sometimes wrong labels. All images were rescaled to have a maximum side length of 512 pixels.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/ethz/food101.PanoInfinigen🗃️ PanoInfinigen Dataset
PanoInfinigen is a synthetic dataset of high-resolution panoramic images in ERP, featuring perfectly aligned RGB, Depth, and Surface Normals. This dataset was generated using a modified Infinigen framework to support wide-angle panoramic geometry, plus the iCity procedural city generator for the urban split.
It serves as the primary training data for PaGeR, a single-step diffusion model for zero-shot panoramic depth… See the full description on the dataset page: https://huggingface.co/datasets/prs-eth/PanoInfinigen.Mem-Gallery
📖 Overview
Mem-Gallery is a comprehensive benchmark dataset designed to evaluate multimodal long-term memory capabilities of MLLM agents across multi-session conversations. The dataset features realistic, persona-driven dialogues spanning 20 scenarios, each enriched with contextual images to test memory retention, recall, and reasoning over extended interactions.
🎯 Key Features
Diverse Scenarios: Covering topics from AI & Robotics to Daily Life… See the full description on the dataset page: https://huggingface.co/datasets/Ethan-Bei/Mem-Gallery.stable-bias-generationsllava-665k
LLaVA-v1.5 Mix665K — Arrow (images embedded)
The LLaVA-v1.5 visual instruction-tuning mixture (llava_v1_5_mix665k) converted to a 🤗 datasets
Arrow dataset with image bytes embedded. 665,298 examples
(624,610 image–text + 40,688 text-only).
⚠️ This repo is a raw save_to_disk Arrow snapshot. The dataset viewer and
load_dataset() do not work here — load it with load_from_disk as shown below.
Loading
from huggingface_hub import snapshot_download
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/Ethlake/llava-665k.ECCV_Event_Video_Depth_Estimation
Event-Guided Video Depth Estimation Workshop Dataset
This dataset is a mirrored and aligned workshop-ready version of the DVD event-guided video depth estimation data.
It packages each scene into a canonical folder tree that aligns:
low-light RGB frames
per-frame event slices
a scene-level lowlight_event.npz
the matched depth ground truth copied from inference_results/*/normal/depth.npz
The dataset is designed for direct upload to Hugging Face as a dataset repository.
The official… See the full description on the dataset page: https://huggingface.co/datasets/Ethanliang99/ECCV_Event_Video_Depth_Estimation.eth3d_omnivggtsynfintabs
SynFinTabs: Synthetic Financial Tables
A dataset of synthetic financial tables for information extraction and table extraction tasks.
Quick Start
To get started with SynFinTabs, load it like any other Hugging Face dataset.
>>> from datasets import load_dataset
>>>
>>> synfintabs_dataset = load_dataset("ethanbradley/synfintabs")
Table Annotations
Table annotations are stored as a list of rows; a row contains a list of cells; and a cell contains a list of… See the full description on the dataset page: https://huggingface.co/datasets/ethanbradley/synfintabs.all-ethereum-contracts
All ethereum contracts
This dataset contains all deployed Ethereum contracts as of block 21850000 (February 15th, 2025), bytecodes of the contracts, and the block numbers the contracts were deployed.
Contract bytecodes are stored as a hash of the bytecode, and another dataset is provided mapping bytecode hashes to bytecodes. This is to reduce the size of the dataset, as many contracts have identical bytecodes.
This dataset was exported from a PostgreSQL database into CSV format.… See the full description on the dataset page: https://huggingface.co/datasets/Zellic/all-ethereum-contracts.z-image-ethnicity-test
Z-Image Turbo Ethnicity Benchmarking Dataset
Overview
This dataset was created to evaluate and test the Z-Image Turbo model's capabilities in accurately rendering various ethnicities. It comprises photorealistic portrait prompts designed to cover a diverse range of ethnic groups and demographic attributes.
Generation Methodology
The prompts in this dataset were synthetically generated using the Mistral-Small-3.2-24B-Instructmodel. The… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/z-image-ethnicity-test.vision-flan
Vision-FLAN (191-task) — Arrow (images embedded)
The Vision-FLAN vision-flan_191-task_1k visual instruction-tuning
set converted to a 🤗 datasets Arrow dataset with image bytes embedded. 186,103 examples
spanning ~191 human-labeled vision tasks.
⚠️ This repo is a raw save_to_disk Arrow snapshot. The dataset viewer and
load_dataset() do not work here — load it with load_from_disk as shown below.
Loading
from huggingface_hub import snapshot_download
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/Ethlake/vision-flan.phenotype-catalog
Ethnic Erotic Phenotype Catalog
A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations.
Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research.
What's in v6
Two columns… See the full description on the dataset page: https://huggingface.co/datasets/EthnicErotic/phenotype-catalog.Genso-teki_Na_Furemu_Ethereal_Framesethiopian-legal-ocr-v1-clean
ethiopian-legal-ocr-v1-clean
This dataset was created using the Claude Dataset Skill.
Continuous-Ethnicity-Face-Recognition
Dataset Card for Ethnicity Fairness in a Continuous Space
Dataset Details
Dataset Description
This dataset provides the training images (from BalancedFace and GlobalFace produced by BUPT) used in the paper: "Balancing Beyond Discrete Categories: Continuous Demographic Labels for Fair Face Recognition".
These have been curated to be balanced in a continuous ethnicity space, following three different strategies: Protocol A, Protocol B and Protocol C.… See the full description on the dataset page: https://huggingface.co/datasets/netopedro/Continuous-Ethnicity-Face-Recognition.stable-bias-professions
Dataset Card for "stable-bias-professions"
More Information needed
tracebench
TraceBench
This dataset contains the public benchmark data, synchronized aggregate results, trajectories, website data, and accepted submission artifacts for TraceBench.
The matching evaluation and reproduction code is available at TommasoBendinelli/TraceBench.
Layout
questions/BallDrop/, questions/BounceBall/, and questions/MassSlide/ contain the three canonical benchmark environments.
results.csv and results.parquet contain the same 420 aggregate result rows… See the full description on the dataset page: https://huggingface.co/datasets/eth-siplab/tracebench.eval_steering_ours_low_4_same_noiseThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "koch_follower",
"total_episodes": 20,
"total_frames": 4259,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ethanCSL/eval_steering_ours_low_4_same_noise.TR-DocVQA-Synth
TR-DocVQA-Synth: A Large-Scale Synthetic Turkish Document Visual Question Answering Dataset
Dataset Summary
TR-DocVQA-Synth is a large-scale synthetic Turkish Document Visual Question Answering dataset designed for training and evaluating multimodal models on Turkish business documents. The dataset contains 15,000 document images and 235,000 question-answer pairs generated from structured ground-truth records.
The dataset focuses on realistic Turkish document… See the full description on the dataset page: https://huggingface.co/datasets/Ethosoft/TR-DocVQA-Synth.RefSeg-CA
RefSeg-CA: command compliance and conservative repair
Research artifacts for RefSeg-CA: evaluating and repairing command compliance in generalized referring segmentation. This is an unpublished R3 manuscript release; independent human parser validation has not yet been conducted.
Contents
Artifact
Size / meaning
Procedural scenes
1,200 unique scene geometries, five seeds
Rendered images
2,400: matched flat and rich versions
Commands
28,800… See the full description on the dataset page: https://huggingface.co/datasets/Ethosoft/RefSeg-CA.et_handwriting_segmentation
Dataset of Text Region and Line Coordinates in Handwritten Estonian Documents
Dataset Description
This dataset contains coordinate annotations for Estonian historical documents, extracted from Transkribus exports. It includes full page images with precise coordinate information for text regions and text lines, designed for text detection, layout analysis, and document structure understanding tasks.
📊 Dataset Summary
Total Examples: 7,664 images
Language: 🇪🇪… See the full description on the dataset page: https://huggingface.co/datasets/Rahvusarhiiv/et_handwriting_segmentation.ZuriPano🗃️ ZüriPano Dataset
ZüriPano is a real-world outdoor panoramic depth benchmark, captured with the
Leica RTC360
LiDAR scanner (8K capture, 130 m effective range, HDR + automated double-scan
for transient-occlusion removal). It contains 100 equirectangular panoramas
across 11 urban locations in Zürich, each paired with a dense metric depth map
and a validity mask. It is used as the outdoor evaluation benchmark for
PaGeR.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/prs-eth/ZuriPano.Mip-NeRF_360_Processed_with_NerfstudioADautoGen-DS
ADautoGen-DS: Multi-Modal Product Advertisement Dataset
A synthetic dataset of 120 product advertisements with AI-generated text and images, designed for multi-modal recommendation systems research.
Dataset Overview
Attribute
Value
Total Samples
120 products
Categories
5 (Food, Tech, Fitness, Beauty, Home)
Target Audiences
5 unique segments
Text Generation Models
3 (Phi-2, Qwen-1.8B, TinyLlama)
Image Generation
Stable Diffusion v1.5
Embedding Model… See the full description on the dataset page: https://huggingface.co/datasets/EthanGabis/ADautoGen-DS.et_handwriting_recognition
Dataset of Transcribed Text Lines in Handwritten Estonian Documents
Dataset Description
This dataset contains 245,975 annotated text regions from Estonian historical documents, designed for optical character recognition (OCR) and document analysis tasks. The dataset includes images of text regions along with their corresponding transcriptions, and source documents.
📊 Dataset Summary
Total Examples: 245,975
Language: 🇪🇪 Estonian
Dataset Size: ~38.9 GB
Task:… See the full description on the dataset page: https://huggingface.co/datasets/Rahvusarhiiv/et_handwriting_recognition.EthereumStatistics
Ethereum Conflict Graph Metrics
This repository contains two CSV files, one for prestateTracer and one for callTracer, each covering over 2 million recent Ethereum blocks. These files provide key metrics extracted from conflict graphs generated by their respective tracers.
You can read about these tracers here:
https://geth.ethereum.org/docs/developers/evm-tracing/built-in-tracers
Dataset Overview
prestateTracer: Contains metrics for RW conflict graphs generated… See the full description on the dataset page: https://huggingface.co/datasets/dbiton/EthereumStatistics.FLIP-Challenge
FLIP Reasoning Challenge Dataset
This repository contains the FLIP dataset, a benchmark for evaluating AI reasoning capabilities based on human verification tasks from the Idena blockchain. The dataset focuses on testing sequential reasoning, visual storytelling, and common sense understanding in multimodal AI systems.
Paper: https://arxiv.org/abs/2504.12256.
Dataset Description
FLIP challenges present users with two orderings (stacks) of 4 images, requiring them to… See the full description on the dataset page: https://huggingface.co/datasets/aplesner-eth/FLIP-Challenge.et_handwriting_complete
Dataset of Full PAGE XML and ALTO Annotations in Handwritten Estonian Documents
Dataset Description
This dataset contains full page-level Transkribus exports from Estonian historical documents. Each example pairs a full page image with the corresponding PAGE XML and ALTO XML for the same page, preserving document structure, layout coordinates, reading order, baselines, and text content where available.
The dataset is intended for OCR research, layout analysis, document… See the full description on the dataset page: https://huggingface.co/datasets/Rahvusarhiiv/et_handwriting_complete.resultsreCAPTCHAv2Data used for training and validation of models in the paper Breaking reCAPTCHAv2.
The code can be accessed on GitHub.
