datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VQAv2_train
Dataset Card for "VQAv2_train"
More Information needed
fatestaynightufotable
Bangumi Image Base of Fate Stay Night [ufotable]
This is the image base of bangumi Fate Stay Night [UFOTABLE], we detected 27 characters, 3899 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/fatestaynightufotable.old-comfyui-freezeFGVC_Aircraft_train
Dataset Card for "FGVC_Aircraft_train"
More Information needed
StanfordCars_test
Dataset Card for "StanfordCars_test"
More Information needed
StanfordCars_train
Dataset Card for "StanfordCars_train"
More Information needed
COCO_captions_train
Dataset Card for "COCO_captions_train"
More Information needed
FGVC_Aircraft_test
Dataset Card for "FGVC_Aircraft_test"
More Information needed
Iris_Database
Synthetic Iris Image Dataset
Overview
This repository contains a dataset of synthetic colored iris images generated using diffusion models based on our paper "Synthetic Iris Image Generation Using Diffusion Networks." The dataset comprises 17,695 high-quality synthetic iris images designed to be biometrically unique from the training data while maintaining realistic iris pigmentation distributions. In this repository we contain about 10000 filtered iris images with the… See the full description on the dataset page: https://huggingface.co/datasets/fatdove/Iris_Database.deepseek-ocr-arabic-v1noor-platform-fatwafatezero
Bangumi Image Base of Fate/zero
This is the image base of bangumi Fate/Zero, we detected 26 characters, 2067 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/fatezero.fatekaleidlinerprismaillya
Bangumi Image Base of Fate - Kaleid Liner Prisma Illya
This is the image base of bangumi Fate - kaleid Liner Prisma Illya, we detected 44 characters, 4621 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/fatekaleidlinerprismaillya.TB3-trojectoryVizWiz_train
Dataset Card for "VizWiz_train"
More Information needed
COCO_captions_validation
Dataset Card for "COCO_captions_validation"
More Information needed
OK-VQA_train
Dataset Card for "OK-VQA_train"
More Information needed
SH-17-Datasetfathomnet-test
Dataset Card for "fathomnet-test"
More Information needed
ytmVQAv2_validation
Dataset Card for "VQAv2_validation"
More Information needed
VQAv2_test
Dataset Card for "VQAv2_test"
More Information needed
zea-fat-layer-2025This dataset contains some acquisitons of the CIRS040 phantom, imaged through a layer of fat. The dataset is intended for speed of sound estimation. All data was acquired with a Verasonics L11-5v transducer.
The files ending in _no_mask are regular acquisitions.
The files ending in _mask have a layer of fat between transducer and phantom.
The files ending in _mask_thick have a slightly thicker layer of fat between transducer and phantom.
The files ending in _mask_half have a layer of fat… See the full description on the dataset page: https://huggingface.co/datasets/zeahub/zea-fat-layer-2025.CIFAR10_train
Dataset Card for "CIFAR10_train"
More Information needed
SNLI-VE_train
Dataset Card for "SNLI-VE_train"
More Information needed
fattah-golden-superset
Fattah Golden
Fattah Golden is a large-scale, model-agnostic supervised fine-tuning (SFT) superset built by Nomeda Labs to train the Fattah family of coding and agentic coding models.
The dataset is designed as a labeled superset with no baked-in training ratios. This means the stored dataset is the complete cleaned and annotated corpus. Researchers and practitioners choose their own mixture at training time by filtering on the boolean capability columns.
Stats… See the full description on the dataset page: https://huggingface.co/datasets/nomeda-lab/fattah-golden-superset.FATURA2-invoicesThe dataset consists of 10000 jpg images with white backgrounds, 10000 jpg images with colored backgrounds (the same colors used in the paper) as well as 3x10000 json annotation files. The images are generated from 50 different templates.
https://zenodo.org/records/10371464
dataset_info:
features:
- name: image
dtype: image
- name: ner_tags
sequence: int64
- name: words
sequence: string
- name: bboxes
sequence:
sequence: int64
splits:
- name: train… See the full description on the dataset page: https://huggingface.co/datasets/mathieu1256/FATURA2-invoices.CIFAR10_test
Dataset Card for "CIFAR10_test"
More Information needed
aerobench
AeroBench: Aviation Document Extraction Benchmark
The first open benchmark for evaluating AI systems that extract structured data from aviation release certificates.
Overview
AeroBench provides real-world EASA Form 1 (Authorised Release Certificate) and FAA Form 8130-3 (Airworthiness Approval Tag) documents with verified ground truth annotations for benchmarking document extraction systems.
These forms are the critical documents in aviation maintenance — every time a part… See the full description on the dataset page: https://huggingface.co/datasets/FathinDos/aerobench.CUB_train
Dataset Card for "CUB_train"
More Information needed
