datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiBanana-Benchmark🍌 MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation 🍌
CVPR 2026 (Main)
This repository provides the datasets for
“MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation” by Yuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo and Hiroki Furuta
Paper Link
https://arxiv.org/abs/2511.22989
Github Repository
For the usage of this benchmark, please see Github… See the full description on the dataset page: https://huggingface.co/datasets/kohsei/MultiBanana-Benchmark.kb-books
open-rdl-books
Dataset Description
Language
dan, dansk, Danish
License
Public Domain, cc0-1.0
Dataset Summary
Documents from the Royal Danish Library published between 1750 and 1930.
The dataset has each page of each document in image and text format. The text was extracted with OCR.
The documents (books of various genres) were obtained from the library. The dataset was assembled to make these public domain Danish texts more accessible.… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/kb-books.tapip3d-kubric
Kubric-MOVi-F 3D Point Tracking Dataset
Kubric MOVi-F synthetic videos with ground-truth 3D point trajectories for training TAPIP3D.
Links:
TAPIP3D: https://tapip3d.github.io/
Code: https://github.com/zbw001/TAPIP3D
KAIST-Multispectral-Pedestrian-Detection-Dataseteden_objaversekaz-vision-50kRFUAV The RFUAV DATASET
This repository contains the RFUAV dataset, presented in the paper "RFUAV: A Benchmark Dataset for Unmanned Aerial Vehicle Detection and Identification". RFUAV provides approximately 1.3 TB of raw frequency data collected from 37 distinct UAVs, offering a comprehensive benchmark for radio-frequency-based drone detection and identification. The dataset addresses limitations of existing datasets by providing a diverse range of drone types, sufficient data volume, coverage… See the full description on the dataset page: https://huggingface.co/datasets/kitofrank/RFUAV.imgKvasir-VQA-x1
Kvasir-VQA-x1
A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy
Kvasir-VQA-x1 on GitHub |
Original Image from Kvasir-VQA(Simula Datasets) |
Paper
🔗 MediaEval Medico 2025 Challenge uses this dataset. We encourage you to check out and participate!
Overview
Kvasir-VQA-x1 is a large-scale dataset designed to benchmark medical visual question answering (MedVQA) in gastrointestinal (GI) endoscopy. It introduces 159,549 new QA… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/Kvasir-VQA-x1.bf-assetsHAVEteleoperation-pick-and-placeaitw-processed-labeled-full
AiTW Processed Full with App Labels
This repository contains a full processed Android in the Wild (AiTW) mirror together with an app-labeled step index, official split assignment by episode_id, major-app statistics, and a ready-to-train Gmail subset.
Why This Exists
AiTW is large and not easy to navigate by app. The original labels contain useful fields such as goal_info, current_activity, and action coordinates, but users often need extra processing before they… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/aitw-processed-labeled-full.Motion-o-MCoT-PLM-motion-keyframes
Motion-o-MCoT (PLM + motion keyframes)
Subset of STGR: STR_plm_rdcap rows with <motion in reasoning_process, plus sharded keyframes under videos/stgr/plm/kfs/.
Train split: 3,168 examples (see export_manifest.json in the repo for exact export stats).
Keyframes: JPEGs are stored under shard subfolders (e.g. videos/stgr/plm/kfs/plm_0150/…) so each directory stays under Hugging Face file-count limits. Each key_frames[].path in the JSON is relative to videos/stgr/plm/kfs/ (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/bishoygaloaa/Motion-o-MCoT-PLM-motion-keyframes.PanScale
PanScale
Dataset Summary
PanScale is a remote-sensing pansharpening dataset with paired multispectral (ms) and panchromatic (pan) TIFF images for cross-scale evaluation.
Total pairs: 7,559
Disk size: ~6.9 GB
Format: 8-bit TIFF
Supported Tasks
Image Fusion (Cross-scale Pansharpening)
Subsets at a Glance
Subset
Splits
# Pairs
MS size
PAN size
PAN/MS scale
jilin
train200, test200/400/800
1,157
200-800
200-800
1.0
landsat
train256… See the full description on the dataset page: https://huggingface.co/datasets/kecao/PanScale.Bcblip3-kale
🥬 BLIP3-KALE:Knowledge Augmented Large-scale Dense Captions
BLIP3-KALE is an open-source dataset of 218 million image-text pairs, featuring knowledge-augmented dense captions combining web-scale knowledge with detailed image descriptions.
Paper: [To be added]
Uses
BLIP3-KALE is designed to facilitate research in multimodal pretraining. The dataset can be used for training large multimodal models that require factually grounded, dense image captions. It has already been an… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-kale.sovaNemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 국가데이터처 국가통계포털(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Korea.SceneText150kendoslamcoco-karpathy
Dataset Card for "yerevann/coco-karpathy"
The Karpathy split of COCO for image captioning.
Brain-Tumour-MRI
Dataset Card for Brain Tumour MRI dataset
A collection of Brain scans covering three different types of tumours and as well as a control class.
Dataset Details
Dataset Description
The Dataset contains ~7000 MRI scans of the brain corresponding to 4 classes: glioma, meningioma, notumor & pituitary.
The dataset has already been split into train/test sets.
Dataset Creation
Source
This dataset was compiled and uploaded to Kaggle by Masoud… See the full description on the dataset page: https://huggingface.co/datasets/Kaynaaf/Brain-Tumour-MRI.retail-products-philippinesmonopoly-assetsvisual-jenga-datasets
Visual Jenga Datasets
This directory contains the original datasets for Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting. Visual Jenga is a novel scene understanding task that involves progressively removing objects from a single image one at a time while keeping the rest of the scene stable. This process reveals object dependencies and provides a new way to evaluate grounded scene understanding by systematically exploring which objects can be removed… See the full description on the dataset page: https://huggingface.co/datasets/konpat/visual-jenga-datasets.kanchigainoateliermeistereiyuupartynomotozatsuyougakarigajitsuwasentouigaigasssrankdattatoiuyok
Bangumi Image Base of Kanchigai No Atelier Meister: Eiyuu Party No Moto Zatsuyougakari Ga, Jitsu Wa Sentou Igai Ga Sss Rank Datta To Iu Yoku Aru Hanashi
This is the image base of bangumi Kanchigai no Atelier Meister: Eiyuu Party no Moto Zatsuyougakari ga, Jitsu wa Sentou Igai ga SSS Rank Datta to Iu Yoku Aru Hanashi, we detected 55 characters, 5388 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kanchigainoateliermeistereiyuupartynomotozatsuyougakarigajitsuwasentouigaigasssrankdattatoiuyok.Kohaku-Delta-Alpha-Characters-Result
Test Index for Kohaku-Delta-Alpha Model
ID
Tag
Copyright
Gender
Posts
CCIP
AIC
BP
Core Tags
2075918
inkling_player_character
splatoon_(series)
female
9646
0.383333
0.998157
0.645165
inkling girl, long hair, tentacle hair, pointy ears, bangs, blunt bangs, red eyes
2040387
doodle_sensei_(blue_archive)
blue_archive
female
6632
0.721212
0.992072
0.625653
halo, bangs, blue eyes, breasts, long hair, black hair, blue hair, hair ornament
1978860
2b_(nier:automata)… See the full description on the dataset page: https://huggingface.co/datasets/AngelBottomless/Kohaku-Delta-Alpha-Characters-Result.aidovecl-vehicle-detection-classification-localization
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
We introduce an annotated AI-generated dataset of eye-level vehicle images using outpainting, offering versatile generation of diverse vehicle classes in varied contexts with pretrained models.
Citation Notice
Please ensure that all publications and presentations using this data reference the following paper:
Kazemi, A., Fatima, Q. ul A., Kindratenko, V., & Tessum, C. W.… See the full description on the dataset page: https://huggingface.co/datasets/amir-kazemi/aidovecl-vehicle-detection-classification-localization.visual-puzzles
