datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WebLINX-full
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
WARNING: This is not the main WebLINX data card! You might want to use the main WebLINX data card instead:
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
Xing Han Lù*, Zdeněk Kasner*, Siva Reddy
💾Code
📄Paper
🌐Website
📓Colab
🤖Models
💻Explorer
🐦Tweets
🏆Leaderboard
Your browser does not support the… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/WebLINX-full.pd12m-fullThis dataset is the downloaded variant of Spawning/PD12M. More specifically, this dataset
is compatible with webdataset. It was made public after obtaining permission
from the original authors of the dataset.
You can use the following to explore the dataset with webdataset:
import webdataset as wds
dataset_path = "pipe:curl -s -f -L https://huggingface.co/datasets/sayakpaul/pd12m-full/resolve/main/{00155..02480}.tar"
dataset = (
wds.WebDataset(dataset_path… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/pd12m-full.fullmetaaitw-processed-labeled-full
AiTW Processed Full with App Labels
This repository contains a full processed Android in the Wild (AiTW) mirror together with an app-labeled step index, official split assignment by episode_id, major-app statistics, and a ready-to-train Gmail subset.
Why This Exists
AiTW is large and not easy to navigate by app. The original labels contain useful fields such as goal_info, current_activity, and action coordinates, but users often need extra processing before they… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/aitw-processed-labeled-full.open-vision-banana-snvc-train-full
SNVC-50M v5_full — Multi-Task Vision Dataset
Description
This dataset is a curated subset of the SenseNova Vision Corpus 50M (SNVC-50M), containing 43,509 samples across 6 vision task families and 31 source datasets. Each sample follows a conversational format with interleaved <image> tokens, designed for training vision-language models (VLMs).
Coverage: 43,509 / 57,878 (75.2%) of the original sampling plan. 23 datasets at 100%, 8 partial, 12 unrecoverable… See the full description on the dataset page: https://huggingface.co/datasets/gatilin/open-vision-banana-snvc-train-full.synthetic_data_v5_finegrain_layout_relight_with_our_synthetic_data_coco_l_full_500kbanana-vidorev3-fullpipe
Banana ViDoRe v3 Fullpipe
Nano Banana Pro full-pipeline synthetic training data for ViDoRe v3 finance and industrial domains.
This repository contains 4670 training records and 40190 unique referenced images across
domain-separated ColFlor/ColQwen training splits. Images are included in the repository and paths in each JSONL are
relative to that domain directory.
Generated at: 2026-06-29T09:39:01.644860+00:00
Layout
finance/train.jsonl
finance/metadata.json… See the full description on the dataset page: https://huggingface.co/datasets/vkehfdl1/banana-vidorev3-fullpipe.sat-image-boundingbox-sft-full
NU-TONIC raw SFT Full
Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning in SFT format. JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters.
Provenance
Locations: GeoGuessr-style POIs (source: stochastic/random_streetview_images_pano_v0.0.2)
Optical: Sentinel-2 multispectral optical COGs from a public STAC catalog, blue/green/red or visual preview, percentile-stretched to uint8.
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft-full.vsi-bench-qa-v3-hm3d-fullMMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.pitt-ads-fullDatBench-Full
DatBench: Discriminative, Faithful, and Efficient VLM Evaluations
DatBench is a curated evaluation suite for vision–language models (VLMs) designed to be faithful, discriminative, and efficient.
📄 DatBench: Discriminative, Faithful, and Efficient VLM Evaluationshttps://arxiv.org/abs/2601.02316
Modern VLM benchmarks often overestimate model capability due to multiple-choice inflation, language-only shortcuts, annotation noise, and redundant low-signal samples. DatBench reframes… See the full description on the dataset page: https://huggingface.co/datasets/DatologyAI/DatBench-Full.FullBenchj2me-full-romset
🎮 J2ME Full Romset
A complete preservation archive of rare J2ME (Java Mobile) games — saved from digital extinction.
This collection was manually gathered, organized, and uploaded as part of the M5 Nostalgia Archive preservation project. Many of these titles are no longer available anywhere online and have been delisted by their original publishers.
📦 Contents
Folder
Files
Size
Roms/
456 .jar games
~303 MB
Photos/
1,381 screenshots
~153 MB… See the full description on the dataset page: https://huggingface.co/datasets/M5-Dev/j2me-full-romset.full_checkbox_dropdown_radiobuttonBirds_of_North_America_FullGradients_Gradients_and_Text_Full_Logic_Captionsbl_books_flickr_fullMMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ericktwo/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.inspector_md_full_v1
maxs-m87/inspector_md_full_v1
Multi-config Hugging Face export of the Inspector MD training datasets.
Configs
detect: Detect with splits {'train': 11394, 'validation': 2240, 'test': 3225}
point: Point with splits {'train': 11394, 'validation': 2240, 'test': 3225}
query_issues: Query issues with splits {'train': 13795, 'validation': 2382, 'test': 3375}
Notes
Images are embedded into the published dataset artifacts; local training paths are not preserved.… See the full description on the dataset page: https://huggingface.co/datasets/maxs-m87/inspector_md_full_v1.fullhdmnist-cleaned-full
Dataset Card for 2025.11.21.16.40.44.970939
This is a FiftyOne dataset with 69807 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Linus-L/mnist-cleaned-full")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Linus-L/mnist-cleaned-full.vision-opd-vqa14k-fullimage-curriculum-v8
Vision-OPD VQA14K Full-Image Curriculum v8
Private single-image visual-question-answering dataset.
Split
Rows
Train
14,000
Diagnostic validation
609
The repository contains 14,609 content-addressed media files (4,580,273,467 bytes). Paths in both Parquet files are relative to the repository root and follow media/<sha256-prefix>/<filename>.
from pathlib import Path
import pyarrow.parquet as pq
from huggingface_hub import snapshot_download
root =… See the full description on the dataset page: https://huggingface.co/datasets/yyy051007/vision-opd-vqa14k-fullimage-curriculum-v8.libero_fullThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 807,
"total_frames": 153511,
"total_tasks": 20,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:807"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/siamaky1984/libero_full.processed_full_rgbuw
Dataset Card for "processed_full_rgbuw"
More Information needed
vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10
Trajectory Ranking Dataset
This dataset contains trajectory ranking results for autonomous navigation scenarios.
Dataset Statistics
Total examples: 39558
Chunks processed: 40
Upload date: 2025-09-13T00:44:30.335177
Features
Image data with terrain analysis
Trajectory rankings and reasoning
Quality and diversity analysis
Terrain and trajectory descriptions
2026-04-01T22-36-32plus00-00_gdpval_full220plantvillage-full
PlantVillage (full)
A curated re-release of the PlantVillage plant-disease image dataset
(Mohanty, Hughes, Salathé 2016), with structured per-image metadata and
a leaf-grouped train/test split. Built for the iResearch Institute 2026
Virtual Lab mentorship's Rationale 2 track. The companion debug-grade
subset is at
geraldmc/plantvillage-tiny.
What's in this dataset
54,304 images of plant leaves, photographed against plain backgrounds
under controlled lighting, across… See the full description on the dataset page: https://huggingface.co/datasets/geraldmc/plantvillage-full.pexels-tagger-v0-w640-ws-full
Pexels Tagger V0 Webdataset Full Dataset
This is the webdataset dataset for animetimm/pexels-wdtagger-w640.
Images here are resized to min(width, height) <= 640.
How to Use It
from datasets import load_dataset
dataset = load_dataset('animetimm/pexels-tagger-v0-w640-ws-full')
print(dataset["train"][0])
Images
3122908 images in total.
Split
Image Count
Total Size
train
2810634
184 GB
test
156409
10.2 GB
val
155865
10.2 GB
Tags… See the full description on the dataset page: https://huggingface.co/datasets/animetimm/pexels-tagger-v0-w640-ws-full.orpo-vlm-pairs-full
ORPO VLM Preference Pairs (Full)
This dataset contains two versions of vision-language preference pairs for training VLM models using ORPO, DPO, or similar preference-based alignment methods.
Dataset Description
File
Rows
Description
orpo_pairs.jsonl
67,754
Refined/filtered pairs (recommended)
orpo_pairs_all.jsonl
94,346
Full dataset before filtering
Images: 11,982 images
Format: JSONL + PNG images
Language: English
Task: Vision-language… See the full description on the dataset page: https://huggingface.co/datasets/mncai/orpo-vlm-pairs-full.
