datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sat-vl-sft-training-ready-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.OmniRef-trainingwebui-training-datah0_post_train_db_2508
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
We introduce Being-H0, the first dexterous Vision-Language-Action model pretrained from large-scale human videos via explicit hand motion modeling.
News
[2025-08-02]: We release the Being-H0 codebase and pretrained models! Check our Hugging Face Model Hub for more details. 🔥🔥🔥
[2025-07-21]: We publish Being-H0! Check our paper here. 🌟🌟🌟
Model Checkpoints… See the full description on the dataset page: https://huggingface.co/datasets/BeingBeyond/h0_post_train_db_2508.DeepEyes_train_4Kvila-q-data-traingraph-captioning-train-onlyMind2Web_train_llava
Mind2Web training set for the paper: Harnessing Webpage Uis For Text Rich Visual Understanding
🌐 Homepage | 🐍 GitHub | 📖 arXiv
Introduction
We introduce MultiUI, a dataset containing 7.3 million samples from 1 million websites, covering diverse multi- modal tasks and UI layouts. Models trained on MultiUI not only excel in web UI tasks—achieving up to a 48% improvement on VisualWebBench and a 19.1% boost in action accuracy on a web agent dataset Mind2Web—but also… See the full description on the dataset page: https://huggingface.co/datasets/neulab/Mind2Web_train_llava.zalo-ai-2025-training-data-v2
Zalo AI Challenge 2025 - RoadBuddy Training Data V2set
This dataset contains training data v2 for the RoadBuddy – Understanding the Road through Dashcam AI challenge from Zalo AI Challenge 2025.
Dataset Description
The challenge aims to build a driving assistant capable of understanding video content from dashcams to quickly answer questions about traffic signs, signals, and driving instructions in Vietnam.
Dataset Structure
Files
frames/: Directory… See the full description on the dataset page: https://huggingface.co/datasets/OpenHay/zalo-ai-2025-training-data-v2.anny-render-corpus-train
anny-render-corpus
A camera-controlled render corpus from the ANNY rig, and the measurements that motivated it.
Code: weftspun/anny-render-corpus, on the 6-datasource side of the hexagon.
Everything here is produced by scripts in that repository and can be regenerated from it.
What this is for
Asked in plain language for eight camera azimuths, OmniGen2 returns a body that does not
turn. Recovered azimuth tracks the request with a slope of 0.04, where 1.00 is… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/anny-render-corpus-train.zalo-ai-2025-new-training-data-v2
Zalo AI Challenge 2025 - RoadBuddy New Train Data V2set
This dataset contains new train data v2 for the RoadBuddy – Understanding the Road through Dashcam AI challenge from Zalo AI Challenge 2025.
Dataset Description
The challenge aims to build a driving assistant capable of understanding video content from dashcams to quickly answer questions about traffic signs, signals, and driving instructions in Vietnam.
Dataset Structure
Files
frames/:… See the full description on the dataset page: https://huggingface.co/datasets/OpenHay/zalo-ai-2025-new-training-data-v2.ridgelora-stage2-imposebase-train160-50k-20260824
Stage-2 ControlNet retraining with the frozen IMPOSE base
This experiment retrains only Stage 2 for RidgeLoRA-FP. Stage 1 is the IMPOSE
checkpoint and is not retrained. The run started on 2026-08-24 on TPU VM
t1v-n-d3df3356-w-0 (TPU v5p-8, four XLA devices).
An initial Stage-1-from-scratch job was stopped at step 575 after correcting
the scope. It produced no scheduled checkpoint and is not used in any result;
its log is retained only as an audit trail.
Frozen IMPOSE… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-stage2-imposebase-train160-50k-20260824.IntroSVG-trainDeepEyes_train_4KMIS_Train
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
Our paper, code, data, models can be found at MIS.
Dataset Structure
Our MIS train set contains 3927 samples with safety CoT labels generated by InternVL2.5-78B. The template is consistent with InternVL2.5 fine-tuning template.
{
"conversations": "list",
"image": "list",
"id": "int",
"category": "str",
"sub_category": "str"
}
train_img_vidcodette_training
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/Raiff1982/codette_training.SQuID-Train
SQuID-Train — 75,687 Quantitative Spatial Reasoning Training Pairs
Training companion to the SQuID benchmark
(https://huggingface.co/datasets/squid-bench-anon/SQuID), under review at the
NeurIPS 2026 Evaluations and Datasets Track.
75,687 question-answer pairs over 1,443 satellite images
Generated by the same pipeline as the benchmark (identical question templates,
geometric definitions, minimum-area thresholds, and GSD handling), applied to the
training partitions of the… See the full description on the dataset page: https://huggingface.co/datasets/squid-bench-anon/SQuID-Train.CrossMath-Traingensin-instruct-read-to-trainMedical-trainlicense: apache-2.0
task_categories:
question-answering
text-generation
language:
en
tags:
medical
biology
configs:
config_name: en
data_files:PNG_DIR
data_files: tset2.json
tags:
audio
logos_train_dataGUI-Critic-Traincontrolnet_trainRiR_Agent_Train_20ktraining_m3diff
M³Diff Instruction-Tuning Data
This repository contains the instruction-tuning annotations used to train
M³Diff, the model introduced in OmniDiff: A Comprehensive Benchmark for
Fine-grained Image Difference Captioning
(ICCV 2025).
The dataset contains 896,015 question-answer records for fine-grained
Image Difference Captioning (IDC). Each record presents two related images and
asks the model to describe the visual changes between them. This repository
contains the JSON… See the full description on the dataset page: https://huggingface.co/datasets/IVC-liuyuan/training_m3diff.train_controlnetfrom datasets import load_dataset, Features, Value, Image
dataset = load_dataset('json', data_files='YMING222/train_controlnet/train.jsonl')
features = Features({
'condition_images': Image(),
'images': Image(),
'text': Value('string')
})
dataset = dataset.cast(features)
ser-re-training-data
SER & RE Training Data
Datasets for Semantic Entity Recognition (SER) and Relation Extraction (RE) training.
Datasets
1. XFUND (~/data/ser_re/xfund/)
Source: https://github.com/doc-analysis/XFUND/releases/tag/v1.0
Languages: zh, ja, es, fr, it, de, pt (7 languages)
Documents: 149 train + 50 val per language = 1,393 total
Images: 1,393 JPG files ({lang}{split}{idx}.jpg)
Annotations: {lang}.train.json / {lang}.val.json
Format: Word-level bboxes, BIO labels… See the full description on the dataset page: https://huggingface.co/datasets/bluecopa/ser-re-training-data.zalo-ai-2025-training-synthetic-filtered
Zalo AI Challenge 2025 - RoadBuddy Training Data Synthetic Filteredset
This dataset contains training data synthetic filtered for the RoadBuddy – Understanding the Road through Dashcam AI challenge from Zalo AI Challenge 2025.
Dataset Description
The challenge aims to build a driving assistant capable of understanding video content from dashcams to quickly answer questions about traffic signs, signals, and driving instructions in Vietnam.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/OpenHay/zalo-ai-2025-training-synthetic-filtered.JARVIS-train
