datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
remote-sensing-sft-data
RSCoVLM: Co-Training Vision Language Models for Remote Sensing Multi-task Learning
Qingyun Li*
Shuran Ma*
Junwei Luo*
Yi Yu*
Yue Zhou
Fengxiang Wang
Xudong Lu
Xiaoxing Wang
Xin He
Yushi Chen
Xue Yang
If you find our work helpful, please consider giving us a ⭐!
ArXiv Paper: https://arxiv.org/abs/2511.21272
Published Paper: https://www.mdpi.com/2072-4292/18/2/222… See the full description on the dataset page: https://huggingface.co/datasets/Qingyun/remote-sensing-sft-data.OmniReasoner-SFT
OmniReasoner-SFT
OmniReasoner-SFT is a mixed-source, research-only supervised fine-tuning dataset
for audio-visual and long-video reasoning. It contains two-stage cold-start SFT
trajectories with interval selection, zoom-in evidence, and final answers.
Contents
data/train.jsonl: HF-ready training JSONL with repo-relative media paths.
media/: raw and derived media referenced by train.jsonl.
manifests/media_manifest.jsonl: media inventory with repo paths, source
family… See the full description on the dataset page: https://huggingface.co/datasets/Rocky131/OmniReasoner-SFT.sat-image-boundingbox-sft-full
NU-TONIC raw SFT Full
Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning in SFT format. JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters.
Provenance
Locations: GeoGuessr-style POIs (source: stochastic/random_streetview_images_pano_v0.0.2)
Optical: Sentinel-2 multispectral optical COGs from a public STAC catalog, blue/green/red or visual preview, percentile-stretched to uint8.
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft-full.Omnimodal-Agent-SFT-2K
OmniGAIA: Omni-Modal General AI Assistant Benchmark
📄 Paper
•
💻 Code & Demo
•
🤗 Dataset & Model
•
📈 Leaderboard
This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.Embodied-R1.5-SFT-Dataset
Embodied-R1.5-SFT-Dataset
🌐 Project Page |
📄 arXiv |
💻 Code |
🧰 EmbodiedEvalKit |
🤗 Models & Datasets
🗓️ Update — 2026-08-20 (20260820). All 34 Stage 1 SFT JSON annotation files have been uploaded to sft_datasets_json/. The complete JSON ↔ image/video data mapping is documented in the Dataset composition table below.
⚠️ Partial release. This repository currently contains only a subset of the full Stage 1 SFT… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1.5-SFT-Dataset.emova-sft-4m
EMOVA-SFT-4M
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-SFT-4M is a comprehensive dataset curated for omni-modal instruction tuning, including textual, visual, and audio interactions. This dataset is created by gathering open-sourced multi-modal instruction datasets and synthesizing high-quality omni-modal conversation data to enhance user experience. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-4m.ETCHR-SFT-400K
ETCHR SFT-400K
📖Paper
| 🏠Homepage
| 🤗ETCHR-FLUX.2-klein-9B Model
| 🤗ETCHR SFT-400K Dataset
| 🤗ETCHR GRPO-10K Dataset
| 🤗DL3DV-2K Benchmark
ETCHR SFT-400K is the SFT training data for transfering a passive instruction-following image editor (built on FLUX.2-klein-base-9B) into an autonomous, question-conditioned visual reasoning assistant. It contains 400,000 samples of five tasks (Fine-grained Perception, Chart Understanding, Maze Solving, Jigsaw Puzzle and… See the full description on the dataset page: https://huggingface.co/datasets/BeichenZhang/ETCHR-SFT-400K.VBVR-Pro-SFT-Image
VBVR-Pro-SFT-Image
The interleaved-image supervised-fine-tuning split of VBVR-Pro: 1.24M programmatically generated reasoning instances across 250 parameterized tasks, one tar.gz per task.
Where VBVR-Pro-SFT-Video asks a model to render the reasoning process as a… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-Pro-SFT-Image.ChartVerse-SFT-600KChartVerse-SFT-600K is a large-scale, high-quality chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the opendatalab/ChartVerse project. For more details about our method, datasets, and full model series, please visit our Project Page.
This dataset contains non-trivial samples filtered by failure rate (r > 0), ensuring that every sample provides meaningful learning signal. Samples that are too easy (r = 0, where the model always answers correctly) are… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/ChartVerse-SFT-600K.We-Math2.0-SFTAgenticOCR-SFT
AgenticOCR SFT Training Data
Supervised fine-tuning data for the AgenticOCR project.
The dataset contains 7,631 training records in sft_combined_0422.json. Image paths in each record are relative to the repository root and point into sft_images/.
Monet-SFT-125K
Introduction
This is the SFT dataset for paper "Monet: Reasoning in Latent Visual Space Beyond Images and Language"
Paper: http://arxiv.org/abs/2511.21395
Code: https://github.com/NOVAglow646/Monet
Citation
If you find this work useful, please use the following BibTeX. Thank you for your support!
@misc{wang2025monetreasoninglatentvisual,
title={Monet: Reasoning in Latent Visual Space Beyond Images and Language},
author={Qixun Wang and Yang Shi and Yifei Wang… See the full description on the dataset page: https://huggingface.co/datasets/NOVAglow646/Monet-SFT-125K.MMFineReason-SFT-586K-Qwen3-VL-235B-Thinking
MMFineReason-SFT-586K
The Hardest 33% — Less Data, More Reasoning
📖 Overview
MMFineReason-SFT-586K is a difficulty-filtered subset of MMFineReason-1.8M, containing the hardest 33% of samples where Qwen3-VL-4B-Thinking do not consistently succeed. (pass rate ≠ 1).
Specifically, this subset removes all easy samples (pass rate = 1) under Qwen3-VL-4B-Thinking, retaining only instances that require non-trivial multimodal reasoning.
🎯 Key Highlights
586K… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-586K-Qwen3-VL-235B-Thinking.Qwen-SFT-Inference-OutputsVC-Tooler-SFT
VC-Tooler-SFT
Supervised cold-start trajectories for VC-Tooler: Learning Compositional and Adaptive Visual Tool Use.
🔗 Links
📄 Paper: arXiv
🌐 Project Page: w1zheng.github.io/VC-Tooler
🤗 Hugging Face: VC-Tooler-SFT (this dataset) · VC-Tooler-RL
🧩 ModelScope: VC-Tooler-SFT (this dataset) · VC-Tooler-RL
This dataset is the Stage I (supervised fine-tuning) trajectory bank used to teach a
vision–language model to use visual tools as a compositional and adaptive… See the full description on the dataset page: https://huggingface.co/datasets/5551z/VC-Tooler-SFT.sat-vl-sft-training-ready-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.scidocbench-sft
scidocbench-sft
This repository contains a packaged export of the sft split of SciDocBench.
Original file: sft_0420_26k_max20_images.json
Records: 24674
Referenced images: 60246
All image paths in the dataset file have been rewritten from cluster-local absolute
paths to repository-relative paths under images/, so the dataset can be moved or
downloaded without depending on the original filesystem layout.
VBVR-Pro-SFT-Video
VBVR-Pro-SFT-Video
The video (I2V) supervised-fine-tuning split of VBVR-Pro: 1.24M programmatically generated reasoning instances across 250 parameterized tasks, one tar.gz per task.
At a glance
Property
Value
Tasks
250
Instances
1,250,000… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-Pro-SFT-Video.trailogy-na-plantae-sft
NA-Plantae SFT data tree
Mixed supervised-finetuning corpus for the on-device Gemma 4 E2B "hike
companion" VLM. Combines a North-American Plantae image-ID slice (sourced
from iNaturalist, label text enriched via GBIF) with general
anti-forgetting buckets (LLaVA-style image QA + refusal/negative).
Layout
inaturalist_na_plantae/ # web-crawled label sources (irreproducible)
observations.jsonl # iNaturalist observation metadata… See the full description on the dataset page: https://huggingface.co/datasets/TimS-ml/trailogy-na-plantae-sft.ChartVerse-SFT-1800KChartVerse-SFT-1800K is an extended large-scale chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the opendatalab/ChartVerse project. For more details about our method, datasets, and full model series, please visit our Project Page.
This dataset contains all verified correct samples without failure rate filtering. Unlike SFT-600K which excludes easy samples (r=0), SFT-1800K includes the complete set of truth-anchored QA pairs for maximum coverage and scale.… See the full description on the dataset page: https://huggingface.co/datasets/0xzanuee/ChartVerse-SFT-1800K.InSight-doc-SFT-18k
InSight-doc-SFT-18k
Agentic Visual Perception for Long-Document Understanding
📄 Paper |
💻 Code |
🤗 Model |
🎯 RL Data |
🎬 Replay Demo |
🚀 Live Demo
Understand the big picture. Focus on the right details. Answer from the evidence.
InSight-doc-SFT-18k is the supervised fine-tuning corpus used to train the
InSight-doc long-document understanding agent. Each example is a complete
multimodal trajectory: the agent starts from low-resolution document… See the full description on the dataset page: https://huggingface.co/datasets/m-Just/InSight-doc-SFT-18k.ChartVerse-SFT-1.8MChartVerse-SFT-1800K is an extended large-scale chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the opendatalab/ChartVerse project. For more details about our method, datasets, and full model series, please visit our Project Page.
This dataset contains all verified correct samples without failure rate filtering. Unlike SFT-600K which excludes easy samples (r=0), SFT-1800K includes the complete set of truth-anchored QA pairs for maximum coverage and scale.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/ChartVerse-SFT-1.8M.OpenMMReasoner-SFT-874K
OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
Introduction
OpenMMReasoner is a dataset introduced in the paper OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe. It supports the development of multimodal reasoning capabilities, utilizing a fully transparent two-stage recipe spanning supervised fine-tuning (SFT) and reinforcement learning (RL). The dataset includes an… See the full description on the dataset page: https://huggingface.co/datasets/OpenMMReasoner/OpenMMReasoner-SFT-874K.gdufs-molmo2-sftumm-sft-data
UMM-SFT Counting Data and Auxiliary Bundle
This dataset is a deterministic, public-source reconstruction of the Counting
data protocol described in Towards Physics of Multimodal Pretraining:
Knowledge Flow, Modality Synergy, Early Unification, and Recipes
(arXiv:2608.05000v2). It pairs the same
image population with image-understanding counting supervision and an
image-generation prompt consisting of an immutable count-fact prefix followed
by a count-free Qwen3-VL scene… See the full description on the dataset page: https://huggingface.co/datasets/xzz789/umm-sft-data.brief-composer-sft-v1
BriefComposer SFT
Multi-image analytical brief rows composed from completed FireWatch, OceanScout, LandShift, and FloodPulse dataset folders (metadata/ + images/). Each sample stitches 1–4 images and metadata-derived headlines into one executive-style assistant reply.
Record counts (this build)
Split
JSONL lines
train
6307
validation
851
test
842
total
8000
Inputs
Source roots: one or more --source-root directories (each must contain… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/brief-composer-sft-v1.gigaverbo-v2-sft
GigaVerbo-v2 SFT: A Large-Scale Portuguese Instruction-Tuning Dataset
Dataset Summary
GigaVerbo-v2 SFT is a large-scale instruction-tuning dataset designed for supervised fine-tuning of language models in Portuguese. The dataset comprises approximately 2.1 billion tokens (~4.4 GB) across 4 million instruction-following examples, organized into 12 distinct task categories. It is entirely composed of high-quality, LLM-generated data that has been carefully curated and… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/gigaverbo-v2-sft.TexOCR-SFT-figuresTexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction [ACL 2026 Main]
This repository provides the figure/image data used in TexOCR SFT training.
Overview
Dataset: TexOCR_SFT_figures
Type: Training Images
Task: Document OCR → LaTeX generation
142_aitw_sft_fbc
aitw_sft_fbc
This model is a fine-tuned version of /mnt/nvme0n1p1/hongxin_li/jingfan/LLaMA-Factory/models/qwen2_vl_lora_sft_aitw_all on the vl_sft_data_aitw dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
learning_rate:… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/142_aitw_sft_fbc.Search-VL-SFT-36K
An Open Recipe for Frontier Multimodal Search Agents
Cold-Start Agentic SFT · Multi-Turn Fatal-Aware GRPO · Visual Tool Use
📑 Table of Contents
📖 Introduction
🗺️ Overview
🍭 Method Overview
📊 Main Results
🔎 Case Study
📁 Repository Layout
🛠️ Prerequisites
🏋️ Agentic SFT · code/SFT
🚀 Agentic RL · code/RL
📊 Inference & Evaluation · code/infer
🚧 TODO
🙌 Acknowledgements
📖 Introduction
OpenSearch-VL is a fully… See the full description on the dataset page: https://huggingface.co/datasets/OpenSearch-VL/Search-VL-SFT-36K.
