datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sat-image-boundingbox-sft-full
NU-TONIC raw SFT Full
Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning in SFT format. JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters.
Provenance
Locations: GeoGuessr-style POIs (source: stochastic/random_streetview_images_pano_v0.0.2)
Optical: Sentinel-2 multispectral optical COGs from a public STAC catalog, blue/green/red or visual preview, percentile-stretched to uint8.
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft-full.brain-mri-dataset-140
Brain Tumor MRI Dataset (Visual Viewer Enabled)
This dataset contains structural MRI cross-sections processed from clinical scans, formatted into interactive image columns for direct streaming.
Dataset Features Map
image: Viewable cross-section slice image.
volume / slice: Core scan extraction reference coordinates.
Age / Survival_Days: Patient clinical records metrics.
Grade: Tumor classifications status (e.g., HGG, LGG).
anatomical_location: Specific scan… See the full description on the dataset page: https://huggingface.co/datasets/Satavisha2026/brain-mri-dataset-140.sat-vl-sft-training-ready-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.SATraj-OS
SATraj-OS: Scaling Agent Trajectories for OSWorld
中文 | English
SATraj-OS is a large-scale multimodal graphical user interface (GUI) interaction trajectory dataset for computer-using agents (CUAs), designed for general capability learning and safety training.
📈 Model Performance
After joint training on Capability and Safety-v2, the SCOPE model series achieves a stronger balance between general capability on OSWorld and safety on OS-Blind. The x-axis below… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/SATraj-OS.SA-Text
SA-Text
Text-Aware Image Restoration with Diffusion Models (arXiv:2506.09993)Large-scale training dataset for the Text-Aware Image Restoration (TAIR) task.
📄 Paper: https://arxiv.org/abs/2506.09993
🌐 Project Page: https://cvlab-kaist.github.io/TAIR/
💻 GitHub: https://github.com/cvlab-kaist/TAIR
🛠 Dataset Pipeline: https://github.com/paulcho98/text_restoration_dataset
Dataset Description
SA-Text is constructed from SA-1B dataset using our official dataset… See the full description on the dataset page: https://huggingface.co/datasets/Min-Jaewon/SA-Text.SAT
SAT: Spatial Aptitude Training for Multimodal Language Models
Project Page
If you wish to use the latest versions of Huggingface and please check out the updated version of SAT here.
It is the same dataset but compatible with the latest versions.
To use this dataset, first make sure you have Python3.10 and Huggingface datasets version 3.0.2 (pip install datasets==3.0.2).
from datasets import load_dataset
import io
split = "val"
dataset = load_dataset("array/SAT"… See the full description on the dataset page: https://huggingface.co/datasets/array/SAT.painting_movements_2SAT-Spatial-VQAmarine-figuressatellite-building-segmentation
Dataset Labels
['building']
Number of Images
{'train': 6764, 'valid': 1934, 'test': 967}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds = load_dataset("keremberke/satellite-building-segmentation", name="full")
example = ds['train'][0]
Roboflow Dataset Page
https://universe.roboflow.com/roboflow-universe-projects/buildings-instance-segmentation/dataset/1
Citation… See the full description on the dataset page: https://huggingface.co/datasets/keremberke/satellite-building-segmentation.DOTAv2DOTA v2 Dataset with OBB, specifically the version from the Ultralytics docs
Website
Full License
Here reproduced from the website webpage
License for Academic Non-Commercial Use Only
This DOTA dataset is made available under the following terms:
The Google Earth images in this dataset are subject to Google Earth's terms of use, which must be adhered to.
The GF-2 and JL-1 satellite images are provided by the China Centre for Resources Satellite Data and Application. The… See the full description on the dataset page: https://huggingface.co/datasets/satyamshorrf/DOTAv2.sat-vl-sft-postprocessed-merged-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-postprocessed-merged-v1.braille_dataset_2sat-image-boundingbox-sft
NU-TONIC raw SFT init
Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning LFM-VL (leap-finetune vlm_sft format). JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters.
Provenance
Locations: GeoGuessr-style POIs (default HF source: stochastic/random_streetview_images_pano_v0.0.2) via download_geoguessr_poi_imagery.py.
Optical: multispectral optical COGs from a public STAC catalog… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft.SAT-v2
SAT-v2 Dataset
Paper
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
This dataset is part of the SAT (Spatial Aptitude Training) project, which introduces a dynamic benchmark for evaluating and improving spatial reasoning capabilities in multimodal language models.
Project Page: https://arijitray.com/SAT/
Paper: arXiv:2412.07755
Dataset Description
SAT-v2 is a comprehensive spatial reasoning benchmark containing over 300,000… See the full description on the dataset page: https://huggingface.co/datasets/array/SAT-v2.painting_movements3ts-satfire
Dataset Card for TS-SatFire
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
The TS-SatFire dataset is a comprehensive multi-temporal remote sensing dataset designed to cover the entire life cycle of wildfires. It provides a unified framework to support three critical and interconnected wildfire monitoring tasks: active fire detection, daily burned area… See the full description on the dataset page: https://huggingface.co/datasets/SamuelWu318/ts-satfire.sat-bbox-metadata-sft-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-bbox-metadata-sft-v1.sat-problems-datasetkitti-sat
kitti-sat
KITTI raw dataset, organized by capture date.
LFM-Orbit-SatData
LFM Orbit SatData
Retagged Earth-observation training data produced by LFM Orbit for the Liquid AI x DPhi Space Hackathon.
The default viewer config is training_assets.jsonl, which contains single-image SFT rows with image, messages, and metadata. Temporal sequence rows live in the temporal_sft config so the Hugging Face Dataset Viewer does not try to cast sequence rows into the single-image schema.
Configs
Config
File
Purpose
default
training_assets.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Shoozes/LFM-Orbit-SatData.safemaize-v2
SafeMaize v2
We use this preliminary public-source dataset for maize screening experiments.
The export contains 46,143 distinct images, including
40,385 core classification images. Files retain their original
bytes and recorded frame-selection rules.
Core class
Images
nlb_tlb_like
20,661
healthy
14,108
faw_feeding_injury
5,616
Core screening task split
Images
train
28,269
val
4,039
calibration
4,038
test
4,039
Core primary source… See the full description on the dataset page: https://huggingface.co/datasets/sathiiii/safemaize-v2.satvian-databaseSAT
Introduction
Disclaimer: This dataset is organized and adapted from array/SAT. The original data has been converted here into a more accessible and easy-to-use format.
SAT: Spatial Aptitude Training for Multimodal Language Models. (150 image QA pairs) Real-image dynamic test set.
Data Fields
Field Name
Type
Description
question_id
string
Unique question ID
question
string
Question text
question_type
string
Type of question
answers
list[string]
Answer… See the full description on the dataset page: https://huggingface.co/datasets/FlagEval/SAT.VIGOR_SAT3DGEN_add_skymask_DSM_satdepth
VIGOR SAT3DGEN Supplement
This repository contains the project-specific supplements for the paper Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image.
Project Page | Code
Dataset Summary
These files are intended to be used alongside the original VIGOR dataset. The supplement includes:
sat_depth/: Satellite depth maps.
pano_sky_mask/: Sky masks for panoramic images.
Seattle_DSM/: High-resolution Digital Surface Model (DSM) data for… See the full description on the dataset page: https://huggingface.co/datasets/qian43/VIGOR_SAT3DGEN_add_skymask_DSM_satdepth.VHR-10The VHR-10 dataset mirrored from https://github.com/chaozhong2010/VHR-10_dataset_coco
NWPU VHR-10 data set is a challenging ten-class geospatial object detection data set. This dataset contains a total of 800 VHR optical remote sensing images, where 715 color images were acquired from Google Earth with the spatial resolution ranging from 0.5 to 2 m, and 85 pansharpened color infrared images were acquired from Vaihingen data with a spatial resolution of 0.08 m. The data set is divided into two… See the full description on the dataset page: https://huggingface.co/datasets/satellite-image-deep-learning/VHR-10.satellite-imagery-projectscreening3-v1
SafeMaize screening3_v1
We use this preliminary dataset for maize screening experiments. It contains
23,990 selected images with normalized public-source labels.
Images retain their original file bytes. No new images or label decisions are
introduced by this export.
Class
Images
nlb_tlb_like
6,251
healthy
12,126
faw_feeding_injury
5,613
Main task split
Images
train
16,793
val
2,399
calibration
2,399
test
2,399
Primary source… See the full description on the dataset page: https://huggingface.co/datasets/sathiiii/screening3-v1.satellite-disruption-triage-aux-v1-3
satellite-disruption-triage-aux-v1-3
Civilian Conflict-Disruption Satellite VLM Dataset — Auxiliary / v1.3
This is an auxiliary dataset for training and evaluating Vision-Language Models (VLMs) to perform civilian conflict-disruption triage from paired satellite imagery. It is not a tactical intelligence dataset and not a canonical expert benchmark.
Scope & Purpose
The target task is detecting macro-scale civilian infrastructure disruption caused by war, armed conflict… See the full description on the dataset page: https://huggingface.co/datasets/ChrisRPL/satellite-disruption-triage-aux-v1-3.SAT_perspective
SAT_perspective Dataset
Paper
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
This dataset is part of the SAT (Spatial Aptitude Training) project, which introduces a dynamic benchmark for evaluating and improving spatial reasoning capabilities in multimodal language models.
Project Page: https://arijitray.com/SAT/
Paper: arXiv:2412.07755
Dataset Description
The SAT_perspective dataset contains 6,527 spatial reasoning questions that… See the full description on the dataset page: https://huggingface.co/datasets/array/SAT_perspective.
