datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UrbanVerse-Training-Scenes
UrbanVerse Training Scenes (Urban Cousins)
A collection of ready-to-simulate urban 3D scenes in OpenUSD
for NVIDIA Isaac Sim / Isaac Lab, released by the
VAIL-UCLA lab. Each scene is a self-contained
USD stage with all of its materials and textures, so it can be opened and
simulated directly.
The scenes are generated with UrbanVerse — Scaling Urban Simulation by
Watching City-Tour Videos (Liu et al., ICLR 2026,
arXiv:2510.15018,
project page) — whose UrbanVerse-Gen
pipeline… See the full description on the dataset page: https://huggingface.co/datasets/UCLA-VAIL/UrbanVerse-Training-Scenes.collected_demos_trainingimage_training_set自用的训练集合集,用于 Stable Diffusion 模型微调。
该仓库仅用于存档,不提供任何技术支持。
Bee-Training-Data-Stage2
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality.
Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Bee-Training-Data-Stage2.sole_training_data
This is the training dataset for SOLE-R1-8B
SOLE-R1-8B is a video-language reward reasoning model for robotics. It is designed to estimate task progress from robot video frames and a natural-language task description, producing both per-timestep reasoning traces and scalar progress predictions that can be used as rewards for online robot reinforcement learning.
This dataset accompanies the paper “SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot RL” by Philip… See the full description on the dataset page: https://huggingface.co/datasets/Philip-MIT/sole_training_data.gaussian_training_datasets
Gaussian Training Datasets (COLMAP) for msplat
COLMAP-format multi-view scenes for training 3D Gaussian Splatting models,
packaged for msplat — a Metal-native 3DGS
trainer for Apple Silicon. Also includes pre-trained .ply splats under
tested_outputs/.
All scenes are redistributed from third-party datasets. Full credit goes to
their original authors — see Licensing & credits and please
cite the original papers. This repo only repackages them in COLMAP layout for
convenience.… See the full description on the dataset page: https://huggingface.co/datasets/alexmkwizu/gaussian_training_datasets.GEWDiff_training_dataset
GEWDiff Training & Evaluation Dataset
📘 Overview
The GEWDiff Training & Evaluation Dataset is derived from the EnMAP Champion and MDAS hyperspectral datasets.It is designed for image enhancement, super-resolution, restoration, and generative remote sensing tasks.The dataset includes Low-Quality (LQ) low-resolution images, corresponding Ground-Truth (GT) high-resolution images, and optional structure information such as masks and edges (partially provided;… See the full description on the dataset page: https://huggingface.co/datasets/zhu-xlab/GEWDiff_training_dataset.sat-vl-sft-training-ready-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.ai2thor-perspective-qa-100k-balanced-training-v1-splitsterra-4m-training-log
Terra 4M Training Log
Terra 4M is a 4.2 million parameter, purely convolutional diffusion model for terrain generation.
This is a log of all the checkpoints and images generated throughout training.
Note that the checkpoint corresponding to step 3,982,014 was selected as the final model.
The training code can be found here.
lora-training-datasetsdeewaiREALCN-training
Repo
git@hf.co:datasets/telcom/deewaiREALCN-training
DeewaiREALCN Training Data
Image–text pairs for training captioning or vision–language models. Each image is a 1024×1024 RGB JPEG portrait with a short English description.
Contents
data/train/: 9,000 pairs for training.
images/: JPEG files (090000.jpg, …).
captions.jsonl: one JSON object per line with file_name and text.
data/val/: 1,000 pairs for validation with the same layout.
Example… See the full description on the dataset page: https://huggingface.co/datasets/javadtaghia/deewaiREALCN-training.e2e-stream-slam-training-dataset
e2e-stream-slam training assets
Reproducibility bundle for the V4 SLAMFormer ablation suite.
Code: https://github.com/SlamMate/e2e-semantic-SLAM/tree/submap (commit 3195a7a)
Contents
Checkpoints
File
Size
Role
checkpoints/v1_paper_ckpt10.pth
3.6 GB
SLAMFormer paper base ckpt (10 ep on the paper datasets). PRETRAINED init for V3 Scale Token training.
checkpoints/v3_scale_token_ckpt2.pth
3.8 GB
V3 Scale Token epoch-2 (3 ep, 3×A6000… See the full description on the dataset page: https://huggingface.co/datasets/qizhangslam/e2e-stream-slam-training-dataset.Pocket-Rocket-1.0-Mid-Training-23MThis repository stores converted raw image WebDataset tar shards from multiple source datasets for streaming training.
latent-image-training
squiggles (metadata-fix)
OC-map FEM rebuild at 35 pixels per wavelength, with corrected geometries,
Helmholtz residuals, and the resolved JCMsuite .jcm / .jcmp files used
for each solve.
Configs
metadata (default)
One row per structure folder (sample_XXXX). Geometry comes from published
optical-constant maps (not the old nested-interface metadata).
validation
One row per FEM incidence (theta in {0, 45}). Self-contained pixel map:… See the full description on the dataset page: https://huggingface.co/datasets/als-rixs/latent-image-training.Bee-Training-Data-Stage2
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality.
Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/toilaluan/Bee-Training-Data-Stage2.OmniRef-trainingfast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde
Fast-Food Cleaning Robot — Floor Mess Dataset
Training dataset for a cleaning robot operating in fast-food-style food-service spaces (break areas / dining). Scenes are staged in break-area environments cluttered with food-service furnishings and food items (pizza, grocery food, cups, spoons) so the robot learns to perceive and act on mess. Covers detection, grasping, navigation, obstacle avoidance and pick-and-place. Renders are 1024x1024 with RGB plus albedo, metric depth and… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde.webui-training-dataVideoChat-Flash-Training-Data-subsetSOC-Training-Data-Visualization
Paper Link
SOS: Synthetic Object Segments Improve Detection, Segmentation, and Grounding
Code repo
Code for Generation
Citation
@misc{huang2025sossyntheticobjectsegments,
title={SOS: Synthetic Object Segments Improve Detection, Segmentation, and Grounding},
author={Weikai Huang and Jieyu Zhang and Taoyang Jia and Chenhao Zheng and Ziqi Gao and Jae Sung Park and Ranjay Krishna},
year={2025},
eprint={2510.09110},
archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/SOC-Training-Data-Visualization.DIV8K_TrainingSetThe training set of DIV8K.
Citation
@inproceedings{gu2019div8k,
title={Div8k: Diverse 8k resolution image dataset},
author={Gu, Shuhang and Lugmayr, Andreas and Danelljan, Martin and Fritsche, Manuel and Lamour, Julien and Timofte, Radu},
booktitle={ICCVW},
year={2019},
}
ISIC_2019_Training_InputScenicOrNot_640x480_training_v1minigpt4_training_for_MMPretrain
Dataset for training MiniGPT4 from scratch in MMPretrain
More information and guide can be found in docs of MMPretrain.
license: cc-by-nc-4.0
gsplat-training-frames_horizontobservatorium
Dataset Card for Gaussian Splatting Drone Frames — Horizontobservatorium (DE)
Direct Use
Training and evaluation of Gaussian Splatting (3DGS/gsplat) and NeRF variants.
3D reconstruction with SfM/MVS (e.g., COLMAP) and validation of photogrammetry pipelines.
Benchmarks/ablations (PSNR/SSIM/LPIPS), pose estimation, approximate intrinsic calibration, metric scaling with GPS.
Out-of-Scope Use
Person/vehicle recognition or surveillance: the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/naruone90/gsplat-training-frames_horizontobservatorium.mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 2B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k.driving_sample_tipsv2_training_images_10kffhq-256_training_facesmhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3.5 9B think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k.
