datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sole_training_data
This is the training dataset for SOLE-R1-8B
SOLE-R1-8B is a video-language reward reasoning model for robotics. It is designed to estimate task progress from robot video frames and a natural-language task description, producing both per-timestep reasoning traces and scalar progress predictions that can be used as rewards for online robot reinforcement learning.
This dataset accompanies the paper “SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot RL” by Philip… See the full description on the dataset page: https://huggingface.co/datasets/Philip-MIT/sole_training_data.gaussian_training_datasets
Gaussian Training Datasets (COLMAP) for msplat
COLMAP-format multi-view scenes for training 3D Gaussian Splatting models,
packaged for msplat — a Metal-native 3DGS
trainer for Apple Silicon. Also includes pre-trained .ply splats under
tested_outputs/.
All scenes are redistributed from third-party datasets. Full credit goes to
their original authors — see Licensing & credits and please
cite the original papers. This repo only repackages them in COLMAP layout for
convenience.… See the full description on the dataset page: https://huggingface.co/datasets/alexmkwizu/gaussian_training_datasets.Bee-Training-Data-Stage2
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality.
Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Bee-Training-Data-Stage2.GEWDiff_training_dataset
GEWDiff Training & Evaluation Dataset
📘 Overview
The GEWDiff Training & Evaluation Dataset is derived from the EnMAP Champion and MDAS hyperspectral datasets.It is designed for image enhancement, super-resolution, restoration, and generative remote sensing tasks.The dataset includes Low-Quality (LQ) low-resolution images, corresponding Ground-Truth (GT) high-resolution images, and optional structure information such as masks and edges (partially provided;… See the full description on the dataset page: https://huggingface.co/datasets/zhu-xlab/GEWDiff_training_dataset.lora-training-datasetse2e-stream-slam-training-dataset
e2e-stream-slam training assets
Reproducibility bundle for the V4 SLAMFormer ablation suite.
Code: https://github.com/SlamMate/e2e-semantic-SLAM/tree/submap (commit 3195a7a)
Contents
Checkpoints
File
Size
Role
checkpoints/v1_paper_ckpt10.pth
3.6 GB
SLAMFormer paper base ckpt (10 ep on the paper datasets). PRETRAINED init for V3 Scale Token training.
checkpoints/v3_scale_token_ckpt2.pth
3.8 GB
V3 Scale Token epoch-2 (3 ep, 3×A6000… See the full description on the dataset page: https://huggingface.co/datasets/qizhangslam/e2e-stream-slam-training-dataset.Bee-Training-Data-Stage2
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality.
Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/toilaluan/Bee-Training-Data-Stage2.VideoChat-Flash-Training-Data-subsetwebui-training-dataSOC-Training-Data-Visualization
Paper Link
SOS: Synthetic Object Segments Improve Detection, Segmentation, and Grounding
Code repo
Code for Generation
Citation
@misc{huang2025sossyntheticobjectsegments,
title={SOS: Synthetic Object Segments Improve Detection, Segmentation, and Grounding},
author={Weikai Huang and Jieyu Zhang and Taoyang Jia and Chenhao Zheng and Ziqi Gao and Jae Sung Park and Ranjay Krishna},
year={2025},
eprint={2510.09110},
archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/SOC-Training-Data-Visualization.GEWDiff_training_dataset
GEWDiff Training & Evaluation Dataset
📘 Overview
The GEWDiff Training & Evaluation Dataset is derived from the EnMAP Champion and MDAS hyperspectral datasets.It is designed for image enhancement, super-resolution, restoration, and generative remote sensing tasks.The dataset includes Low-Quality (LQ) low-resolution images, corresponding Ground-Truth (GT) high-resolution images, and optional structure information such as masks and edges (partially provided;… See the full description on the dataset page: https://huggingface.co/datasets/zhang5566/GEWDiff_training_dataset.DeepVRM-Training-DataAndrew_Alpha_training_data
Bark Beetle Grouped Images for AI Classification
Dataset Summary
This dataset comprises high-resolution photographs of bark and ambrosia beetles captured under controlled laboratory conditions. Each image contains multiple beetle specimens arranged on a uniform white background while submerged in 70% ethanol. This approach speeds up data collection and ensures reproducible imaging conditions. Individual beetle images can later be extracted from these grouped photographs… See the full description on the dataset page: https://huggingface.co/datasets/ChristopherMarais/Andrew_Alpha_training_data.Meta-CoT-Training-Data-Exampletraining
Training Data for Detecting Drones
This dataset was generated for a bachelor degree in computer science.
The aim was to improve the detection score on the drone versus bird dataset.
The images in each folder are from videos with a frame rate of 25 FPS.
They are intended to be used for the training of recurrent computer vision models.
laser_gui_grounding_training_dataChest_CT_GMPO_Training_Datasetscta-htr-training-data
Guiding Principles of this Dataset
Segmentation
The goal is here to segment the "main text" while ignoring paratext. This means page numbers, catch words, marginal annotations, are purposefully not segmented
Transcription
The guiding transcription principles are to
expand abbreviations
preserve orthography
MINDI-1.5-training-data
MINDI 1.5 Training Data
Training dataset for MINDI 1.5 Vision-Coder by MINDIGENOUS.AI
Dataset Statistics
Metric
Value
Total examples
1,449,428
Total tokens
859,694,776
Avg tokens/example
593
Avg quality score
6.49
Sources
9
Splits
Split
Examples
Percentage
Train
1,304,486
90.0%
Validation
72,471
5.0%
Test
72,471
5.0%
Sources
Source
Examples
Kept %
starcoderdata
569,350
94.9%
websight
250,987… See the full description on the dataset page: https://huggingface.co/datasets/Mindigenous/MINDI-1.5-training-data.Med_training_datatraining_datasettraining_densepose_data_finalz-image-training-datasetilids-faint-patterns-training-datasetRePIC-training-dataThis dataset accompanies the paper RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models.
dreambench_eval_results_internvl2_5_78b_mpo_awq_init_1_prompt_collect_training_dataHammershoi-Training-Dataset
Dataset Card for Hammershoi Training Dataset
A curated dataset of 244 image-caption pairs based on the works of Danish painter Vilhelm Hammershøi (1864–1916), created for fine-tuning text-to-image diffusion models. Used to train the Hammershoi Flux1 LoRA.
Dataset Details
Dataset Description
A structured image-caption dataset of cropped and framed reproductions from Hammershøi's paintings, annotated with detailed captions designed for LoRA training. Each image… See the full description on the dataset page: https://huggingface.co/datasets/jejunepixels/Hammershoi-Training-Dataset.Brain-Tumor-MRI-Dataset-Trainingbratota_pilot_training_datazalo-ai-2025-training-data-v2
Zalo AI Challenge 2025 - RoadBuddy Training Data V2set
This dataset contains training data v2 for the RoadBuddy – Understanding the Road through Dashcam AI challenge from Zalo AI Challenge 2025.
Dataset Description
The challenge aims to build a driving assistant capable of understanding video content from dashcams to quickly answer questions about traffic signs, signals, and driving instructions in Vietnam.
Dataset Structure
Files
frames/: Directory… See the full description on the dataset page: https://huggingface.co/datasets/OpenHay/zalo-ai-2025-training-data-v2.
