datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
morpheus-real-world
Morpheus — Real-World Physics Videos
Real-world reference footage for Morpheus, a benchmark that tests whether
video generative models (Wan, CogVideo, LTX-Video, COSMOS-predict1/2,
Pyramid-Flow, Veo3, Kling-Turbo, ...) obey Newtonian mechanics. Object
trajectories are extracted via SAM2 tracking and tested against physical laws
(energy/momentum conservation, equations of motion) rather than pixel-matched
to a single "correct" video.
This repo contains the filmed real-world… See the full description on the dataset page: https://huggingface.co/datasets/physics-from-video/morpheus-real-world.cross-channel-toxoplasma-from-cellmask
Cross-channel toxoplasma from cellmask dataset
3030 paired fields. images/ is the input channel, masks_pv/ the instance-labelled
parasitophorous-vacuole ground truth (same stem = same field), regenerated with the promoted
PV model cpsam_v2_toxo_r5.
Train/test annotation: fields.csv gives name,split,n_objects for every field -
2567 train (31637 objects) / 463 test (6116 objects).
The split is by well (split_by_well.csv), so no well contributes to both sides.
Prepared with spaCR… See the full description on the dataset page: https://huggingface.co/datasets/einarolafsson/cross-channel-toxoplasma-from-cellmask.diffusion_db_dedupe_from50k_train
Dataset Card for "diffusion_db_dedupe_from50k_train"
More Information needed
220k-GPT4Vision-captions-from-LIVIS
220k-GPT4Vision-captions-from-LVIS
by: Christoph Schuhmann, Peter Bevan, 21 Nov, 2023
This dataset comprises 220,000 captioned images from the LVIS dataset. The captions were generated by summarising the LVIS-Instruct4V dataset released by X2FD. The instructions are converted into captions using Mistral-7B-OpenOrca.
PROMPT
"""<<SYS>> You are a highly intelligent, empathic, helpful, respectful, and honest assistant with high emotional intelligence. Always… See the full description on the dataset page: https://huggingface.co/datasets/laion/220k-GPT4Vision-captions-from-LIVIS.sync_bigjob_8_finalised_processed_with_error_handling_from_51th_splitHUBT_from_a_drones_perspective
HUTB From a Drone's Perspective
Dataset Summary
HUTB From a Drone's Perspective is a synthetic, multi-map UAV dataset generated
with OpenHUTB/CARLA. It provides synchronized visible RGB, metric depth,
surface normals, target semantic segmentation, LiDAR, and vehicle/pedestrian
detection annotations from elevated oblique viewpoints.
The current release contains:
4,081 synchronized frames at 1920 x 1080 pixels.
8 simulated maps.
6 weather and illumination… See the full description on the dataset page: https://huggingface.co/datasets/yutiangu/HUBT_from_a_drones_perspective.cross-channel-nuclei-from-cellmask
Cross-channel nuclei-from-cellmask dataset
3029 paired fields. images/ is the
cell-mask input channel, masks/ the matching instance-labelled nuclei ground truth (same stem =
same field). split_by_well.csv gives the train/test assignment; the split is by well, so no
well contributes to both sides - field-level splitting would leak.
Prepared with spaCR
(PyPI | conda-forge).
Used to train einarolafsson/cross-channel-nuclei-from-cellmask-cpsam.
20260725_pinch_from_bowl_with_fingers_right_onlyPerson_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform
Person Detection and Re-Identification from Low Altitude UAV-based Platform
Dataset Description
This dataset was collected as part of a master's thesis on person detection and re-identification using low-altitude UAV (drone) footage. It contains labeled aerial images captured from a DJI Mini drone, annotated in YOLOv8 format.
The dataset supports two tasks:
Person Detection — detecting people in aerial drone footage
Person Re-Identification (Re-ID) — recognizing and… See the full description on the dataset page: https://huggingface.co/datasets/Mikiee/Person_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform.diffusion_db_dedup_from50k_train_v2
Dataset Card for "diffusion_db_dedup_from50k_train_v2"
More Information needed
easyr1-126k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-aug-jitter
easyr1-126k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-aug-jitter
Augmented version of datasets/easyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP with coordinate jitter.
For each original example, 1 additional copies were created. Each copy
randomly jitters the target coordinate by ±1 pixel in both X and Y. The
assistant coordinate in messages is updated, and bbox/normalized_bbox
are shifted when present.… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-126k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-aug-jitter.easyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP
easyr1-63k-hard-qwen7b-easy-gta1-nores-jedi-fix-synced-ui-vision-grounding-pro-apps-manually-labeled-icon-data-from-yt-4MP
Merged dataset composed of the following sources:
/Users/anasawadalla/Desktop/easyr1-57k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced (57011 samples in split train)
ui-vision-grounding-4MP (5790 samples in split train)
easyr1-v2-pro-apps-manually-labeled-icon-data-from-yt-4MP (230 samples in split train)
Summary
Generated on: 2025-09-10… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP.SFT-Data-From-Thyme
SFT Data From Thyme
GroundingDINO-assisted visual tool-call trajectories derived from Thyme-SFT.
Files
dataset_accepted_first10000_v12.jsonl: 9,795 accepted trajectories.
images/: source and generated crop images.
metadata/failed_after_three_rounds_first10000_v12.jsonl: 205 failed trajectories, excluded from training data.
metadata/aggregate_audit_first10000_v12.json: source aggregation audit.
Image references in the uploaded JSONL use repository-relative… See the full description on the dataset page: https://huggingface.co/datasets/albert13200/SFT-Data-From-Thyme.cross-channel-cell-from-hoechst
Cross-channel cell from hoechst dataset
3029 paired fields. images/ input channel, masks/ instance-labelled ground truth
(same stem = same field).
Train/test annotation: fields.csv gives name,split,n_objects for every field -
2578 train (237957 objects) / 451 test (45098 objects).
The split is by well ( split_by_well.csv ), so no well contributes to both sides.
Prepared with spaCR
(PyPI | conda-forge).
Used to train einarolafsson/cross-channel-cell-from-hoechst-cpsam.
laion-gpt4v-from-lavisdl3dv_2dMode_1Sub2D_TrueAlpha_fromRe10KPretraineddiffusion_db_dedupe_from50k_val
Dataset Card for "diffusion_db_dedupe_from50k_val"
More Information needed
diffusion_db_dedup_from10k_train_v2
Dataset Card for "diffusion_db_dedup_from10k_train_v2"
More Information needed
easyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-answer-keyFrom_Reasoning_Structure_to_the_Ancient_Problem_of_Primes
From Reasoning Structure to the Ancient Problem of Primes
Author: Zixi Li (Oz Lee)
Date: 2025
Publisher: Hugging Face
Citation
@misc{oz_lee_2025,
author = { Oz Lee },
title = { From_Reasoning_Structure_to_the_Ancient_Problem_of_Primes (Revision d9034a1) },
year = 2025,
url = { https://huggingface.co/datasets/OzTianlu/From_Reasoning_Structure_to_the_Ancient_Problem_of_Primes },
doi = { 10.57967/hf/7156 }… See the full description on the dataset page: https://huggingface.co/datasets/OzTianlu/From_Reasoning_Structure_to_the_Ancient_Problem_of_Primes.RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts
Dataset Attribution
The original dataset is available on Kaggle.
This dataset has been curated solely for ease of use within the Hugging Face ecosystem, with no intention of plagiarizing or copying the original work.
Please cite the original authors if you use this dataset.
Citation
@INPROCEEDINGS{8978005,
author={Chamchong, Rapeeporn and Gao, Wei and McDonnell, Mark D.},
booktitle={2019 International Conference on Document Analysis and Recognition (ICDAR)}… See the full description on the dataset page: https://huggingface.co/datasets/fwgpiyawudk/RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts.easyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-dense-rewardsummarize_from_feedback_oai_preprocessing_1711138793
Dataset Card for "summarize_from_feedback_oai_preprocessing_1711138793"
More Information needed
summarize_from_feedback_oai_preprocessing_1711138084
Dataset Card for "summarize_from_feedback_oai_preprocessing_1711138084"
More Information needed
Person_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform
Person Detection and Re-Identification from Low Altitude UAV-based Platform
Dataset Description
This dataset was collected as part of a master's thesis on person detection and re-identification using low-altitude UAV (drone) footage. It contains labeled aerial images captured from a DJI Mini drone, annotated in YOLOv8 format.
The dataset supports two tasks:
Person Detection — detecting people in aerial drone footage
Person Re-Identification (Re-ID) — recognizing… See the full description on the dataset page: https://huggingface.co/datasets/yuotub/Person_Detection_and_Re-Identification_from_Low_Altitude_UAV-based_platform.diffusion_db_dedup_from50k_val_v2
Dataset Card for "diffusion_db_dedup_from50k_val_v2"
More Information needed
Rice_Diagnosis_Leaf_Images_FromKaggle
AutoTrain Dataset for project: rice_diagnosis
Dataset Description
This dataset has been automatically processed by AutoTrain for project rice_diagnosis.
Originally from Kaggle, this shows rice leaves (leaf) up close pictures labeled with the disease of which they show symptoms.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image":… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/Rice_Diagnosis_Leaf_Images_FromKaggle.smolvla_remove_block_from_plate_V2Wildfire_Images_From_Satilite_For_Azerbaijansocio_economic_from_space
Predicting Socio-Economic Development in Poland (2020–2024) – Dataset
This dataset provides the resources necessary for the project Predicting Socio-Economic Development in Poland in 2020–2024 Using Satellite Imagery.
The corresponding codebase is available here: GitHub Repository.
Contents
Municipal incomes (2018–2024)Official statistics on municipal-level revenues in Poland from the Statistics Poland Local Data Bank (BDL).
Shapefiles of Polish municipalities… See the full description on the dataset page: https://huggingface.co/datasets/LukaszJanisiow/socio_economic_from_space.
