datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
robocurate-synth100
synth100 — 100 generated clips for validating Pre-Contact Level Filtering
100 episodes drawn (seed 20260824) from the 952-episode multi-object generation set, packaged so
Stage-5 filtering can be run on them without re-deriving anything. Every input the filter needs
travels with the package, in the space it is consumed in.
Read section 1 before using this. The single most important fact about this data is not in the
file layout, and getting it wrong invalidates any score… See the full description on the dataset page: https://huggingface.co/datasets/glory-hyeok/robocurate-synth100.Aperture_Lab_Synthetic_Aperture_Sonar_v1
ApertureLab Synthetic SAS Dataset
Version 1.0 (September 2026). Author: Isaac Gerg. Made with
ApertureLab; samples, statistics and the
generation pipeline are described on the
dataset page.
1000 simulated synthetic aperture sonar (SAS) images, each an 80 m along-track
by 200 m range swath from a HISAS 1030-class 100 kHz sonar on a straight
track, beamformed by time-domain back-projection at 2.5 cm pixels and
delivered as dynamic-range-compressed (DRC) TIFF LZW images with COCO… See the full description on the dataset page: https://huggingface.co/datasets/idg101/Aperture_Lab_Synthetic_Aperture_Sonar_v1.Synth-Text-Eng-512x128
Synthetic Text Images (English)
A synthetic dataset of rendered text images with rich per-sample
annotations: the text itself, its rendering attributes, background
description, applied post-processing, and a natural-language caption.
Each image is generated by compositing English text over a procedurally
generated background with random font, color, position, rotation, blur,
brightness and noise. All samples are accompanied by a structured
metadata.csv and a ready-to-use… See the full description on the dataset page: https://huggingface.co/datasets/Nininkkka/Synth-Text-Eng-512x128.SyntheticGenV5
SyntheticGenV5
SyntheticGenV5 is a synthetic remote-sensing semantic segmentation dataset (from the paper https://huggingface.co/papers/2602.04749) built for Urban–Rural domain-aware learning.
It keeps the original folder layout and uses Train/metadata.csv to connect each image with its semantic mask and RGB mask.
Why use this dataset?
🌆 Two domains: Urban and Rural
🛰️ Designed for remote-sensing semantic segmentation
🧪 Useful for synthetic augmentation and… See the full description on the dataset page: https://huggingface.co/datasets/buddhi19/SyntheticGenV5.ccs_synthetic_translated_arabic_processedccs_synthetic_translated_arabicThe columns inside the dataset as follows:
index
url
caption_en
caption_ar
The dataset size is 12556500 rows × 4 columns
SyntheticKonkle
