datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
srtm-global-void-filledThis dataset mirrors the
Shuttle Radar Topography Mission (SRTM) Void Filled digital elevation data from USGS.
It consists of the 15,417 GeoTIFFs available on USGS EarthExplorer in the "SRTM Void Filled" (srtm_v2) dataset.
Each GeoTIFF covers 1x1 degrees.
The data is in WGS84, with a resolution of 1 arc-second/pixel in the United States and 3 arc-seconds/pixel elsewhere.
Coverage is limited to "80% of the Earth's land surface between 60° north and 56° south latitude".
The data is attributed to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/srtm-global-void-filled.VOID-Quadmask-Dataset
VOID-Compatible Quadmask Counterfactual Video Dataset
The first publicly available, pre-built quadmask-annotated counterfactual video dataset for physics-aware video object removal, inspired by and fully compatible with Netflix/VOID (arXiv:2604.02296).
What is this?
VOID introduced a powerful framework for removing objects from videos while correcting downstream physical interactions. Their key innovation is the quadmask — a 4-value segmentation mask that tells the… See the full description on the dataset page: https://huggingface.co/datasets/ErenAta00/VOID-Quadmask-Dataset.testVoidLinuxISOSPCA-EVALlunar-debris-and-voids
Lunar Debris and Voids
Instance segmentation dataset of geological features on the lunar surface,
derived from Lunar Reconnaissance Orbiter Camera (LROC) NAC imagery.
Classes
id
name
description
1
pit
Volcanic collapse / lava tube skylight
2
stone
Surface boulder / rock
3
crater
Impact crater rim / bowl
Format
COCO Detection 2017. Annotations under annotations/instances_{split}.json.
Images under images/{split}/.… See the full description on the dataset page: https://huggingface.co/datasets/F1nnSBK/lunar-debris-and-voids.grab50_2camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 51,
"total_frames": 16357,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/Voidx21/grab50_2cam.staring-into-the-void-runs
Staring Into the Void — cloud validation runs
Monte Carlo validation artifacts for
bshepp/staring-into-the-void,
a pipeline applying persistent homology to ensembles of forced-photometry
light curves. Each runs/<TIMESTAMP>/ folder contains the NPZ null
distribution, a JSON results manifest, the run log, and diagnostic plots
from one Hugging Face Jobs run of scripts/hf_jobs/null_sweep.py.
⚠️ Provenance notice (added 2026-07-15)
Both runs published to date… See the full description on the dataset page: https://huggingface.co/datasets/bshepp/staring-into-the-void-runs.nsfw1024elementstw-ocr
tw-ocr
Filtered OCR/document parsing records for Traditional Chinese OCR experiments.
The public Hugging Face upload is parquet-only for dataset rows. Every parquet
row embeds non-empty image bytes and is written with Hugging Face datasets
Image feature metadata. The physical parquet storage is the standard Image
struct ({bytes, path}) with image.path = null; when loaded through
datasets, the image column is an Image feature instead of a generic dict
column. The source path is… See the full description on the dataset page: https://huggingface.co/datasets/voidful/tw-ocr.Noob-Edit
Acknowledgements
The computing resources are sponsored by FishAudio & Feelin, All rights reserved to 39 AI Inc.
About data
Current version's data is for Noob-Edit-v0.1, including several editing tasks and natural language to anime image. Please refer to README in the folder for more info (Chinese only).
EikonomnemeEikontopos
