datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ade20k-panoptic-demo
Dataset Card for "ade20k-panoptic-demo"
More Information needed
PaddleOCR-VL_democollected_demos_traininghdr-demo-clips
HDR Demo Clips (Lightricks SDR→HDR)
Paired SDR (input) / HDR (output) frame sequences from the Lightricks SDR-to-HDR pipeline (IC-LoRA on LTX-2).
Each clip contains:
hdr_exr/frame_XXXXX.exr — HDR output (f16, linear Rec.709/sRGB primaries, scene-referred)
sdr_png/frame_XXXXX.png — SDR input (8-bit sRGB, display-referred)
thumbnail.jpg — 280px preview from the middle frame
Dimensions: HDR is symmetrically cropped from SDR to match model-friendly dimensions (typically 28–56px… See the full description on the dataset page: https://huggingface.co/datasets/oumoumad/hdr-demo-clips.Demo_videoCore-DEM
Major TOM Core-DEM
Major TOM Core-DEM contains a global coverage of Copernicus DEM, each of size 356 x 356 pixels.
This dataset was created to support the development of the MESA terrain generation model. It is also featured in the paper EarthEmbeddingExplorer: A Web Application for Cross-Modal Retrieval of Global Satellite Images and is part of the Major TOM: Expandable Datasets for Earth Observation ecosystem.
Official Viewer App:: Major TOM Viewer
Major TOM GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/Major-TOM/Core-DEM.Prague-REALMAP-Demo
Dataset description
Mosaic REALMAP Demo dataset with 9114 images collected with panoramic camera with global shutter sensors. You can find more about full dataset in this article: Mosaic Prague REALMAP Each image has its position written in EXIF metadata populated from IMU device paired with the camera.
Images from each sensor are electrically syncronized up to several nanoseconds and the whole device can be considered as a multicamera rig. Additional camera information, including… See the full description on the dataset page: https://huggingface.co/datasets/ishipachev/Prague-REALMAP-Demo.demonslayer
Bangumi Image Base of Demon Slayer
This is the image base of bangumi Demon Slayer, we detected 78 characters, 5890 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/demonslayer.processed_demo
Dataset Card for "processed_demo"
More Information needed
PiLoT-demo
PiLoT Demo
Open demo assets for the PiLoT project (query images, poses, pretrained weights, and scene models used by the demo scripts).
Download
pip install -U huggingface_hub
hf download choyaa/PiLoT-demo --repo-type dataset --local-dir ./PiLoT-demo
Layout:
data_demo/
├── 3dgs_model/
├── smbu_model/
├── pretrained_model/
└── query/
See data_demo/README.md for pose formats and case details.
Related
Training data (gated):… See the full description on the dataset page: https://huggingface.co/datasets/choyaa/PiLoT-demo.demo_source_datacs2-demo-archive
CS2 Demo to Dataset — Pipeline Output Samples
Sample archives produced by the open-source
cs2-demo-to-dataset
pipeline, which converts a single CS2 .dem replay file into per-round,
per-player first-person video aligned to tick-level state, input and event
tables.
This release is not a dataset contribution. The point of the upload is to
demonstrate that the pipeline produces a coherent, reproducible archive
format. Please see the
GitHub repository for the
recorder code… See the full description on the dataset page: https://huggingface.co/datasets/Vasy7777/cs2-demo-archive.collected_demosmeteor-demo-scenes
METEOR demo scenes for Autoware (meteor-demo-scenes)
Six short driving scenes, one per road type (147–148 frames each, 8 synchronised cameras, ego motion, LiDAR raster) to run
the released METEOR model (AutowareFoundation/meteor) and the demo renderers of
https://github.com/tier4/METEOR without access to the training corpus. These are the exact scene
roots the Orin demos and benchmarks in the repository refer to (valday, valcurve, fast).
The scenes come from the validation split… See the full description on the dataset page: https://huggingface.co/datasets/AutowareFoundation/meteor-demo-scenes.demo_openai_clip_index
CLIP index — demo_openai_clip
Precomputed image embeddings (openai/clip-vit-base-patch32) for the static Space
diegoolguinw/demo_openai_clip.
Current contents: 9000 images from the train split of
detection-datasets/coco.
(HF datasets cap a directory at 10 000 files, so thumbs/ stays below that.)
File
Description
embeddings.f16.bin
[N, 512] row-major float16, L2-normalized
metadata.json
index-aligned list: {id, file, thumb, width, height}
manifest.json
model, count… See the full description on the dataset page: https://huggingface.co/datasets/diegoolguinw/demo_openai_clip_index.robomme-demo-frames
RoboMME demonstration-prefix frames with foreground masks
The demonstration prefix of RoboMME episodes -- the video-instruction frames, where info/is_video_demo
is True -- together with a binary foreground mask for each frame.
These frames are worth publishing separately because they are missing from the usual export: the pickle
conversion that training pipelines consume writes a file only when is_demo is False, so the demo prefix
(292k of RoboMME's 769k frames) exists only in… See the full description on the dataset page: https://huggingface.co/datasets/hmkang/robomme-demo-frames.lingbot-map-demoEnvironmental_geology-dem
Digital Elevation Model
Overview
This repository contains the Digital Elevation Model dataset, part of the Environmental Geology suite maintained by NORA Research Lab. The data is processed into Cloud-Optimized GeoTIFF (COG) format for efficient geospatial analysis.
Dataset Details
Variable: Digital Elevation Model
Units: metres
Format: Cloud-Optimized GeoTIFF (COG)
Compression: ZSTD
Coordinate System: EPSG:4326 (WGS84)
Spatial Resolution: ~30… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/Environmental_geology-dem.DemirBlackroadcbt-dataChartBench-Demo
ChartBench: A Benchmark for Complex Visual Reasoning in Charts
Introduction
We propose the challenging ChartBench to evaluate the chart recognition of MLLMs.
We improve the Acc+ metric to avoid the randomly guessing situations.
We collect a larger set of unlabeled charts to emphasize the MLLM's ability to interpret visual information without the aid of annotated data points.
Todo
Open source all data of ChartBench.
Open source the evaluate… See the full description on the dataset page: https://huggingface.co/datasets/SincereX/ChartBench-Demo.video-demodemichanwakataritai
Bangumi Image Base of Demi-chan Wa Kataritai
This is the image base of bangumi Demi-chan wa Kataritai, we detected 16 characters, 1889 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/demichanwakataritai.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.Uranus-Demo-Data
Download
Download and unpack the samples from either Hugging Face or ModelScope into examples/data/.
Hugging Face
pip install -U huggingface_hub
hf download D-Robotics/Uranus-Demo-Data \
--repo-type dataset \
--local-dir ./examples/data
ModelScope
pip install -U modelscope
modelscope download \
--dataset D-Robotics/Uranus-Demo-Data \
--local_dir ./examples/data
Each episode lands in ./examples/data/<episode_id>/ and can be passed… See the full description on the dataset page: https://huggingface.co/datasets/D-Robotics/Uranus-Demo-Data.vla0-context-trace-full-demo-datasetdemultiplexversion https://git-lfs.github.com/spec/v1
oid sha256:282388825b64081bb30e34ff900e8db5f7b467c754c38798c3aa80e992cc3e59
size 712
10132025_human_demonstrations_tiktokmug_oceantablecloth_reorientThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 151,
"total_frames": 6731,"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:151"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/10132025_human_demonstrations_tiktokmug_oceantablecloth_reorient.RoboTwin-adjust_bottle-official-demo_clean50-Pi0_processed-datademo-maltesersThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha-stationary",
"total_episodes": 90,
"total_frames": 32336,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ccop/demo-maltesers.
