datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vqgan-pairs
VQGAN Pairs
This dataset contains ~2.4 million image pairs intended for improvement of image quality in VQGAN predictions. Each pair consists of:
A 512x512 crop of an image taken from Open Images.
A 256x256 image encoded and decoded using VQGAN, corresponding to the same image crop as the original.
This is the VQGAN implementation that was used for encoding and decoding: https://github.com/patil-suraj/vqgan-jax
License
This dataset is created using Open Images… See the full description on the dataset page: https://huggingface.co/datasets/dalle-mini/vqgan-pairs.RACER-Mini
RACER-Mini
RACER (Rationale-Aware Captioning of Edge-Case Driving Scenarios) is a reasoning caption dataset designed for training vision-language-action (VLA) models in autonomous driving.
This repository provides approximately 1,000 samples, as a small subset of the RACER dataset. Each sample consists of a temporal sequence of front camera images, the ego vehicle’s future trajectory, and a corresponding reasoning caption.
For details, please refer to our techblog RACER:… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/RACER-Mini.vggt-1m-wdsZOD-Mini-2D-Road-Scenes
ZOD-Mini-2D-Road-Scenes
The ZOD-Mini-2D-Road-Scenes dataset is derived from the Zenseact Open Dataset (ZOD), property of Zenseact AB (© 2022 Zenseact AB), and is licensed under the permissive CC BY-SA 4.0. Any public use, distribution, or display of this dataset must contain this entire notice:
For this dataset, Zenseact AB has taken all reasonable measures to remove all personally identifiable information, including faces and license plates. To the extent that you like to request… See the full description on the dataset page: https://huggingface.co/datasets/8bits-ai/ZOD-Mini-2D-Road-Scenes.open-sora-mini-hf-layered
Open-Sora Mini HF (Layered Format)
Converted from wangxingjun778/open-sora-mini-hf
Structure
├── videos/ # Video-only TAR shards (immutable)
│ ├── mixkit_shard_0000.tar
│ └── ...
├── annotations/ # Annotations (can add new versions)
│ └── captions_v1.parquet # Vision-generated captions
└── manifest.parquet # Index: sample_id -> shard mapping
Stats
Metric
Value
Total samples with captions
8,247… See the full description on the dataset page: https://huggingface.co/datasets/zengxianyu/open-sora-mini-hf-layered.open-sora-mini-hfmininet-1kmining-YT-datasetCoVLA-Dataset-Minintusldl2024_miniproject_2
Description
This comprehensive dataset is specifically designed for mini project 2 of the course SLDL at National Taiwan University. It is sourced from publicly available data provided by the Intellectual Property Office, MOEA (智慧財產局) in Taiwan. The images within this dataset are of different trademarks (商標). Each image is annotated by humans and belongs to multiple labels, with the labels representing the class that the trademark belongs to.
license: mit… See the full description on the dataset page: https://huggingface.co/datasets/dodofk/ntusldl2024_miniproject_2.match_mini
Remaining Source Images for MBridge
This dataset contains tar shards for the non-X2I2 image sources used by the
M experiments. OmniGen2/X2I2 video images are not
mirrored here; use the upstream OmniGen2/X2I2 dataset for video_edit and
video_icgen images.
Contents
images/<source>/<source>-NNNNNN.tar: image shards.
metadata/image_manifest.jsonl: one row per packed image.
metadata/pack_summary.json: packing configuration and counts.
Summary
{… See the full description on the dataset page: https://huggingface.co/datasets/youyou321/match_mini.GPUDrive-NuPlan-MiniSetruslan-stressed-mini
RUSLAN stressed — mini sanity-check dataset
This is a 200-sample mini version of stilletto/ruslan-stressed used to
verify that the WebDataset tar layout is parsed correctly by the HuggingFace
dataset viewer before the full 22,200-sample dataset is repacked the same way.
Layout (WebDataset):
mini_part_001.tar # samples 000000…000099 (wav + paired txt)
mini_part_002.tar # samples 000100…000199 (wav + paired txt)
Each tar contains paired files sharing a basename:
000000_RUSLAN.wav… See the full description on the dataset page: https://huggingface.co/datasets/stilletto/ruslan-stressed-mini.mini-nsd-boldvggt-1m-wds_valLME-graph_m-4o-miniLME-graph_s-4o-mini
