datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagenet1k-256-wdsThis is imagenet1k in webdataset format. Images are stored as jpg files. Every image has been resized to a maximum side length of 256. That means that if an image in the original dataset was 1000 by 500, the new size will be 256 by 128. Images with a maximum side length of under 256 were not resized.
The total size of all dataset files is 57.8 GB, there are 1,281,167 rows in the training split and 50,000 rows in the validation split.
Stereo4D_vlbm
Stereo4D (converted to VLBM format)
This dataset contains 4,687 sequences from the Stereo4D dataset converted to the VLBM-compatible format using preprocess_stereo4d.py. The sequences have been compressed into .tar.gz archives in chunks of 50 sequences per archive.
Scale
Metric
Value
Total sequences
4,687
Image resolution
512 x 512 px
Depth type
Sparse (projected from tracked 3D points)
Dataset Structure
Each sequence directory follows this… See the full description on the dataset page: https://huggingface.co/datasets/ZhengGuangze/Stereo4D_vlbm.scientific-stuff-1hi-stt-preprocessed-webdatasetSTRIDE-QA-Dataset
STRIDE-QA Dataset
📦 Dataset
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
Category
Description
Object-centric Spatial QA
Spatial relations between two… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset.nyu-depthv2-wds
Dataset Card for nyu-depthv2-wds
This is the NYU DepthV2 dataset, converted into the webdataset format. https://huggingface.co/datasets/sayakpaul/nyu_depth_v2/
There are 47584 samples in the training split, and 654 samples in the validation split.
I shuffled both the training samples, and the validation samples.
I also cropped 16 pixels from all sides of the image, and depth image. I did this because there is a white border around all images.
This is an example of the border… See the full description on the dataset page: https://huggingface.co/datasets/adams-story/nyu-depthv2-wds.MegaSynth-webdatasetscientific-stuff-2scientific-stuff-4stereo4d-lefteye-perspective
Dataset Summary
This dataset contains the left-eye rectified perspective views from the Stereo4D dataset (Paper). Each video is generated using the rectify.py script, which processes VR180 stereo videos to produce 512×512 video clips with a 60° field of view perspective camera. This dataset is intended to be used alongside the Stereo4D dataset annotations which can be found here.
This dataset is provided as-is for non-commercial research purposes only.
Download
git clone… See the full description on the dataset page: https://huggingface.co/datasets/KevinMathew/stereo4d-lefteye-perspective.stark-image
Dataset Card for Stark
🏠 Homepage | 💻 Github | 📄 Arxiv | 📕 PDF
List of Provided Model Series
Ultron-Summarizer-Series: 🤖 Ultron-Summarizer-1B | 🤖 Ultron-Summarizer-3B | 🤖 Ultron-Summarizer-8B
Ultron 7B: 🤖 Ultron-7B
🚨 Disclaimer: All models and datasets are intended for research purposes only.
Dataset Summary
Stark is a publicly available, large-scale, long-term multi-modal conversation dataset that encompasses a diverse range of social personas… See the full description on the dataset page: https://huggingface.co/datasets/passing2961/stark-image.ImagePulseV2-Edit-Structure
ImagePulseV2 Dataset - Image Structure
The ImagePulseV2 dataset is a collection we constructed for training the Diffusion Templates series of models. It comprises multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio
Model… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Structure.RoboTwin-Randomized-targzoverhead-traffic-anomalies
Overhead Traffic Anomalies (OTA)
Developed by: H. Lichtenberg, Starwit Technologies GmbH, 2025
License: Creative Commons Attribution Non Commercial Share Alike 4.0
Master Thesis: H. Lichtenberg, Anomaly Detection in Traffic Applications: A Probabilistic Forecasting Approach Based on Object Tracking, 2025
This dataset contains traffic anomalies from static overhead cameras (intersections/roundabouts).
It includes 32 anomaly categories with a relevance mapping for severity-aware… See the full description on the dataset page: https://huggingface.co/datasets/starwit/overhead-traffic-anomalies.StreamGaze_v2
StreamGaze Dataset
StreamGaze is a comprehensive streaming video benchmark for evaluating MLLMs on gaze-based QA tasks across past, present, and future contexts.
Companion dataset: The EgoGazeVQA dataset is hosted separately at Peanuttoad/gaze_dataset.
📁 Dataset Structure
streamgaze/
├── metadata/
│ ├── egtea.csv # EGTEA fixation metadata
│ ├── egoexolearn.csv # EgoExoLearn fixation metadata
│ └── holoassist.csv # HoloAssist… See the full description on the dataset page: https://huggingface.co/datasets/Peanuttoad/StreamGaze_v2.stylebench-sSIGGRAPH 2026 / ACM TOG Journal Track
yuxuan_good_dataset_sttstark-image-url
Dataset Card for Stark
🏠 Homepage | 💻 Github | 📄 Arxiv | 📕 PDF
List of Provided Model Series
Ultron-Summarizer-Series: 🤖 Ultron-Summarizer-1B | 🤖 Ultron-Summarizer-3B | 🤖 Ultron-Summarizer-8B
Ultron 7B: 🤖 Ultron-7B
🚨 Disclaimer: All models and datasets are intended for research purposes only.
Dataset Summary
Stark is a publicly available, large-scale, long-term multi-modal conversation dataset that encompasses a diverse range of social personas… See the full description on the dataset page: https://huggingface.co/datasets/passing2961/stark-image-url.mc_gui_cold_startjapanese-str-dataset-v1
STR Dataset
Japanese STR (Scene Text Recognition) dataset in WebDataset format.
This dataset is composed of:
Images of Japanese named entities (full names and their affiliations)
Images of sentences retrieved from Aozora Bunko (青空文庫)
and their corresponding ground truth texts.
All images are synthesized using TRDG.
Dataset Structure
Split
Samples
Shards
train
10,000,000
1000
valid
50,000
5
test
50,000
5
Total
10,100,000
1010
Usage
import… See the full description on the dataset page: https://huggingface.co/datasets/nagohachi/japanese-str-dataset-v1.streethazardscomponent-static-buildsyuxuan_dataset_sttstormer-40yrstrain: 1979~2018 val: 2019 test: 2020
HF_HUB_ENABLE_HF_TRANSFER=1 hf download —repo-type dataset —local-dir /workspace/stormer/ KyleBae1017/stormer-40yrs
cat wb2_h5df.tar.part-* | tar -xvf -
Warwick-STEM
Warwick STEM Dataset (WebDataset)
A collection of 19,769 experimental scanning transmission electron microscopy (STEM) images from the University of Warwick, spanning hundreds of diverse materials projects collected between 2010 and 2018.
Dataset Description
This dataset contains experimental STEM images originally published as part of the Warwick Electron Microscopy Datasets by Jeffrey Ede. The images cover a wide range of materials and imaging conditions, making them… See the full description on the dataset page: https://huggingface.co/datasets/Stemson-AI/Warwick-STEM.wds_dollar_streetnsfw-video-still-caption-grid-onlystortinget_speech_corpus_v1.0
Dataset Card for Stortinget Speech Corpus V1.0
Overview
This is the WebDataset version of the Stortinget Speech Corpus V1.0, originally created by the National Library of Norway. We re-organize it into WebDataset format for better usability.
The Stortinget Speech Corpus (SSC) is a 5000+ hours speech dataset for weak supervision ASR created from audio andaligned proceedings text from Stortinget, the Norwegian Parliament. For more information, please refer to the original… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/stortinget_speech_corpus_v1.0.ElysiumTrack-1M
Dataset Card
ElysiumTrack-1M dataset is a million-scale object perception video dataset. It supports the following tasks:
Single Object Tracking (SOT): Predicting the location of a specific object in consecutive frames by referencing its initial position in the first frame.
Referring Single Object Tracking (RSOT): Identifying and locating a specific object within an entire video based on the given language expression. This task provides a more flexible tracking format and… See the full description on the dataset page: https://huggingface.co/datasets/sty-yyj/ElysiumTrack-1M.ruslan-stressed
RUSLAN with Word Stress Marks · RUSLAN с проставленными ударениями
English / Русский
English
What is this?
A drop-in replacement for the metadata of the RUSLAN Russian single-speaker
TTS corpus, with word-stress marks added to every multi-syllabic Russian
word in the transcripts. Audio is bundled unchanged.
The motivation is to train Russian TTS models (e.g. Kokoro, Tacotron, VITS,
StyleTTS, XTTS) that pronounce words with correct lexical stress.
Vanilla… See the full description on the dataset page: https://huggingface.co/datasets/stilletto/ruslan-stressed.
