web-data
Datasets
All datasets matching “web-data”waymo_webdatasetbiomedica_webdataset_24M
Dataset Card for Dataset Name
Arxiv: Arxiv
|
Website: Biomedica
|
Training instructions: OpenCLIP
|
Tutorial: Google Colab
BIOMEDICA Dataset is a large-scale, deep-learning-ready biomedical dataset containing over 24M imagecaption pairs and 30M image-references from 6M unique open-source articles. Each… See the full description on the dataset page: https://huggingface.co/datasets/BIOMEDICA/biomedica_webdataset_24M.conceptual-captions-12m-webdatasetWEB-Dataset
WorldEngine Bimanual Dataset for Post-training
A large-scale, language-annotated real-robot bimanual manipulation dataset for
post-training robotics foundation models. It spans 90 everyday manipulation tasks
collected with a bimanual YAM follower arm teleoperated by a GELLO leader,
recording joint state, action, and three synchronized camera streams at 60 Hz.
Shared lineage, different story. This dataset shares its hardware, teleoperation
setup, and recording pipeline with the… See the full description on the dataset page: https://huggingface.co/datasets/WorldEngineAI/WEB-Dataset.infinity-mm-stage1-webdataset
WebDataset Image-Text Dataset
This dataset contains image-text pairs in WebDataset format.
Dataset Structure
Each .tar.wds file contains entries with JSON data including image, text, and metadata.
conceptual-captions-12m-webdataset-metadata
Conceptual Captions 12M — Webshart metadata indices
Per-shard webshart metadata indices for
laion/conceptual-captions-12m-webdataset:
1,100 JSON files under data/, one per source tar shard, mirroring the source's shard layout.
Each index records every tar member's byte offset and length (enabling ranged reads without
downloading whole shards), image geometry (width/height for aspect bucketing), and — as of
August 2026 — embedded captions for all 10,994,853 samples, coalesced… See the full description on the dataset page: https://huggingface.co/datasets/webshart/conceptual-captions-12m-webdataset-metadata.
