datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vlm_evaluation_v1.0
Datacard
This dataset is the evaluation VLM dataset used in VLABench. It is designed to evaluate the planning capabilities of Vision-Language Models (VLMs) in embodied scenarios.
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code: https://github.com/OpenMOSS/VLABench
Uses
The dataset structure is as follows:
vlm_evaluation_v1.0/
├── CommenSence/
├── add_condiment_common_sense/
├──… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlm_evaluation_v1.0.vlabench_primitive_ft_lerobotomnid_vlabench_dataset_v_0vlabench_select_paintingvlm_evaluation_v0_testThe v0 dataset is designed to evaluate the capabilities of VLMs in a non-interactive manner. This initial version primarily serves to help readers understand the structure and design of our benchmark.
Each ability dimension is represented by a dedicated directory, and within each ability directory are multiple task-specific subdirectories. For each task, there are numerous data examples. Each example includes original multi-view images, segmented images for visual prompts, a corresponding… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlm_evaluation_v0_test.vlabench_primitive_ft_lerobot_224
vlabench_primitive_ft_lerobot_224
VLABench/vlabench_primitive_ft_lerobot (the official pi0.5 VLABench fine-tuning set, LeRobot format) with every
camera image re-encoded from 480x480 to 224x224 (bit-identical to openpi's resize_with_pad applied at load time),
so that training does not pay for the 480px decode + resize and the LeRobot Arrow cache stays small.
Same episodes, same schema and metadata; only the image resolution differs. Made with… See the full description on the dataset page: https://huggingface.co/datasets/SeonghoonYu/vlabench_primitive_ft_lerobot_224.vlabench_primitive_etc
VLABench Primitive ETC
This release contains VLABench primitive ETC assets from two independent parts:
primitive/
primitive_track2/
Each part is kept self-contained at the top level. Its annotations, PNG tar
shards, index, previews, and manifest are stored under the corresponding
directory. The two parts are not merged.
PNG images are stored as uncompressed tar/WebDataset-style shards instead of
hundreds of thousands of individual PNG files. This avoids Hugging Face
repository… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlabench_primitive_etc.vlabench_select_single_painting_wo_button
