datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EgoHaFL
EgoHaFL: Egocentric 3D Hand Forecasting Dataset with Language Instruction
EgoHaFL is a dataset designed for egocentric (first-person) 3D hand forecasting with accompanying natural language instructions.
It was introduced in the paper SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting.
The dataset contains short video clips, text descriptions, camera intrinsics, and detailed MANO-based 3D hand annotations.
The dataset supports research in 3D hand… See the full description on the dataset page: https://huggingface.co/datasets/ut-vision/EgoHaFL.PVL-BA-Bench
PVL-BA-Bench
PVL-BA-Bench is a large-scale bundle adjustment benchmark dataset from Polar-vision Lab. It provides public release metadata, interactive browser viewers, and downloadable PVL-BA, COLMAP, and BAL packages for photogrammetric optimization research.
Release Links
Static release index and interactive viewers: https://pub-2c28bdf6e62548919c47727a9b969dda.r2.dev/index.html
Downloadable packages in this dataset repository: packages/{pvl-ba,colmap… See the full description on the dataset page: https://huggingface.co/datasets/Polar-vision/PVL-BA-Bench.blv-assistive-vision-unified-viInfiniBench_challlenge
Download
git lfs install
git clone https://huggingface.co/datasets/Vision-CAIR/InfiniBench_challlenge
cd InfiniBench_challlenge
rm -rf .git/
How to Decompress the Videos
To extract the videos from the compressed files, follow these steps:
Open a terminal and navigate to the videos directory:
cd videos
Combine all the split archive parts into a single .tar.gz file:
cat test_videos.tar.gz.part_* > test_videos.tar.gz
Extract the contents of the archive:
tar -xvf… See the full description on the dataset page: https://huggingface.co/datasets/Vision-CAIR/InfiniBench_challlenge.dino-data-vision-tooling-preview
Dino Data Vision Tooling Preview
What This Dataset Is
This dataset is a focused vision-tooling preview built from two Dino Data capability slices:
image context understanding
image tooling
The goal is to train or inspect assistant behavior for image-related tasks where visual context, multimodal interpretation, or tool-aware image handling is relevant.
Included Capability Slices
Source lane
Public task name
What it teaches… See the full description on the dataset page: https://huggingface.co/datasets/DinoDS/dino-data-vision-tooling-preview.zamai-pashto-vision
ZamAI-Pashto Vision
This repository contains the dataset and scripts for the ZamAI-Pashto vision project, focusing exclusively on Pashto-related visual data, including scene captions, signs, and cultural objects.
Project Structure
images/: Raw and cleaned images, including signs, boards, and documents with Pashto text.
annotations/: Pashto captions, object bounding boxes, and scene labels.
metadata/: Cultural and location-specific tags.
scripts/: Image… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-vision.brfss-vision-module-data-vision-and-eye-health
BRFSS Vision Module Data – Vision & Eye Health
Description
2005-2016. This dataset includes data from the retired BRFSS Vision Module. From 2005-2011 the BRFSS employed a ten question vision module regarding vision impairment, access and utilization of eye care, and self-reported eye diseases. In 2013 and subsequently, one question in the core of BRFSS asks about vision: “Are you blind or do you have serious difficulty seeing, even when wearing glasses?” The latest data… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/brfss-vision-module-data-vision-and-eye-health.Language-Vision-Hallucinations
Dataset for Techen Project 095280
A comprehensive dataset for the Techen Project, focused on examining hallucinations in multi-modal AI-generated text by investigating model uncertainty, text generation patterns, and linguistic factors.
Columns Overview
image_link: URL to the image associated with each data row.
temperature: Temperature setting for text generation, controlling output randomness.
description: Text generated by the model for each image, using the… See the full description on the dataset page: https://huggingface.co/datasets/wrom/Language-Vision-Hallucinations.humanoid-object-vision-text-dataset
Humanoid Object Vision Text Dataset
Text-based object recognition data for humanoid perception and action mapping.
humanoid-vision-depth-dataset
Humanoid Vision Depth Dataset
Description
Depth camera statistics mapped to navigation actions for humanoid robots.
Designed for vision-based perception and safe movement in cluttered environments.
robot-vision-sample-dataset
Robot Vision Sample Dataset
This dataset contains simple labeled image references
for robotics vision and object recognition training.
Columns:
image_name
label
Size: Small experimental dataset
Version: 1.0
data_llama_vision_5secdata_llama_vision_10secdata_llama_vision_15secdata_llama_vision_20secdata_llama_vision_20_asrhumanoid-vision-distance-labels-v10humanoid-vision-lighting-labels-v9vision_model_init_datahumanoid-vision-context-labels-v6airletters-qualcomm
