datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
3d_optical_flow_droid
3D Optical Flow DROID Dataset
Processed DROID robotics dataset with optical flow and scene flow annotations.
Dataset Structure
Organized by lab, each trajectory in separate tar.gz archive:
IPRL/IPRL+2023-06-19+Mon_Jun_19_23:27:48_2023.tar.gz
CLVR/CLVR+2023-...tar.gz
... (15 labs, ~33K trajectories)
Each trajectory contains:
metadata.json - Trajectory metadata
trajectory.h5 - Robot state and actions
camera_left/, camera_right/ - Camera data
rgb/ - RGB images
depth/ -… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/3d_optical_flow_droid.blip3-kale
🥬 BLIP3-KALE:Knowledge Augmented Large-scale Dense Captions
BLIP3-KALE is an open-source dataset of 218 million image-text pairs, featuring knowledge-augmented dense captions combining web-scale knowledge with detailed image descriptions.
Paper: [To be added]
Uses
BLIP3-KALE is designed to facilitate research in multimodal pretraining. The dataset can be used for training large multimodal models that require factually grounded, dense image captions. It has already been an… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-kale.UniDoc-Bench
UNIDOC-BENCH Dataset
A unified benchmark for document-centric multimodal retrieval-augmented generation (MM-RAG).
Dataset Description
UNIDOC-BENCH is the first large-scale, realistic benchmark for multimodal retrieval-augmented generation (MM-RAG) and Visual Question Answering (VQA) built from 70,000 real-world PDF pages across eight domains. The dataset extracts and links evidence from text, tables, and figures, then generates 1,700+ multimodal QA pairs spanning… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/UniDoc-Bench.blip3-ocr-200m
BLIP3-OCR-200M Dataset
Overview
The BLIP3-OCR-200M dataset is designed to address the limitations of current Vision-Language Models (VLMs) in processing and interpreting text-rich images, such as documents and charts. Traditional image-text datasets often struggle to capture nuanced textual information, which is crucial for tasks requiring complex text comprehension and reasoning.
Key Features
OCR Integration: The dataset incorporates Optical Character… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-ocr-200m.MTA-Vision-DeepSearchgrounding_dataset
Grounding Dataset
A comprehensive, high-quality dataset for GUI element grounding tasks, curated from multiple authoritative sources to provide diverse, well-annotated interface interactions.
Overview
This dataset combines and standardizes annotations from five major GUI interaction datasets:
Aria-UI
OmniAct
Widget Caption
UI-Vision
OS-Atlas
Dataset Schema
Each sample contains the following fields:
Field
Type
Description
Example
dataset
string… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/grounding_dataset.blip3-grounding-50m
BLIP3-GROUNDING-50M Dataset
Overview
The BLIP3-GROUNDING-50M dataset is designed to enhance the ability of Vision-Language Models (VLMs) to ground semantic concepts in visual features, which is crucial for tasks like object detection, semantic segmentation, and understanding referring expressions (e.g., "the object to the left of the dog"). Traditional datasets often lack the necessary granularity for such tasks, making it challenging for models to accurately localize and… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-grounding-50m.ST-Evidence-Bench
ST-Evidence Benchmark Dataset
ST-Evidence is a comprehensive benchmark for evaluating Spatial-Temporal Evidence generation in video understanding. It contains two tasks: Generation (Gen) and Multiple Choice Question (MCQ).
This was released for research purposes only, in support of the academic paper Evidence-Backed Video Question Answering.
Dataset Overview
Total Videos: ~1,300 videos at 6fps
Annotations: Question-Answer pairs with temporal segments and spatial… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/ST-Evidence-Bench.PROVE
Trust but Verify: Programmatic VLM Evaluation in the Wild
Viraj Prabhu, Senthil Purushwalkam, An Yan, Caiming Xiong, Ran Xu
Explorer
| Paper
| Quickstart
Vision-Language Models (VLMs) often generate plausible but incorrect responses to visual queries. However, reliably quantifying the effect of such hallucinations in free-form responses to open-ended queries is challenging as it requires visually verifying each claim within the response. We propose Programmatic VLM… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/PROVE.program-cota-mantis
🌮 TACO: Learning Multi-modal Action Models with Synthetic Chains-of-Thought-and-Action
🌐 Website | 📑 Arxiv | 💻 Code| 🤗 Datasets
If you like our project or are interested in its updates, please star us :) Thank you! ⭐
Summary
TLDR: CoTA is a large-scale dataset of synthetic Chains-of-Thought-and-Action (CoTA) generated by programs.
Load data
from datasets import load_dataset
dataset = load_dataset("Salesforce/program-cota-mantis"… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/program-cota-mantis.CogAlign
Dataset Card for CogAlign
Dataset Description
Citation
Dataset Description
CogAlign is a post-training strategy for Vision Language Models (VLMs) aimed at enhancing their visual arithmetic capabilities. This repository presents the training data for CogAlign, a synthetic dataset containing 64,000 examples designed to facilitate this post-training process.
CogAlign is inspired by Piaget's theory of cognitive development and focuses on improving a VLM's understanding of… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/CogAlign.cartoon-captioned-datasets-salesforce-blip
Dataset Card for "cartoon-captioned-datasets-salesforce-blip"
More Information needed
amazon_sales_datasetsynthetic_vqa_dataset_21.4k_images_salesforce_blip_vqa_baseclothing-sales-dsNYC_Sales_EDA
NYC Property Sales – Exploratory Data Analysis
Author: David Wilfand
Notebook: https://colab.research.google.com/drive/13FqJBe0YAL3wtOM0oRuMqPS5M2bVfTaF#scrollTo=OI0MZzohKwfE
Dataset source: NYC Property Sales (originally from Kaggle, re-hosted on this Hugging Face dataset)
Video walkthrough: https://youtu.be/jBrw9EJpCOo
Overview
This project performs an Exploratory Data Analysis (EDA) on tens of thousands of New York City property sales.The goal is to understand… See the full description on the dataset page: https://huggingface.co/datasets/KalsusEvening/NYC_Sales_EDA.NYC_Property_Sales_EDA
NYC Property Sales – Exploratory Data Analysis
Author: TODONotebook: TODO – link to your Colab / .ipynb on this datasetDataset source: NYC Property Sales (originally from Kaggle, re-hosted on this Hugging Face dataset)Video walkthrough: TODO – Loom / Zoom / YouTube link once recorded
Overview
This project performs an Exploratory Data Analysis (EDA) on tens of thousands of New York City property sales.The goal is to understand which property characteristics are most… See the full description on the dataset page: https://huggingface.co/datasets/KalsusEvening/NYC_Property_Sales_EDA.filesystem-huggingface-terminal-snowflake-5102-sales-report-4jtzxg
Quarterly Sales Performance Report
Report Period: Q2 2026
Generated At: 2026-08-16 07:22:52 UTC
Executive Summary
Total Revenue (Completed Orders): $177,675.79
Total Orders: 187.00
Average Order Value: $1,090.04
Revenue Target: $155,000.00
Target Status: TARGET MET
Top Region: EAST
Top Product Category: ELECTRONICS
Underperforming Regions: WEST
Business Recommendation
Maintain the current sales strategy and focus on expanding EAST leadership.… See the full description on the dataset page: https://huggingface.co/datasets/TianfuXinqu/filesystem-huggingface-terminal-snowflake-5102-sales-report-4jtzxg.clothing-sales-datafilesystem-huggingface-terminal-snowflake-5102-sales-report-cc9svz
Quarterly Sales Performance Report
Report Period: Q2 2026
Generated At: 2026-08-15 13:53
Executive Summary
Total Revenue (Completed Orders): $157,168.66
Total Orders: 187
Average Order Value: $1,099.08
Revenue Target: $140,000.00
Target Status: TARGET MET
Top Region: EAST
Top Product Category: ELECTRONICS
Underperforming Regions: WEST
Business Recommendation
Maintain the current sales strategy and focus on expanding EAST leadership.… See the full description on the dataset page: https://huggingface.co/datasets/TianfuXinqu/filesystem-huggingface-terminal-snowflake-5102-sales-report-cc9svz.vehicle_sales_data--
clothing-sales-data-embeddingssales-prediction-dsfilesystem_huggingface_terminal_emails_5084_sales_reportahmedsayed564_amazon-sales-dataset
Amazon Sales Dataset 📦
The Dataset Contains 1K+ Amazon Product's Ratings and Reviews.
Dataset Info
Source: Kaggle
Original Size: 1.71 MB
Kaggle Downloads: 4,158
Files: 1
Files
Amazon.csv
Mirrored from Kaggle
salessynthetic_vqa_dataset_100_images_salesforce_blip_vqa_baseSales
