datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikitext
Dataset Card for "wikitext"
Dataset Summary
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified
Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License.
Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over
110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/wikitext.lotsa_data
LOTSA Data
The Large-scale Open Time Series Archive (LOTSA) is a collection of open time series datasets for time series forecasting.
It was collected for the purpose of pre-training Large Time Series Models.
See the paper and codebase for more information.
Citation
If you're using LOTSA data in your research or applications, please cite it using this BibTeX:
BibTeX:
@article{woo2024unified,
title={Unified Training of Universal Time Series Forecasting Transformers}… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lotsa_data.xlam-function-calling-60k
APIGen Function-Calling Datasets
Paper | Website | Models
This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness.
We conducted human evaluation over 600 sampled data points… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k.3d_optical_flow_droid
3D Optical Flow DROID Dataset
Processed DROID robotics dataset with optical flow and scene flow annotations.
Dataset Structure
Organized by lab, each trajectory in separate tar.gz archive:
IPRL/IPRL+2023-06-19+Mon_Jun_19_23:27:48_2023.tar.gz
CLVR/CLVR+2023-...tar.gz
... (15 labs, ~33K trajectories)
Each trajectory contains:
metadata.json - Trajectory metadata
trajectory.h5 - Robot state and actions
camera_left/, camera_right/ - Camera data
rgb/ - RGB images
depth/ -… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/3d_optical_flow_droid.blip3-kale
🥬 BLIP3-KALE:Knowledge Augmented Large-scale Dense Captions
BLIP3-KALE is an open-source dataset of 218 million image-text pairs, featuring knowledge-augmented dense captions combining web-scale knowledge with detailed image descriptions.
Paper: [To be added]
Uses
BLIP3-KALE is designed to facilitate research in multimodal pretraining. The dataset can be used for training large multimodal models that require factually grounded, dense image captions. It has already been an… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-kale.cos_e
Dataset Card for "cos_e"
Dataset Summary
Common Sense Explanations (CoS-E) allows for training language models to
automatically generate explanations that can be used during training and
inference in a novel Commonsense Auto-Generated Explanation (CAGE) framework.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
v1.0
Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/cos_e.APIGen-MT-5k
Summary
APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications
This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
Code: https://github.com/apigen-mt/apigen-mt.github.io
The repo contains 5000 multi-turn trajectories collected by APIGen-MT
This dataset is a subset of the data used to train the xLAM-2 model… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/APIGen-MT-5k.self-improve-fragilityUniDoc-Bench
UNIDOC-BENCH Dataset
A unified benchmark for document-centric multimodal retrieval-augmented generation (MM-RAG).
Dataset Description
UNIDOC-BENCH is the first large-scale, realistic benchmark for multimodal retrieval-augmented generation (MM-RAG) and Visual Question Answering (VQA) built from 70,000 real-world PDF pages across eight domains. The dataset extracts and links evidence from text, tables, and figures, then generates 1,700+ multimodal QA pairs spanning… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/UniDoc-Bench.Salesforce-xlam-function-calling-60kblip3-ocr-200m
BLIP3-OCR-200M Dataset
Overview
The BLIP3-OCR-200M dataset is designed to address the limitations of current Vision-Language Models (VLMs) in processing and interpreting text-rich images, such as documents and charts. Traditional image-text datasets often struggle to capture nuanced textual information, which is crucial for tasks requiring complex text comprehension and reasoning.
Key Features
OCR Integration: The dataset incorporates Optical Character… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-ocr-200m.CRMArenaPro
Dataset Card for CRMArena-Pro
Dataset Description
Paper Information
Citation
Dataset Description
CRMArena-Pro is a benchmark for evaluating LLM agents' ability to perform real-world work tasks in realistic environment. It expands on CRMArena with nineteen expert-validated tasks across sales, service, and "configure, price, and quote" (CPQ) processes, for both Business-to-Business (B2B) and Business-to-Customer (B2C) scenarios. CRMArena-Pro distinctively incorporates… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/CRMArenaPro.LiveResearchBench
Dataset Overview
LiveResearchBench provides expert-curated, real-world tasks spanning daily life, enterprise, and academia, each requiring extensive, real-time web search, multi-source reasoning, and cross-domain synthesis. DeepEval offers human-aligned protocols for reliable, systematic evaluation of agentic systems on open-ended deep research tasks.
📌 Quick Links
Project Page
Paper
Codebase
Dataset Fields
Subsets:
question_with_checklist: Full dataset with… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/LiveResearchBench.FinTrain
💰 Demystifying Domain-adaptive Post-training for Financial LLMs
This is the training data used in the recipe described in our paper:📄 Demystifying Domain-adaptive Post-training for Financial LLMs
For more details, please check the following resources:
🌐 Project Page: https://vincent950129.github.io/adapt-llm/
📚 Trained Model: https://huggingface.co/Salesforce/Llama-Fin-8b
🧠 Evaluation Data: https://huggingface.co/datasets/Salesforce/FinEval
💻 Code Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FinTrain.CRMArena
Dataset Card for CRMArena
Dataset Description
Paper Information
Citation
Dataset Description
CRMArena is a benchmark for evaluating LLM agents' ability to perform real-world work tasks in realistic environment. This benchmark is introduced in the paper "CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments". We include 16 commonly-used industrial objects (e.g., account, order, knowledge article, case) with… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/CRMArena.FaithEval-counterfactual-v1.0
FaithEval
FaithEval is a new and comprehensive benchmark dedicated to evaluating contextual faithfulness in LLMs across three diverse tasks: unanswerable, inconsistent, and counterfactual contexts.
[Paper] FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows", ICLR 2025, https://arxiv.org/abs/2410.03727
[Code and Detailed Instructions] https://github.com/SalesforceAIResearch/FaithEval
Disclaimer and Ethical Considerations… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FaithEval-counterfactual-v1.0.Salesforce-xlam-function-calling-60kfineweb_deduplicated
TL;DR
Fineweb is a popular and high quality open dataset. This dataset is a deduplicated version of Fineweb - removing rows with duplicate text, collecting counts.
Motivation
Fineweb is an open text dataset intended for training language models. It's one of the highest quality and most popular open datasets available. It has been produced by a reputable AI lab - HuggingFace and has been downloaded tens of thousands of times.
Fineweb dataset is 93.4 TB and has 15T… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/fineweb_deduplicated.FaithEval-unanswerable-v1.0
FaithEval
FaithEval is a new and comprehensive benchmark dedicated to evaluating contextual faithfulness in LLMs across three diverse tasks: unanswerable, inconsistent, and counterfactual contexts.
[Paper] FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows", ICLR 2025, https://arxiv.org/abs/2410.03727
[Code and Detailed Instructions] https://github.com/SalesforceAIResearch/FaithEval
Disclaimer and Ethical Considerations… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FaithEval-unanswerable-v1.0.FaithEval-inconsistent-v1.0
FaithEval
FaithEval is a new and comprehensive benchmark dedicated to evaluating contextual faithfulness in LLMs across three diverse tasks: unanswerable, inconsistent, and counterfactual contexts.
[Paper] FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows", ICLR 2025, https://arxiv.org/abs/2410.03727
[Code and Detailed Instructions] https://github.com/SalesforceAIResearch/FaithEval
Disclaimer and Ethical Considerations… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FaithEval-inconsistent-v1.0.ContextualBench
ContextualBench - A comprehensive toolkit to evaluate LM on different Contextual datasets
Evaluation Code: SalesforceAIResearch/SFR-RAG
Description
ContextualBench is a powerful evaluation framework designed to assess the performance of Large Language Models (LLMs) on contextual datasets. It provides a flexible pipeline for evaluating various LLM families across different tasks, with a focus on handling large context inputs.
Users need to make their own assessment… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/ContextualBench.cota-mantis
🌮 TACO: Learning Multi-modal Action Models with Synthetic Chains-of-Thought-and-Action
🌐 Website | 📑 Arxiv | 💻 Code| 🤗 Datasets
If you like our project or are interested in its updates, please star us :) Thank you! ⭐
Summary
TLDR: CoTA is a large-scale dataset of synthetic Chains-of-Thought-and-Action (CoTA) generated by multi-modal large language models.
Load data
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/cota-mantis.grounding_dataset
Grounding Dataset
A comprehensive, high-quality dataset for GUI element grounding tasks, curated from multiple authoritative sources to provide diverse, well-annotated interface interactions.
Overview
This dataset combines and standardizes annotations from five major GUI interaction datasets:
Aria-UI
OmniAct
Widget Caption
UI-Vision
OS-Atlas
Dataset Schema
Each sample contains the following fields:
Field
Type
Description
Example
dataset
string… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/grounding_dataset.GiftEvalParquet
GiftEval Parquet Collection
This repository hosts the parquet formatted GiftEval test data for ease of evaluating with LLM backboned models. Each dataset in the original GiftEval dataset can be loaded separately using the config names: datasetName_freq_term. Each row is a sample window from the test split of data, generated using the original GiftEval proressing script.
Each entry contains the following fields:
item_id (string): e.g. "item_0_dim0_window0/2018-04-12 20:00:00"… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/GiftEvalParquet.ST-Evidence-Instruct
ST-Evidence-Instruct Dataset
This dataset contains spatiotemporal evidence-based video question answering data for training.
This dataset was generated using Gemini and should not be used to develop models that compete with Google.
This project also uses the Segment Anything Model 3 (SAM 3) distributed by Meta Platforms, Inc. Use of SAM 3 is subject to the SAM License.
This was released for research purposes only, in support of the academic paper Evidence-Backed Video Question… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/ST-Evidence-Instruct.ProVision-10M
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
ProVision is an extendable data generation engine which produces instruction data for large multimodal language models (MLMs).
In particular, it synthesizes instruction data via data generators (Python programs) and scene graphs rather than proprietary models. It also includes a scene graph generation pipeline consisting of various state-of-the-art models (eg, object detection model). Thus… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/ProVision-10M.FinEval
💰 Demystifying Domain-adaptive Post-training for Financial LLMs
This is the evaluation data used in the recipe described in our paper:📄 Demystifying Domain-adaptive Post-training for Financial LLMs
For more details, please check the following resources:
🌐 Project Page: https://vincent950129.github.io/adapt-llm/
📚 Trained Model: https://huggingface.co/Salesforce/Llama-Fin-8b
🧠 Training Data: https://huggingface.co/datasets/Salesforce/FinTrain
💻 Code Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FinEval.blip3-grounding-50m
BLIP3-GROUNDING-50M Dataset
Overview
The BLIP3-GROUNDING-50M dataset is designed to enhance the ability of Vision-Language Models (VLMs) to ground semantic concepts in visual features, which is crucial for tasks like object detection, semantic segmentation, and understanding referring expressions (e.g., "the object to the left of the dog"). Traditional datasets often lack the necessary granularity for such tasks, making it challenging for models to accurately localize and… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-grounding-50m.Hard2Verify
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
Large language model (LLM)-based reasoning systems have recently achieved gold medal-level performance in the IMO 2025 competition, writing mathematical proofs where, to receive full credit, each step must be not only correct but also sufficiently supported. To train LLM-based reasoners in such challenging, open-ended settings, strong verifiers capable of catching step-level mistakes are necessary… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/Hard2Verify.hermes_salesforce_apigen_tool_use
