datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
blip3-kale
🥬 BLIP3-KALE:Knowledge Augmented Large-scale Dense Captions
BLIP3-KALE is an open-source dataset of 218 million image-text pairs, featuring knowledge-augmented dense captions combining web-scale knowledge with detailed image descriptions.
Paper: [To be added]
Uses
BLIP3-KALE is designed to facilitate research in multimodal pretraining. The dataset can be used for training large multimodal models that require factually grounded, dense image captions. It has already been an… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-kale.FinTrain
💰 Demystifying Domain-adaptive Post-training for Financial LLMs
This is the training data used in the recipe described in our paper:📄 Demystifying Domain-adaptive Post-training for Financial LLMs
For more details, please check the following resources:
🌐 Project Page: https://vincent950129.github.io/adapt-llm/
📚 Trained Model: https://huggingface.co/Salesforce/Llama-Fin-8b
🧠 Evaluation Data: https://huggingface.co/datasets/Salesforce/FinEval
💻 Code Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FinTrain.fineweb_deduplicated
TL;DR
Fineweb is a popular and high quality open dataset. This dataset is a deduplicated version of Fineweb - removing rows with duplicate text, collecting counts.
Motivation
Fineweb is an open text dataset intended for training language models. It's one of the highest quality and most popular open datasets available. It has been produced by a reputable AI lab - HuggingFace and has been downloaded tens of thousands of times.
Fineweb dataset is 93.4 TB and has 15T… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/fineweb_deduplicated.FinEval
💰 Demystifying Domain-adaptive Post-training for Financial LLMs
This is the evaluation data used in the recipe described in our paper:📄 Demystifying Domain-adaptive Post-training for Financial LLMs
For more details, please check the following resources:
🌐 Project Page: https://vincent950129.github.io/adapt-llm/
📚 Trained Model: https://huggingface.co/Salesforce/Llama-Fin-8b
🧠 Training Data: https://huggingface.co/datasets/Salesforce/FinTrain
💻 Code Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FinEval.RealUserSim
RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
Behavioral user profiles and evaluation benchmark for realistic LLM-powered user simulation, derived from the WildChat dataset.
Dataset Summary
This release contains:
7,273 behavioral user profiles extracted from real conversations, each containing demographics and executable linguistic style commands
600 evaluation test cases (6 splits x 100) for measuring user simulation fidelity… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/RealUserSim.lalm-judge-validation-full-duplex
LALM Judge Validation on Full-Duplex Voice Agents
Companion dataset for the paper A Reliability Assessment of
LALM Audio Judges for Full-Duplex Voice Agents.
This repository contains the anonymised ratings, adversarial-defect
recall tables, JSON schemas, and analysis scripts used to produce
every headline number, table, and figure in that paper.
Summary
209 rated stereo sessions: 152 full-duplex agent-client
conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.Salesforce__LLaMA-3-8B-SFR-Iterative-DPO-R-details
Dataset Card for Evaluation run of Salesforce/LLaMA-3-8B-SFR-Iterative-DPO-R
Dataset automatically created during the evaluation run of model Salesforce/LLaMA-3-8B-SFR-Iterative-DPO-R
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Salesforce__LLaMA-3-8B-SFR-Iterative-DPO-R-details.shared-imagination
Dataset Card for Shared Imagination
This dataset contains the problems used in the paper Shared
Dataset Description
This dataset contains the questions generated for the investigations described in the TMLR paper Shared Imagination: LLMs Hallucinate Alike.
If you want to use this dataset to assess new models, please use the default config (i.e., datasets.load_dataset('Salesforce/shared-imagination')).
This config contains questions for which the four candidate choices… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/shared-imagination.salesforce-dataset-prepInstruSum
InstruSum
This is the dataset corresponding to our paper "Benchmarking Generation and Evaluation Capabilities of Large Language
Models for Instruction Controllable Summarization".
dataset
The dataset subset contains 100 human-written data examples by us.
Each example contains an article, a summary instruction, a LLM-generated summary, and a hybrid LLM-human summary.
human_eval
This subset contains human evaluation results for the 100 examples in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/InstruSum.salesforce_scifact_queriessalesforce-omnistudio-dataset-preptextvqa_mini_validation_Salesforce_blip2-flan-t5-xxl_ns_100textvqa_valid_Salesforce_blip2-flan-t5-xxl_ns_5000salesforce_cqadupstack_queries
