datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
blip3-kale
🥬 BLIP3-KALE:Knowledge Augmented Large-scale Dense Captions
BLIP3-KALE is an open-source dataset of 218 million image-text pairs, featuring knowledge-augmented dense captions combining web-scale knowledge with detailed image descriptions.
Paper: [To be added]
Uses
BLIP3-KALE is designed to facilitate research in multimodal pretraining. The dataset can be used for training large multimodal models that require factually grounded, dense image captions. It has already been an… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-kale.FinTrain
💰 Demystifying Domain-adaptive Post-training for Financial LLMs
This is the training data used in the recipe described in our paper:📄 Demystifying Domain-adaptive Post-training for Financial LLMs
For more details, please check the following resources:
🌐 Project Page: https://vincent950129.github.io/adapt-llm/
📚 Trained Model: https://huggingface.co/Salesforce/Llama-Fin-8b
🧠 Evaluation Data: https://huggingface.co/datasets/Salesforce/FinEval
💻 Code Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FinTrain.saas-sales-conversations
saas-sales-conversations
Dataset Description
This is a synthetic dataset of sales conversations for SaaS (Software as a Service) companies, designed for training sales conversion prediction models. The dataset was created following the methodology presented in "SalesRLAgent: A Reinforcement Learning Approach for Real-Time Sales Conversion Prediction and Optimization" (Nandakishor M, 2025).
The dataset contains realistic dialogues between sales representatives and… See the full description on the dataset page: https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations.fineweb_deduplicated
TL;DR
Fineweb is a popular and high quality open dataset. This dataset is a deduplicated version of Fineweb - removing rows with duplicate text, collecting counts.
Motivation
Fineweb is an open text dataset intended for training language models. It's one of the highest quality and most popular open datasets available. It has been produced by a reputable AI lab - HuggingFace and has been downloaded tens of thousands of times.
Fineweb dataset is 93.4 TB and has 15T… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/fineweb_deduplicated.FinEval
💰 Demystifying Domain-adaptive Post-training for Financial LLMs
This is the evaluation data used in the recipe described in our paper:📄 Demystifying Domain-adaptive Post-training for Financial LLMs
For more details, please check the following resources:
🌐 Project Page: https://vincent950129.github.io/adapt-llm/
📚 Trained Model: https://huggingface.co/Salesforce/Llama-Fin-8b
🧠 Training Data: https://huggingface.co/datasets/Salesforce/FinTrain
💻 Code Repository:… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FinEval.dubai-real-estate-sales-transactions
Dubai Real Estate Sales Transactions (DLD, 2004–2026)
1,364,226 registered property sale transactions in Dubai, sourced directly from the Dubai Land Department (DLD) via the Dubai Pulse open-data platform. This is not a listings scrape and it is not asking prices: every row is a transaction actually registered with the government.
Published by Dubai Real Estate Data (dubairealestatedata.com), an independent analytics site that turns this exact DLD data into medians, trends… See the full description on the dataset page: https://huggingface.co/datasets/dubairealestatedata/dubai-real-estate-sales-transactions.kapibala-sales-dialogues
Kapibala Sales Dialogues
A sales-conversation dataset with outcome, conversation-level and sentence-level labels
🤗 Hugging Face · Annotation details · 中文
630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset:
L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/Thomasgudan/kapibala-sales-dialogues.saas-sales-bge-open
saas-sales-bge-open Dataset (Initial Setup)
This dataset is currently being uploaded. Please check back later for the complete dataset.
Gym_Salesman_Dataset
Gym Salesman Dataset
11,997 synthetic gym-membership sales conversations, each labelled SUCCESS or FAILURE.
🔗 Project Links
| Live App — practice against an AI customer | Hugging Face Space |
| Telegram Bot — practice on the go | @ido_salescoach_bot |
| Dataset — 11,997 labelled conversations | elg4/Gym_Salesman_Dataset |
| Data Generation — how the data was built | notebook |
| Recommendation — the embedding retriever | notebook |
Every conversation is a… See the full description on the dataset page: https://huggingface.co/datasets/elg4/Gym_Salesman_Dataset.RealUserSim
RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
Behavioral user profiles and evaluation benchmark for realistic LLM-powered user simulation, derived from the WildChat dataset.
Dataset Summary
This release contains:
7,273 behavioral user profiles extracted from real conversations, each containing demographics and executable linguistic style commands
600 evaluation test cases (6 splits x 100) for measuring user simulation fidelity… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/RealUserSim.sales-orders-us
US Sales Orders & Invoices Dataset
Complete sales pipeline for a simulated US retail SME: customers, orders, order lines, invoices, and payment allocations. Customer behaviour is stochastic with seasonal patterns — not uniform random. Payment timing varies by customer reliability.
Tables
Table
Rows
companies
1
customers
163
payment_allocations
2,271
sales_invoice_lines
11,065
sales_invoices
2,441
sales_order_lines
11,065
sales_orders
2,441
Total… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/sales-orders-us.Video-Games-Sales-EDA
🎮 Video Game Sales — Exploratory Data Analysis (EDA)
Video Presentation
If I didn’t cover everything it’s because I didn’t have enough time
Your browser does not support the video tag.
Executive Summary
The video game industry is a multi‑billion dollar market characterized by extreme unpredictability—a single "mega‑hit" can generate more revenue than thousands of average games combined. This project analyzes historical video game sales… See the full description on the dataset page: https://huggingface.co/datasets/0tizm0/Video-Games-Sales-EDA.store-sales-time-series-forecasting
taken from this Kaggle competition:
Dataset Description
In this competition, you will predict sales for the thousands of product families sold at Favorita stores located in Ecuador. The training data includes dates, store and product information, whether that item was being promoted, as well as the sales numbers. Additional files include supplementary information that may be useful in building your models.
File Descriptions and Data Field Information… See the full description on the dataset page: https://huggingface.co/datasets/mrcksggcfc/store-sales-time-series-forecasting.cafe-sales-cleaned
☕ Cafe Sales Dataset (Cleaned)
A cleaned and preprocessed version of a synthetic cafe sales transaction dataset. Originally a dirty, real-world-style dataset riddled with missing values, corrupt entries, and wrong data types — now fully cleaned and ready for analysis or ML tasks.
📋 Dataset Overview
Property
Value
Original Rows
10,000
Cleaned Rows
9,540
Columns
8
Time Period
January 2023 – December 2023
Domain
Retail / Food & Beverage… See the full description on the dataset page: https://huggingface.co/datasets/adityasuyal/cafe-sales-cleaned.videogames_salesvideo-games-salesabhishekgupta56447_video-games-sales-from-zenodo
Video Games Sales (from Zenodo)
Video game sales data including platform, genre, publisher, and global sales.
Dataset Info
Source: Kaggle
Original Size: 0.37 MB
Kaggle Downloads: 819
Files: 1
Files
vgsales.csv
Mirrored from Kaggle
sales-orders-uk
UK Sales Orders & Invoices Dataset
Complete sales pipeline for a simulated UK retail SME: customers, orders, order lines, invoices, and payment allocations. Customer behaviour is stochastic with seasonal patterns — not uniform random. Payment timing varies by customer reliability.
Tables
Table
Rows
companies
1
customers
113
payment_allocations
1,594
sales_invoice_lines
7,681
sales_invoices
1,712
sales_order_lines
7,681
sales_orders
1,712
Total
20… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/sales-orders-uk.uae-sales-table-qa
🇦🇪 UAE Sales Table QA (Arabic)
⚠️ Note: All data in this dataset is synthetically generated using random values for learning and experimentation purposes. It does not represent real-world business data.
🧠 Overview
UAE Sales Table QA (Arabic) is an Arabic Question–Answering dataset for table reasoning and data analysis, generated from 21 UAE-style CSV tables.Each example includes:
Question — a natural-language query about the data
Steps — human-readable reasoning… See the full description on the dataset page: https://huggingface.co/datasets/zSynctic/uae-sales-table-qa.kapibala-sales-dialogues
Kapibala Sales Dialogues
A sales-conversation dataset with outcome, conversation-level and sentence-level labels
🤗 Hugging Face · Annotation details · 中文
630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset:
L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/kapibala-ai/kapibala-sales-dialogues.lalm-judge-validation-full-duplex
LALM Judge Validation on Full-Duplex Voice Agents
Companion dataset for the paper A Reliability Assessment of
LALM Audio Judges for Full-Duplex Voice Agents.
This repository contains the anonymised ratings, adversarial-defect
recall tables, JSON schemas, and analysis scripts used to produce
every headline number, table, and figure in that paper.
Summary
209 rated stereo sessions: 152 full-duplex agent-client
conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.Salesforce__LLaMA-3-8B-SFR-Iterative-DPO-R-details
Dataset Card for Evaluation run of Salesforce/LLaMA-3-8B-SFR-Iterative-DPO-R
Dataset automatically created during the evaluation run of model Salesforce/LLaMA-3-8B-SFR-Iterative-DPO-R
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Salesforce__LLaMA-3-8B-SFR-Iterative-DPO-R-details.coffee_sales_datacar_sales_dataagungpambudi_trends-product-coffee-shop-sales-revenue-dataset
Maven Roasters: Coffee Shop Sales & Revenue Data
Unveiling Trends: Time Analysis, Transaction & Revenue in Coffee Shop Sales Data
Dataset Info
Source: Kaggle
Original Size: 2.54 MB
Kaggle Downloads: 4,364
Files: 2
Files
coffee-shop-sales-revenue.csv
coffee-shop-sales-revenue.parquet
Mirrored from Kaggle
Walmart-sales
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/large-traversaal/Walmart-sales.superkart-sales-datasetsales-forecasting-datashared-imagination
Dataset Card for Shared Imagination
This dataset contains the problems used in the paper Shared
Dataset Description
This dataset contains the questions generated for the investigations described in the TMLR paper Shared Imagination: LLMs Hallucinate Alike.
If you want to use this dataset to assess new models, please use the default config (i.e., datasets.load_dataset('Salesforce/shared-imagination')).
This config contains questions for which the four candidate choices… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/shared-imagination.ahmedmohamed2003_retail-store-sales-dirty-for-data-cleaning
Retail Store Sales: Dirty for Data Cleaning
Dirty Retail Store Sales Dataset
Dataset Info
Source: Kaggle
Original Size: 0.22 MB
Kaggle Downloads: 15,157
Files: 1
Files
retail_store_sales.csv
Mirrored from Kaggle
