datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikitext
Dataset Card for "wikitext"
Dataset Summary
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified
Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License.
Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over
110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/wikitext.xlam-function-calling-60k
APIGen Function-Calling Datasets
Paper | Website | Models
This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness.
We conducted human evaluation over 600 sampled data points… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k.APIGen-MT-5k
Summary
APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications
This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
Code: https://github.com/apigen-mt/apigen-mt.github.io
The repo contains 5000 multi-turn trajectories collected by APIGen-MT
This dataset is a subset of the data used to train the xLAM-2 model… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/APIGen-MT-5k.salesbench-100
SalesBench-100
SalesBench-100 is a synthetic long-horizon sales-agent benchmark with 100 original workflows across Salesforce, HubSpot, Gong, and a seeded evidence room. Each task begins with a high-level employee request and has its own authored causal rule and provider transition. Identity, operating facts, authority, governed policy, live-system indexes, and exceptions are separated so no mounted business asset publishes a selected option or precomputed change. Every task has… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/salesbench-100.saas-sales-conversations
saas-sales-conversations
Dataset Description
This is a synthetic dataset of sales conversations for SaaS (Software as a Service) companies, designed for training sales conversion prediction models. The dataset was created following the methodology presented in "SalesRLAgent: A Reinforcement Learning Approach for Real-Time Sales Conversion Prediction and Optimization" (Nandakishor M, 2025).
The dataset contains realistic dialogues between sales representatives and… See the full description on the dataset page: https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations.iraqi-arabic-sales-dialogue-dataset
Iraqi Arabic Sales Dialogue Dataset
A large synthetic dataset of Iraqi (Baghdadi-based) Arabic dialogue, centered on
retail sales, haggling, and everyday conversation.
النسخة العربية متوفرة بالكامل بالأسفل — Arabic version available in full below.
What this is
210,832 template-generated conversations, of which 171,601 (81%) are exact-unique
message sequences, spanning 20 topical categories in colloquial Iraqi Arabic. The
core of the dataset (10 categories) is… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi-arabic-sales-dialogue-dataset.MASBench
🎼 MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
This is the proposed MAS evaluation data used in the recipe described in our paper:📄 MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
For more details, please check the following resources:
🌐 Project Page: https://mas-orchestra.salesforceresearch.ai/mas_r1/index.html
📚 Live… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/MASBench.kapibala-sales-dialogues
Kapibala Sales Dialogues
A sales-conversation dataset with outcome, conversation-level and sentence-level labels
🤗 Hugging Face · Annotation details · 中文
630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset:
L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/Thomasgudan/kapibala-sales-dialogues.Gym_Salesman_Dataset
Gym Salesman Dataset
11,997 synthetic gym-membership sales conversations, each labelled SUCCESS or FAILURE.
🔗 Project Links
| Live App — practice against an AI customer | Hugging Face Space |
| Telegram Bot — practice on the go | @ido_salescoach_bot |
| Dataset — 11,997 labelled conversations | elg4/Gym_Salesman_Dataset |
| Data Generation — how the data was built | notebook |
| Recommendation — the embedding retriever | notebook |
Every conversation is a… See the full description on the dataset page: https://huggingface.co/datasets/elg4/Gym_Salesman_Dataset.RealUserSim
RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
Behavioral user profiles and evaluation benchmark for realistic LLM-powered user simulation, derived from the WildChat dataset.
Dataset Summary
This release contains:
7,273 behavioral user profiles extracted from real conversations, each containing demographics and executable linguistic style commands
600 evaluation test cases (6 splits x 100) for measuring user simulation fidelity… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/RealUserSim.ReasoningJudgeBench
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
Austin Xu, Yilun Zhou, Xuan-Phi Nguyen, Caiming Xiong, Shafiq Joty
To run evaluation, please see our Github repo.
💻 Github: https://github.com/SalesforceAIResearch/ReasoningJudgeBench
📜 Paper: https://arxiv.org/abs/2505.13346
ReasoningJudgeBench
ReasoningJudgeBench is a 1,483 sample pairwise benchmark introduced in the paper J4R: Learning to Judge with Equivalent Initial State… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/ReasoningJudgeBench.sales-textbook_for_convincing_and_selling
Dataset Card for sales-textbook_for_convincing_and_selling
A textbook create for the purpose of training a sales chatbot.
Inspiration come from: Textbooks is all you need https://arxiv.org/abs/2306.11644
The data was generated by gpt-3.5-turbo
#Structure
A simpel textbook that has subheadlines and headlines.
Chapters and Subheadlines are mentioned in the dataset. Look at the first two examples.
Data Generation
The following code was used for the text generation:… See the full description on the dataset page: https://huggingface.co/datasets/goendalf666/sales-textbook_for_convincing_and_selling.Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedEDR-200
Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
Paper: Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
Code: https://github.com/SalesforceAIResearch/enterprise-deep-research
Dataset Overview
EDR-200 contains 201 complete agentic research trajectories generated by Enterprise Deep Research—99 queries from DeepResearch Bench and 102 queries from DeepConsult. Unlike prior benchmarks that only… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/EDR-200.vibepass
VIBEPASS: Can Vibe Coders Really Pass the Vibe Check?
Authors: Srijan Bansal, Jiao Fangkai, Yilun Zhou, Austin Xu, Shafiq Joty, Semih Yavuz
TL;DR: As LLMs shift programming toward human-guided "vibe coding", agentic tools increasingly rely on models to self-diagnose and repair their own subtle faults—a capability central to autonomous software engineering yet never systematically evaluated. VIBEPASS presents the first empirical benchmark that decomposes fault-targeted reasoning into… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/vibepass.kapibala-sales-dialogues
Kapibala Sales Dialogues
A sales-conversation dataset with outcome, conversation-level and sentence-level labels
🤗 Hugging Face · Annotation details · 中文
630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset:
L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/kapibala-ai/kapibala-sales-dialogues.real_estate_sales
房地产销冠话术 - 多轮对话
SalesAgent-Consultant-V_0.0.1
SalesAgent-Consultant Dataset V0.0.1
🏢 Brought to you by Dauji AI
AI models for Sales, CRM and Consultancy
🤖 SALES AGENT CONSULTANT DATASET - VERSION 0.0.1 🤖
🎯 Overview
The SalesAgent-Consultant Dataset V0.0.1 is a comprehensive sales training dataset containing 124,954 high-quality sales conversations with detailed metadata across the complete sales cycle. This dataset is specifically designed for training AI sales agents and consultants with deep… See the full description on the dataset page: https://huggingface.co/datasets/Dauji-AI/SalesAgent-Consultant-V_0.0.1.lalm-judge-validation-full-duplex
LALM Judge Validation on Full-Duplex Voice Agents
Companion dataset for the paper A Reliability Assessment of
LALM Audio Judges for Full-Duplex Voice Agents.
This repository contains the anonymised ratings, adversarial-defect
recall tables, JSON schemas, and analysis scripts used to produce
every headline number, table, and figure in that paper.
Summary
209 rated stereo sessions: 152 full-duplex agent-client
conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.sales-methodology-sft-100k
Sales Methodology SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level B2B sales execution across discovery, objection handling, enterprise pricing, prospecting, account management, and sales leadership.
Motivation
Enterprise sales is one of the highest-leverage skills in business — great salespeople and sales leaders drive disproportionate revenue. Models commonly fail at sales tasks by:
Generic frameworks without execution detail: Describing… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/sales-methodology-sft-100k.Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedsales-support-108-perfect
sales-support-108-perfect
Описание
Датасет для обучения AI моделей продаж и поддержки клиентов.
Включает 11 специализированных сфер × 108 примеров:
ПРОДАЖИ:
019: Sales Manager - выявление потребностей и закрытие сделок
020: Account Manager - долгосрочные отношения и рост аккаунта
021: PreSales Engineer - техническая экспертиза на этапе продажи
022: Business Developer - поиск новых возможностей
ПОДДЕРЖКА:
023: Customer Support - быстрое решение проблем
024: Technical… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/sales-support-108-perfect.dealscope-salesforce-ai-brief-dataset-v1
DealScope Salesforce AI Brief Dataset v1
Dataset Summary
This dataset contains 25 structured Salesforce-record brief examples in the DealScope output format.
Each record is shaped like a real DealScope API response and includes:
record metadata
buying signals
risks
stakeholders
a draft follow-up email
a multi-line summary
The dataset is intended as a public retrieval and reference asset for Salesforce-focused AI brief workflows.
What Is In This Release
2… See the full description on the dataset page: https://huggingface.co/datasets/DealScopeAI/dealscope-salesforce-ai-brief-dataset-v1.SalesA_SMEs_Datasetssales_analysis_queriesB2B-Sales-Acceleration-Intelligence
B2B Sales & CRM Intelligence Dataset (Expert Edition)
This repository contains a premium, expert-verified dataset of 500+ instruction-response pairs designed to fine-tune AI agents for B2B sales acceleration.
💰 Access & Licensing
Access to this dataset is strictly gated for commercial and professional use.
To gain access:
Click the "Apply for Commercial Access" button above and provide your details.
Purchase the Commercial License here:
Once the transaction is… See the full description on the dataset page: https://huggingface.co/datasets/Hridhi/B2B-Sales-Acceleration-Intelligence.
