datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xlam-function-calling-60k
APIGen Function-Calling Datasets
Paper | Website | Models
This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness.
We conducted human evaluation over 600 sampled data points… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k.APIGen-MT-5k
Summary
APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications
This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
Code: https://github.com/apigen-mt/apigen-mt.github.io
The repo contains 5000 multi-turn trajectories collected by APIGen-MT
This dataset is a subset of the data used to train the xLAM-2 model… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/APIGen-MT-5k.Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedvibepass
VIBEPASS: Can Vibe Coders Really Pass the Vibe Check?
Authors: Srijan Bansal, Jiao Fangkai, Yilun Zhou, Austin Xu, Shafiq Joty, Semih Yavuz
TL;DR: As LLMs shift programming toward human-guided "vibe coding", agentic tools increasingly rely on models to self-diagnose and repair their own subtle faults—a capability central to autonomous software engineering yet never systematically evaluated. VIBEPASS presents the first empirical benchmark that decomposes fault-targeted reasoning into… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/vibepass.Vietnamese-Salesforce-xlam-function-calling-60k-gg-translateddealscope-salesforce-ai-brief-dataset-v1
DealScope Salesforce AI Brief Dataset v1
Dataset Summary
This dataset contains 25 structured Salesforce-record brief examples in the DealScope output format.
Each record is shaped like a real DealScope API response and includes:
record metadata
buying signals
risks
stakeholders
a draft follow-up email
a multi-line summary
The dataset is intended as a public retrieval and reference asset for Salesforce-focused AI brief workflows.
What Is In This Release
2… See the full description on the dataset page: https://huggingface.co/datasets/DealScopeAI/dealscope-salesforce-ai-brief-dataset-v1.
