datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xlam-function-calling-60k
APIGen Function-Calling Datasets
Paper | Website | Models
This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness.
We conducted human evaluation over 600 sampled data points, and… See the full description on the dataset page: https://huggingface.co/datasets/lockon/xlam-function-calling-60k.xlam-function-calling-60k
APIGen Function-Calling Datasets
Paper | Website | Models
This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness.
We conducted human evaluation over 600 sampled data points… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k.DS-1000 DS-1000 in simplified format
🔥 Check the leaderboard from Eval-Arena on our project page.
See testing code and more information (also the original fill-in-the-middle/Insertion format) in the DS-1000 repo.
Reformatting credits: Yuhang Lai, Sida Wang
xlam-function-calling-60kosworld_v2_tasks
OSWorld V2 Task Classes
This gated dataset contains the official root-level task_*.py Python task classes for OSWorld V2.
The public GitHub repository keeps the task loader, helper utilities, and documentation. The task implementations are gated to reduce benchmark leakage and to help prevent evaluated agents from finding task answers, setup logic, or evaluator details online while executing a task.
Download from the public repository root with:
uvx --from huggingface_hub hf… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/osworld_v2_tasks.xlam-function-calling-60k-shareGPTShareGPT converted version of Salesforce/xlam-function-calling-60k
xlam-irrelevance-7.5k
xlam-irrelevance-7.5k
Overview
The xlam-irrelevance-7.5k is a specialized dataset designed to activate the ability of irrelevant function detection for large language models (LLMs).
Source and Construction
This dataset is built upon xlam-function-calling-60k dataset, from which we random sampled 7.5k instances, removed the ground truth function from the provided tool list, and relabel them as irrelevant. For more details, please refer to Hammer: Robust… See the full description on the dataset page: https://huggingface.co/datasets/MadeAgents/xlam-irrelevance-7.5k.AgentTrek
AgentTrek Data Collection
AgentTrek dataset is the training dataset for the Web agent AgentTrek-1.0-32B. It consists of a total of 52,594 dialogue turns, specifically designed to train a language model for performing web-based tasks, such as browsing and web shopping. The dialogues in this dataset simulate interactions where the agent assists users in tasks like searching for information, comparing products, making purchasing decisions, and navigating websites.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/AgentTrek.xlam-function-calling-60k_langchainReformatted dataset from "Salesforce/xlam-function-calling-60k" (from Hugging Face) for the purposes of fine tuning LLMs for tool calling for the LangChain and LangGraph frameworks
license: mit
spider2-litesommelier-xlam-single-call-splits
sommelier xlam single-call splits
Deterministic, deduplicated, single-tool-call train/validation/test
splits derived from
Salesforce/xlam-function-calling-60k
(APIGen, CC-BY-4.0), produced by the
sommelier pipeline for
reproducible tool-calling fine-tuning. These are the exact splits used to
train and evaluate
abdelstark/llama-3.1-nemotron-nano-8b-xlam-tool-calling-lora.
Why single-call
The upstream dataset mixes single-call and multi-call examples (~52.6%… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits.XLAM-AtroposA hermes function call formatted version of the XLAM function calling dataset for use in Atropos - Nous' LLM RL Environments Framework
Vietnamese-Salesforce-xlam-function-calling-60k-gg-translateddnhkng__RYS-XLarge-details
Dataset Card for Evaluation run of dnhkng/RYS-XLarge
Dataset automatically created during the evaluation run of model dnhkng/RYS-XLarge
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/dnhkng__RYS-XLarge-details.xlam-function-callsXL-AlpacaEval
Dataset Card for XL-AlpacaEval
XL-AlpacaEval is a benchmark for evaluating the cross-lingual open-ended generation capabilities of Large Language Models (LLMs), introduced in the paper XL-Instruct: Synthetic Data for Cross-Lingual Open-Ended Generation. It is designed to evaluate a model's ability to respond in a target language that is different from the source language of the user's query.
For evaluating multilingual (i.e., non-English, but monolingual) generation, see the sister… See the full description on the dataset page: https://huggingface.co/datasets/viyer98/XL-AlpacaEval.Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedxlam-function-calling-60k_langchain_langgraphXLA-CACHE-NEW-VERSION-CORE-ONLYxlam-function-calling-60k-ita
ReDiX Function Calling ITA
This dataset is the italian translation of Salesforce/xlam-function-calling-60k
sommelier-xlam-single-call-splits-fr
sommelier-xlam-single-call-splits-fr
French paired variant of the single call tool calling rows selected by the Sommelier reference pipeline from Salesforce/xlam-function-calling-60k. Only the user query is translated. Tool schemas and gold answers are byte identical to the English source rows, so the two languages measure the same task with the same scoring.
How it was built
The Sommelier data translate tool (source) translated the exact 17,000 rows the reference… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits-fr.sommelier-xlam-single-call-splits-he-hymt-sanitized
Sommelier xLAM single-call Hebrew paired rows (Hy-MT2, sanitized release)
This CC-BY-4.0 dataset is derived from
Salesforce/xlam-function-calling-60k.
Sommelier filters the source corpus to single-tool-call examples, deterministically
splits it, and machine-translates only each natural-language query into Hebrew.
The exact training snapshot kept tool schemas and gold answers byte-identical to
the English root. For public release, 15 GitHub-PAT-shaped substrings inherited
from… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits-he-hymt-sanitized.xlam2_sub_samples_with_rubrics_v1
xlam2_sub_samples_with_rubrics_v1 (messages + tools)
Private dataset for tool-use SFT derived from xlam2_sub_samples_with_rubrics_v1.json.
Samples: 2,044
Domains: e-commerce customer support and airline booking
Format: one JSON object per line with the following keys:
messages: OpenAI-style chat with roles system, user, assistant, tool.
Tool calls are represented as assistant messages with tool_calls: [{id, type: "function", function: {name, arguments}}].
Tool responses are tool… See the full description on the dataset page: https://huggingface.co/datasets/yentinglin/xlam2_sub_samples_with_rubrics_v1.Foxtool-XLAM-Atroposxlam_function_calldnhkng__RYS-XLarge-base-details
Dataset Card for Evaluation run of dnhkng/RYS-XLarge-base
Dataset automatically created during the evaluation run of model dnhkng/RYS-XLarge-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/dnhkng__RYS-XLarge-base-details.dnhkng__RYS-XLarge2-details
Dataset Card for Evaluation run of dnhkng/RYS-XLarge2
Dataset automatically created during the evaluation run of model dnhkng/RYS-XLarge2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/dnhkng__RYS-XLarge2-details.mirror-XLAM-7.5k-Irrelevance
xlam-irrelevance-7.5k
Overview
The xlam-irrelevance-7.5k is a specialized dataset designed to activate the ability of irrelevant function detection for large language models (LLMs).
Source and Construction
This dataset is built upon xlam-function-calling-60k dataset, from which we random sampled 7.5k instances, removed the ground truth function from the provided tool list, and relabel them as irrelevant. For more details, please refer to Hammer:… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-XLAM-7.5k-Irrelevance.xlam-irrelevance-7.5k
xlam-irrelevance-7.5k
Overview
The xlam-irrelevance-7.5k is a specialized dataset designed to activate the ability of irrelevant function detection for large language models (LLMs).
Source and Construction
This dataset is built upon xlam-function-calling-60k dataset, from which we random sampled 7.5k instances, removed the ground truth function from the provided tool list, and relabel them as irrelevant. For more details, please refer to Hammer: Robust… See the full description on the dataset page: https://huggingface.co/datasets/LyounJAP/xlam-irrelevance-7.5k.
