datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xlam-function-calling-60k
APIGen Function-Calling Datasets
Paper | Website | Models
This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness.
We conducted human evaluation over 600 sampled data points, and… See the full description on the dataset page: https://huggingface.co/datasets/lockon/xlam-function-calling-60k.xlam-function-calling-60k
APIGen Function-Calling Datasets
Paper | Website | Models
This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness.
We conducted human evaluation over 600 sampled data points… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k.xlam-function-calling-60k-parsed
[PARSED] APIGen Function-Calling Datasets (xLAM)
This dataset contains the full data from the original Salesforce/xlam-function-calling-60k
Subset name
multi-turn
parallel
multiple definition
Last turn type
number of dataset
xlam-function-calling-60k
no
yes
yes
tool_calls
60000
This is a re-parsing formatting dataset for the xLAM official dataset.
Load the dataset
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/xlam-function-calling-60k-parsed.CUA-Gym
CUA-Gym
CUA-Gym is a collection of verifiable computer-use agent tasks for reinforcement learning with verifiable rewards (RLVR). Each task pairs a natural-language instruction with executable setup artifacts and a Python reward function that checks task completion programmatically. For details, see the paper CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents.
This release contains the full public CUA-Gym task set after the necessary data review.… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/CUA-Gym.OpenGrad-ToolPolicy-Canonical-v2-minus-xlam
This is the byte-identical training view for a joint xLAM-plus-CALL_PREDICTION removal
experiment. xLAM is currently the corpus's only source of that supervision contract, so this is
not a pure source-content ablation. It carries no result of its own and is not a recommended
mixture. It is part of OpenGrad Study 001.
What this is
OpenGrad-ToolPolicy-Canonical-v2
with one source removed: xLAM/APIGen. Three sources remain, 115,895 canonical records, 118
shards.
It is the exact… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v2-minus-xlam.CoT-XLangRU:CoT-XLang — это многоязычный датасет, состоящий из текстовых примеров с пошаговыми рассуждениями (Chain-of-Thought, CoT) на различных языках, включая английский, русский, японский и другие. Он используется для обучения и тестирования моделей в задачах, требующих пояснений решений через несколько шагов. Датасет включает около 2,419,912 примеров, что позволяет эффективно обучать модели, способные генерировать пошаговые рассуждения.
Рекомендация:Используйте датасет для обучения моделей… See the full description on the dataset page: https://huggingface.co/datasets/Egor-3926/CoT-XLang.sommelier-xlam-single-call-splits
sommelier xlam single-call splits
Deterministic, deduplicated, single-tool-call train/validation/test
splits derived from
Salesforce/xlam-function-calling-60k
(APIGen, CC-BY-4.0), produced by the
sommelier pipeline for
reproducible tool-calling fine-tuning. These are the exact splits used to
train and evaluate
abdelstark/llama-3.1-nemotron-nano-8b-xlam-tool-calling-lora.
Why single-call
The upstream dataset mixes single-call and multi-call examples (~52.6%… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits.Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedxlam-interleave-thinking-40k
XLAM Interleave Thinking 40k
Dataset Description
This dataset is a variation of Salesforce/xlam-function-calling-60k, distilled from Minimax m2.1 with a pipeline to rewrite queries into more complex multi-turn tool calling. It specifically features interleaved thinking traces (Chain of Thought) within the assistant responses to enhance reasoning capabilities in function-calling scenarios.
Dataset Summary
Original Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/hxssgaa/xlam-interleave-thinking-40k.XL-AlpacaEval
Dataset Card for XL-AlpacaEval
XL-AlpacaEval is a benchmark for evaluating the cross-lingual open-ended generation capabilities of Large Language Models (LLMs), introduced in the paper XL-Instruct: Synthetic Data for Cross-Lingual Open-Ended Generation. It is designed to evaluate a model's ability to respond in a target language that is different from the source language of the user's query.
For evaluating multilingual (i.e., non-English, but monolingual) generation, see the sister… See the full description on the dataset page: https://huggingface.co/datasets/viyer98/XL-AlpacaEval.Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedxlam-function-calling-60k-ita
ReDiX Function Calling ITA
This dataset is the italian translation of Salesforce/xlam-function-calling-60k
sommelier-xlam-single-call-splits-fr
sommelier-xlam-single-call-splits-fr
French paired variant of the single call tool calling rows selected by the Sommelier reference pipeline from Salesforce/xlam-function-calling-60k. Only the user query is translated. Tool schemas and gold answers are byte identical to the English source rows, so the two languages measure the same task with the same scoring.
How it was built
The Sommelier data translate tool (source) translated the exact 17,000 rows the reference… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits-fr.sommelier-xlam-single-call-splits-he-hymt-sanitized
Sommelier xLAM single-call Hebrew paired rows (Hy-MT2, sanitized release)
This CC-BY-4.0 dataset is derived from
Salesforce/xlam-function-calling-60k.
Sommelier filters the source corpus to single-tool-call examples, deterministically
splits it, and machine-translates only each natural-language query into Hebrew.
The exact training snapshot kept tool schemas and gold answers byte-identical to
the English root. For public release, 15 GitHub-PAT-shaped substrings inherited
from… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits-he-hymt-sanitized.
