CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lockon /xlam-function-calling-60k APIGen Function-Calling Datasets Paper | Website | Models This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness. We conducted human evaluation over 600 sampled data points, and… See the full description on the dataset page: https://huggingface.co/datasets/lockon/xlam-function-calling-60k.textquestion-answering10K<n<100K1 likes39k downloads2y agoHugging Face02Salesforce /xlam-function-calling-60kgated APIGen Function-Calling Datasets Paper | Website | Models This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness. We conducted human evaluation over 600 sampled data points… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k.textquestion-answering10K<n<100K720 likes37k downloads2y agoHugging Face03minpeter /xlam-function-calling-60k-parsed [PARSED] APIGen Function-Calling Datasets (xLAM) This dataset contains the full data from the original Salesforce/xlam-function-calling-60k Subset name multi-turn parallel multiple definition Last turn type number of dataset xlam-function-calling-60k no yes yes tool_calls 60000 This is a re-parsing formatting dataset for the xLAM official dataset. Load the dataset from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/xlam-function-calling-60k-parsed.texttext-generation10K<n<100K3 likes20k downloads1y agoHugging Face04xlangai /CUA-Gym CUA-Gym CUA-Gym is a collection of verifiable computer-use agent tasks for reinforcement learning with verifiable rewards (RLVR). Each task pairs a natural-language instruction with executable setup artifacts and a Python reward function that checks task completion programmatically. For details, see the paper CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents. This release contains the full public CUA-Gym task set after the necessary data review.… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/CUA-Gym.tabularreinforcement-learning10K<n<100K29 likes2.6k downloads4mo agoHugging Face05arjhinety /OpenGrad-ToolPolicy-Canonical-v2-minus-xlam This is the byte-identical training view for a joint xLAM-plus-CALL_PREDICTION removal experiment. xLAM is currently the corpus's only source of that supervision contract, so this is not a pure source-content ablation. It carries no result of its own and is not a recommended mixture. It is part of OpenGrad Study 001. What this is OpenGrad-ToolPolicy-Canonical-v2 with one source removed: xLAM/APIGen. Three sources remain, 115,895 canonical records, 118 shards. It is the exact… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/OpenGrad-ToolPolicy-Canonical-v2-minus-xlam.texttext-generation100K<n<1M0 likes297 downloads13d agoHugging Face06Egor-3926 /CoT-XLangRU:CoT-XLang — это многоязычный датасет, состоящий из текстовых примеров с пошаговыми рассуждениями (Chain-of-Thought, CoT) на различных языках, включая английский, русский, японский и другие. Он используется для обучения и тестирования моделей в задачах, требующих пояснений решений через несколько шагов. Датасет включает около 2,419,912 примеров, что позволяет эффективно обучать модели, способные генерировать пошаговые рассуждения. Рекомендация:Используйте датасет для обучения моделей… See the full description on the dataset page: https://huggingface.co/datasets/Egor-3926/CoT-XLang.texttext-generation1M<n<10M7 likes284 downloads2y agoHugging Face07abdelstark /sommelier-xlam-single-call-splits sommelier xlam single-call splits Deterministic, deduplicated, single-tool-call train/validation/test splits derived from Salesforce/xlam-function-calling-60k (APIGen, CC-BY-4.0), produced by the sommelier pipeline for reproducible tool-calling fine-tuning. These are the exact splits used to train and evaluate abdelstark/llama-3.1-nemotron-nano-8b-xlam-tool-calling-lora. Why single-call The upstream dataset mixes single-call and multi-call examples (~52.6%… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits.texttext-generation10K<n<100K0 likes157 downloads3mo agoHugging Face085CD-AI /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K8 likes108 downloads2y agoHugging Face09viyer98 /XL-AlpacaEval Dataset Card for XL-AlpacaEval XL-AlpacaEval is a benchmark for evaluating the cross-lingual open-ended generation capabilities of Large Language Models (LLMs), introduced in the paper XL-Instruct: Synthetic Data for Cross-Lingual Open-Ended Generation. It is designed to evaluate a model's ability to respond in a target language that is different from the source language of the user's query. For evaluating multilingual (i.e., non-English, but monolingual) generation, see the sister… See the full description on the dataset page: https://huggingface.co/datasets/viyer98/XL-AlpacaEval.texttext-generation1K<n<10K1 likes37 downloads1y agoHugging Face10ChaosAIVision /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K0 likes34 downloads9mo agoHugging Face11ReDiX /xlam-function-calling-60k-itagated ReDiX Function Calling ITA This dataset is the italian translation of Salesforce/xlam-function-calling-60k texttext-generation10K<n<100K0 likes28 downloads2y agoHugging Face12abdelstark /sommelier-xlam-single-call-splits-fr sommelier-xlam-single-call-splits-fr French paired variant of the single call tool calling rows selected by the Sommelier reference pipeline from Salesforce/xlam-function-calling-60k. Only the user query is translated. Tool schemas and gold answers are byte identical to the English source rows, so the two languages measure the same task with the same scoring. How it was built The Sommelier data translate tool (source) translated the exact 17,000 rows the reference… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits-fr.texttext-generation10K<n<100K0 likes26 downloads3mo agoHugging Face13abdelstark /sommelier-xlam-single-call-splits-he-hymt-sanitized Sommelier xLAM single-call Hebrew paired rows (Hy-MT2, sanitized release) This CC-BY-4.0 dataset is derived from Salesforce/xlam-function-calling-60k. Sommelier filters the source corpus to single-tool-call examples, deterministically splits it, and machine-translates only each natural-language query into Hebrew. The exact training snapshot kept tool schemas and gold answers byte-identical to the English root. For public release, 15 GitHub-PAT-shaped substrings inherited from… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits-he-hymt-sanitized.texttext-generation10K<n<100K1 likes20 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.