CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Salesforce /xlam-function-calling-60kgated APIGen Function-Calling Datasets Paper | Website | Models This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness. We conducted human evaluation over 600 sampled data points… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k.textquestion-answering10K<n<100K719 likes37k downloads2y agoHugging Face02Salesforce /cos_e Dataset Card for "cos_e" Dataset Summary Common Sense Explanations (CoS-E) allows for training language models to automatically generate explanations that can be used during training and inference in a novel Commonsense Auto-Generated Explanation (CAGE) framework. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances v1.0 Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/cos_e.textquestion-answering10K<n<100K13 likes6.6k downloads3y agoHugging Face03Salesforce /APIGen-MT-5k Summary APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay Code: https://github.com/apigen-mt/apigen-mt.github.io The repo contains 5000 multi-turn trajectories collected by APIGen-MT This dataset is a subset of the data used to train the xLAM-2 model… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/APIGen-MT-5k.textquestion-answering1K<n<10K115 likes4.7k downloads1y agoHugging Face04Salesforce /UniDoc-Bench UNIDOC-BENCH Dataset A unified benchmark for document-centric multimodal retrieval-augmented generation (MM-RAG). Dataset Description UNIDOC-BENCH is the first large-scale, realistic benchmark for multimodal retrieval-augmented generation (MM-RAG) and Visual Question Answering (VQA) built from 70,000 real-world PDF pages across eight domains. The dataset extracts and links evidence from text, tables, and figures, then generates 1,700+ multimodal QA pairs spanning… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/UniDoc-Bench.imagequestion-answering1K<n<10K15 likes3.3k downloads10mo agoHugging Face05Salesforce /ConvoMem Conversational Memory Benchmark A comprehensive benchmark for evaluating conversational memory in large language models, featuring 75,336 question-answer pairs across six evidence categories. This benchmark addresses the critical challenge of memory management in conversational AI systems, where models must retain, update, and utilize information across extended multi-turn dialogues. 📚 Resources Paper: ConvoMem Benchmark: Why Your First 150 Conversations Don't Need RAG… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/ConvoMem.question-answering10K<n<100K4 likes2.7k downloads10mo agoHugging Face06Salesforce /FaithEval-unanswerable-v1.0 FaithEval FaithEval is a new and comprehensive benchmark dedicated to evaluating contextual faithfulness in LLMs across three diverse tasks: unanswerable, inconsistent, and counterfactual contexts. [Paper] FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows", ICLR 2025, https://arxiv.org/abs/2410.03727 [Code and Detailed Instructions] https://github.com/SalesforceAIResearch/FaithEval Disclaimer and Ethical Considerations… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FaithEval-unanswerable-v1.0.textquestion-answering1K<n<10K5 likes751 downloads2y agoHugging Face07Salesforce /ProVision-10M ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models ProVision is an extendable data generation engine which produces instruction data for large multimodal language models (MLMs). In particular, it synthesizes instruction data via data generators (Python programs) and scene graphs rather than proprietary models. It also includes a scene graph generation pipeline consisting of various state-of-the-art models (eg, object detection model). Thus… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/ProVision-10M.textquestion-answering10M<n<100M19 likes407 downloads2y agoHugging Face08Salesforce /Hard2Verify Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math Large language model (LLM)-based reasoning systems have recently achieved gold medal-level performance in the IMO 2025 competition, writing mathematical proofs where, to receive full credit, each step must be not only correct but also sufficiently supported. To train LLM-based reasoners in such challenging, open-ended settings, strong verifiers capable of catching step-level mistakes are necessary… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/Hard2Verify.textquestion-answeringn<1K7 likes272 downloads11mo agoHugging Face09Salesforce /dialogstudiogated DialogStudio: Unified Dialog Datasets and Instruction-Aware Models for Conversational AI Author: Jianguo Zhang, Kun Qian Paper|Github|[GDrive] 🎉 March 18, 2024: Update for AI Agent. Check xLAM for the latest data and models relevant to AI Agent! 🎉 March 10 2024: Update for dataset viewer issues: Please refer to https://github.com/salesforce/DialogStudio for view of each dataset, where we provide 5 converted examples along with 5 original examples under each data folder. For… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/dialogstudio.question-answering228 likes258 downloads2y agoHugging Face10Salesforce /HERB Dataset Card for HERB Dataset Description HERB is a benchmark for evaluating LLM agents’ ability to perform Deep Search and Long Context Reasoning. It is generated using a synthetic data pipeline that simulates business workflows across product planning, development, and support stages, generating interconnected content with realistic noise and multi-hop questions with guaranteed ground-truth answers. Directory Structure data/ ├── metadata/ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/HERB.text-retrieval10K<n<100K1 likes242 downloads1y agoHugging Face115CD-AI /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K8 likes107 downloads2y agoHugging Face12Salesforce /EDR-200 Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics Paper: Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics Code: https://github.com/SalesforceAIResearch/enterprise-deep-research Dataset Overview EDR-200 contains 201 complete agentic research trajectories generated by Enterprise Deep Research—99 queries from DeepResearch Bench and 102 queries from DeepConsult. Unlike prior benchmarks that only… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/EDR-200.textquestion-answeringn<1K15 likes104 downloads11mo agoHugging Face13Salesforce /shared-imagination Dataset Card for Shared Imagination This dataset contains the problems used in the paper Shared Dataset Description This dataset contains the questions generated for the investigations described in the TMLR paper Shared Imagination: LLMs Hallucinate Alike. If you want to use this dataset to assess new models, please use the default config (i.e., datasets.load_dataset('Salesforce/shared-imagination')). This config contains questions for which the four candidate choices… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/shared-imagination.tabularmultiple-choice10K<n<100K0 likes44 downloads1y agoHugging Face14ChaosAIVision /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K0 likes34 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.