CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Salesforce /wikitext Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/wikitext.texttext-generation1M<n<10M809 likes1.9m downloads3y agoHugging Face02Salesforce /xlam-function-calling-60kgated APIGen Function-Calling Datasets Paper | Website | Models This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness. We conducted human evaluation over 600 sampled data points… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k.textquestion-answering10K<n<100K719 likes37k downloads2y agoHugging Face03Salesforce /APIGen-MT-5k Summary APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay Code: https://github.com/apigen-mt/apigen-mt.github.io The repo contains 5000 multi-turn trajectories collected by APIGen-MT This dataset is a subset of the data used to train the xLAM-2 model… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/APIGen-MT-5k.textquestion-answering1K<n<10K115 likes4.7k downloads1y agoHugging Face04Salesforce /ConvoMem Conversational Memory Benchmark A comprehensive benchmark for evaluating conversational memory in large language models, featuring 75,336 question-answer pairs across six evidence categories. This benchmark addresses the critical challenge of memory management in conversational AI systems, where models must retain, update, and utilize information across extended multi-turn dialogues. 📚 Resources Paper: ConvoMem Benchmark: Why Your First 150 Conversations Don't Need RAG… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/ConvoMem.question-answering10K<n<100K4 likes2.7k downloads10mo agoHugging Face05Salesforce /dialogstudiogated DialogStudio: Unified Dialog Datasets and Instruction-Aware Models for Conversational AI Author: Jianguo Zhang, Kun Qian Paper|Github|[GDrive] 🎉 March 18, 2024: Update for AI Agent. Check xLAM for the latest data and models relevant to AI Agent! 🎉 March 10 2024: Update for dataset viewer issues: Please refer to https://github.com/salesforce/DialogStudio for view of each dataset, where we provide 5 converted examples along with 5 original examples under each data folder. For… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/dialogstudio.question-answering228 likes258 downloads2y agoHugging Face06Salesforce /MASBench 🎼 MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks This is the proposed MAS evaluation data used in the recipe described in our paper:📄 MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks For more details, please check the following resources: 🌐 Project Page: https://mas-orchestra.salesforceresearch.ai/mas_r1/index.html 📚 Live… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/MASBench.texttext-generation10K<n<100K3 likes235 downloads4mo agoHugging Face07Salesforce /RealUserSim RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation Behavioral user profiles and evaluation benchmark for realistic LLM-powered user simulation, derived from the WildChat dataset. Dataset Summary This release contains: 7,273 behavioral user profiles extracted from real conversations, each containing demographics and executable linguistic style commands 600 evaluation test cases (6 splits x 100) for measuring user simulation fidelity… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/RealUserSim.tabulartext-generation1K<n<10K1 likes162 downloads5mo agoHugging Face08Salesforce /ReasoningJudgeBench J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization Austin Xu, Yilun Zhou, Xuan-Phi Nguyen, Caiming Xiong, Shafiq Joty To run evaluation, please see our Github repo. 💻 Github: https://github.com/SalesforceAIResearch/ReasoningJudgeBench 📜 Paper: https://arxiv.org/abs/2505.13346 ReasoningJudgeBench ReasoningJudgeBench is a 1,483 sample pairwise benchmark introduced in the paper J4R: Learning to Judge with Equivalent Initial State… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/ReasoningJudgeBench.texttext-generation1K<n<10K5 likes131 downloads1y agoHugging Face09Salesforce /tracelab-comprehend tracelab COMPREHEND synthetic corpus Twelve seeded synthetic long-horizon agent sessions (JSONL event streams, ~11 MB total) released with the paper "Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers" (Pakhomov & Nijkamp, Salesforce AI Research; arXiv:2609.01466, https://arxiv.org/abs/2609.01466). 📄 Paper: https://arxiv.org/abs/2609.01466 💻 Code, generator, benchmarks, and traces (BSD-3-Clause): https://github.com/SalesforceAIResearch/tracelab… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/tracelab-comprehend.text-generationn<1K0 likes129 downloads23d agoHugging Face105CD-AI /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K8 likes107 downloads2y agoHugging Face11Salesforce /EDR-200 Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics Paper: Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics Code: https://github.com/SalesforceAIResearch/enterprise-deep-research Dataset Overview EDR-200 contains 201 complete agentic research trajectories generated by Enterprise Deep Research—99 queries from DeepResearch Bench and 102 queries from DeepConsult. Unlike prior benchmarks that only… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/EDR-200.textquestion-answeringn<1K15 likes104 downloads11mo agoHugging Face12Salesforce /vibepass VIBEPASS: Can Vibe Coders Really Pass the Vibe Check? Authors: Srijan Bansal, Jiao Fangkai, Yilun Zhou, Austin Xu, Shafiq Joty, Semih Yavuz TL;DR: As LLMs shift programming toward human-guided "vibe coding", agentic tools increasingly rely on models to self-diagnose and repair their own subtle faults—a capability central to autonomous software engineering yet never systematically evaluated. VIBEPASS presents the first empirical benchmark that decomposes fault-targeted reasoning into… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/vibepass.texttext-generationn<1K2 likes82 downloads6mo agoHugging Face13Salesforce /lalm-judge-validation-full-duplex LALM Judge Validation on Full-Duplex Voice Agents Companion dataset for the paper A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents. This repository contains the anonymised ratings, adversarial-defect recall tables, JSON schemas, and analysis scripts used to produce every headline number, table, and figure in that paper. Summary 209 rated stereo sessions: 152 full-duplex agent-client conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.tabularaudio-classification1K<n<10K2 likes57 downloads2mo agoHugging Face14ChaosAIVision /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K0 likes34 downloads9mo agoHugging Face15DealScopeAI /dealscope-salesforce-ai-brief-dataset-v1 DealScope Salesforce AI Brief Dataset v1 Dataset Summary This dataset contains 25 structured Salesforce-record brief examples in the DealScope output format. Each record is shaped like a real DealScope API response and includes: record metadata buying signals risks stakeholders a draft follow-up email a multi-line summary The dataset is intended as a public retrieval and reference asset for Salesforce-focused AI brief workflows. What Is In This Release 2… See the full description on the dataset page: https://huggingface.co/datasets/DealScopeAI/dealscope-salesforce-ai-brief-dataset-v1.texttext-generationn<1K0 likes10 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.