CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AngieYYF /SPADE-customer-service-dialogue SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection Paper | Code SPADE contains a repository of customer service line synthetic user dialogues with goals, augmented from MultiWOZ 2.1 using GPT-3.5 and Llama 70B. The datasets are intended for training and evaluating machine generated text detectors in dialogue settings. There are 15 English datasets generated using 5 different augmentation methods and 2 large language models… See the full description on the dataset page: https://huggingface.co/datasets/AngieYYF/SPADE-customer-service-dialogue.tabulartext-generation10K<n<100K3 likes187 downloads1y agoHugging Face02ConsumerDividends /Customer-Churn-Dataset-V2 Customer Churn Conversation Dataset - 500 (Benchmark-Anchored) Free 500-record sample. Licensed CC BY-NC 4.0. Commercial use requires a license. The generator is the product This sample was produced by our synthetic customer-churn dialogue generator. The generator is what we license: it produces a labeled 10,000-record dataset anchored to published subscription-industry benchmarks, with a cleaner and refiner pipeline built in. Real churn conversations are locked… See the full description on the dataset page: https://huggingface.co/datasets/ConsumerDividends/Customer-Churn-Dataset-V2.tabulartext-classificationn<1K0 likes51 downloads2mo agoHugging Face03dzcorpora /algerian-darija-customer-service-samplegated Algerian Darija customer messages — stratified sample 500 spontaneous Algerian Darija messages, written by real customers, drawn from a first-party corpus of 869,166 customer messages. Every message here is unique after normalization, de-identified, and typed by a human — nothing elicited, translated, scraped or generated. Algerian Darija (ISO 639-3 arq) is spoken by around 45 million people and is one of the worst-covered varieties in current language models. For scale: PADIC… See the full description on the dataset page: https://huggingface.co/datasets/dzcorpora/algerian-darija-customer-service-sample.tabulartext-generation1K<n<10K1 likes47 downloads10d agoHugging Face04AngieYYF /Frames-synthetic-customer-service-dialogue Frames Synthetic Customer Service Dialogues This contains a repository of customer service line synthetic user dialogues with goals, augmented from Frames using Qwen2.5-32B. The datasets are intended for training and evaluating machine generated text detectors in dialogue settings. Dataset Structure The datasets are of parquet file format and contain the following columns: Column Description dia_no Unique ID for each dialogue. Dialogues with the same ID… See the full description on the dataset page: https://huggingface.co/datasets/AngieYYF/Frames-synthetic-customer-service-dialogue.tabulartext-generation1K<n<10K3 likes42 downloads1y agoHugging Face05kaushik-systalyze /customer-transcript-long-dialog Customer Transcript Long Dialogue Curated customer-support and transcript-analytics prompts mapped to a single fixed "analyze this transcript -> compact JSON" prompt, for benchmarking batched offline LLM inference on realistic workloads. Motivation and intended use This dataset provides a realistic transcript-analytics workload for batched offline-inference experiments: throughput benchmarking and predicted-vs-observed throughput validation. Rows carry token… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/customer-transcript-long-dialog.tabulartext-generation1K<n<10K0 likes42 downloads3mo agoHugging Face06kaushik-systalyze /customer-transcript-analytics Customer Transcript Analytics Curated customer-support and meeting transcripts mapped to a single fixed "analyze this transcript → compact JSON" prompt, for benchmarking batched offline LLM inference on realistic workloads. Motivation and intended use This dataset provides a realistic transcript-analytics workload for batched offline-inference experiments: throughput benchmarking and predicted-vs-observed throughput validation. Rows range from short support chats… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/customer-transcript-analytics.tabulartext-generation1K<n<10K0 likes37 downloads4mo agoHugging Face07Ghanashyaam /CallAgentAI-Hinglish-Customer-Service CallAgent AI: Hinglish Business Conversations Dataset This dataset contains synthetic, high-quality "Hinglish" (Hindi + English code-switching) customer service interactions. It was generated by CallAgent AI (callagentai.in) — India's leading AI voice receptionist platform designed specifically for Indian SMBs. Why this dataset exists Global voice AI models often fail to capture the unique nuances of Indian business calls, which heavily rely on fluid language… See the full description on the dataset page: https://huggingface.co/datasets/Ghanashyaam/CallAgentAI-Hinglish-Customer-Service.tabulartext-generationn<1K0 likes37 downloads25d agoHugging Face08eagle0504 /multireward-grpo-fintech-customer-comms Multi-Reward GRPO — Synthetic Fintech Customer Communications Synthetic multi-turn customer-service conversations for a fictional bank ("Bank of XYZ"), generated for the empirical Section of "Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis". Each conversation ends with m parallel sampled bot replies, each scored on three verifiable reward channels designed for fintech customer service. This is the multi-reward GRPO group structure on a real generation… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-fintech-customer-comms.tabulartext-generation1K<n<10K0 likes24 downloads4mo agoHugging Face09kaushik-systalyze /customer-transcript-short-control Customer Transcript Short Control Curated customer-support and transcript-analytics prompts mapped to a single fixed "analyze this transcript -> compact JSON" prompt, for benchmarking batched offline LLM inference on realistic workloads. Motivation and intended use This dataset provides a realistic transcript-analytics workload for batched offline-inference experiments: throughput benchmarking and predicted-vs-observed throughput validation. Rows carry token… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/customer-transcript-short-control.tabulartext-generation1K<n<10K0 likes22 downloads3mo agoHugging Face10kaushik-systalyze /customer-transcript-holdout-eval Customer Transcript Holdout Eval Curated customer-support and transcript-analytics prompts mapped to a single fixed "analyze this transcript -> compact JSON" prompt, for benchmarking batched offline LLM inference on realistic workloads. Motivation and intended use This dataset provides a realistic transcript-analytics workload for batched offline-inference experiments: throughput benchmarking and predicted-vs-observed throughput validation. Rows carry token… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/customer-transcript-holdout-eval.tabulartext-generation1K<n<10K0 likes17 downloads3mo agoHugging Face11kaushik-systalyze /customer-transcript-source Customer Transcript Source Curated customer-support and transcript-analytics prompts mapped to a single fixed "analyze this transcript -> compact JSON" prompt, for benchmarking batched offline LLM inference on realistic workloads. Motivation and intended use This dataset provides a realistic transcript-analytics workload for batched offline-inference experiments: throughput benchmarking and predicted-vs-observed throughput validation. Rows carry token accounting… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/customer-transcript-source.tabulartext-generation1K<n<10K0 likes13 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.