CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ServiceNow /drbench DRBench: A Realistic Benchmark for Enterprise Deep Research 📄 Paper | 💻 GitHub | 💬 Discord DRBench is the first of its kind benchmark designed to evaluate deep research agents on complex, open-ended enterprise deep research tasks. It tests an agent's ability to conduct multi-hop, insight-driven research across public and private data sources, just like a real enterprise analyst. ✨ Key Features 🔎 Real Deep Research Tasks: Not simple fact lookups. Tasks… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/drbench.documentquestion-answeringn<1K5 likes831 downloads6mo agoHugging Face02ServiceNow-AI /AgentJudgeBench AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling A benchmark for systematically evaluating how reliably LLM judges assess agentic tool-calling workflows across structured, dependency-driven tasks. Why this benchmark? AgentJudgeBench measures how reliably LLM judges assess agentic tool-calling outputs. It provides 3,808 benchmark records spanning six DAG topologies and three difficulty… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/AgentJudgeBench.tabularquestion-answering100K<n<1M0 likes461 downloads25d agoHugging Face03ServiceNow /Dr-CiK Dr-CiK: A Testbed for Foresight-Driven Agents Dr-CiK is a benchmark for evaluating whether agents can retrieve forecasting-relevant context from a noisy document corpus, filter out distractors, distill the retrieved context into forecast-useful evidence, and produce forecasts grounded in that evidence. Real-world time-series forecasting often depends not only on historical observations but also on external context that must be actively discovered from heterogeneous, noisy… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/Dr-CiK.tabulartime-series-forecasting10K<n<100K3 likes333 downloads3mo agoHugging Face04serval-uni-lu /orc-bench ORC-bench Task 1: Topological Path Finding Task 2: Topological Connectivity Task 3: Linear Power Flow Task 4: Contingency Analysis Task 5: Power Grid ControlTask 6: Power Flow Optimization Task 1: Topological Path Finding Problem Formulation This task assesses the spatial reasoning ability of the model by asking it to determine the shortest path between two specific buses in a given power grid state. The grid state… See the full description on the dataset page: https://huggingface.co/datasets/serval-uni-lu/orc-bench.textquestion-answering10K<n<100K0 likes205 downloads5mo agoHugging Face05RichardSakaguchiMS /brazilian-customer-service-conversations Brazilian Customer Service Conversations Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR). De um like me apoie em manter esse dataset! Descricao Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de: Chatbots de atendimento Classificacao de intencao (intent classification) Analise de sentimento em conversas Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.texttext-classificationn<1K5 likes141 downloads10mo agoHugging Face06sermonindex /bible-reference Bible Reference Corpus Thirteen aligned reference datasets for study of the biblical text: Greek and Hebrew lexicons keyed to Strong's numbers, an interlinear word map, the critical apparatus of eight Greek editions, cross-reference and topical indexes, and geolocated places. Published by SermonIndex. Everything in this repository is public domain or CC BY 4.0. Sources with share-alike terms are kept in a separate repository, sermonindex/bible-reference-sa, so that a share-alike… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-reference.tabulartext-retrieval100K<n<1M0 likes135 downloads15d agoHugging Face07sermonindex /early-church-fathers Early Church Fathers — Scripture Citation Index 68,240 passages from 349 Church Fathers, each keyed to the Bible verse it comments on. Drawn from 20,253 distinct works and covering all 66 books. This is a patristic catena in machine-readable form: given a verse, it returns what the Fathers said about it. Nothing comparable exists as an open dataset — the underlying translations are freely available, but the verse-level alignment is the work, and that is what this releases.… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/early-church-fathers.tabulartext-retrieval10K<n<100K0 likes67 downloads15d agoHugging Face08imran-siddique /context-as-a-service CaaS Benchmark Corpus v1 A diverse collection of synthetic enterprise documents for benchmarking context extraction and RAG systems. Dataset Description This dataset contains 16 representative enterprise documents spanning multiple formats and domains, designed to evaluate: Structure-aware indexing - Can the system identify high-value vs. low-value content? Time decay relevance - Does the system properly weight recent vs. old information? Pragmatic truth detection - Can… See the full description on the dataset page: https://huggingface.co/datasets/imran-siddique/context-as-a-service.tabulartext-retrievaln<1K0 likes63 downloads8mo agoHugging Face09Chamaka8 /Serendip-sft-sinhala Serendip-SFT-Sinhala Dataset 🇱🇰 📊 Dataset Summary Serendip-SFT-Sinhala is a large-scale Sinhala instruction-tuning dataset with 293,613 high-quality examples for supervised fine-tuning (SFT) of large language models. Created to train SerendipLLM, a Sinhala language model designed to excel at instruction-following, question-answering, summarization, and text classification. 🌟 Highlights 🇱🇰 293,613 Sinhala examples (largest Sinhala SFT dataset) 📚 4 task… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/Serendip-sft-sinhala.texttext-generation100K<n<1M0 likes34 downloads7mo agoHugging Face10W-L /Customer-service-tickets-qwen-qa Customer Support Tickets QA (English) — Qwen SFT Dataset This dataset is formatted for supervised fine-tuning (SFT) of Qwen-style chat models on customer support email tasks. source dataset: Tobi-Bueck/customer-support-tickets It is designed for training models to read a customer ticket, understand its context, and generate an appropriate support response. Depending on the prompt design, the same data can also support auxiliary tasks such as queue prediction, priority prediction… See the full description on the dataset page: https://huggingface.co/datasets/W-L/Customer-service-tickets-qwen-qa.texttext-generation10K<n<100K1 likes32 downloads5mo agoHugging Face11Rabe3 /sera-phase2-saudi-dialect-rag SERA Saudi Dialect RAG Dataset (Phase 2 - Domain) Domain-specific RAG fine-tuning dataset for Saudi Arabic dialect, focused on Saudi Electricity Regulatory Authority (SERA) documents. Format LlamaFactory Alpaca format: Field Description instruction System prompt + real document chunk as context + question in Saudi dialect input Always empty output Answer in Saudi dialect Usage with LlamaFactory Copy the JSON files into your LlamaFactory… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/sera-phase2-saudi-dialect-rag.textquestion-answering1K<n<10K0 likes25 downloads7mo agoHugging Face12pkchwy /turkish-comprehensive-movie-series-dataset Beyazperde Film & Series Dataset This dataset contains a comprehensive collection of Turkish films and TV series from Beyazperde.com, including detailed information about movies, series, cast, reviews, and ratings. Dataset Summary Total Movies: 27,227 Total Series: 11,240 Total Entries: 38,467 File Size: ~222 MB Format: JSONL (JSON Lines) Language: Turkish Source: Beyazperde.com Data Structure Each line in the JSONL file contains a JSON object… See the full description on the dataset page: https://huggingface.co/datasets/pkchwy/turkish-comprehensive-movie-series-dataset.imagetext-classification10K<n<100K4 likes22 downloads1y agoHugging Face13rileyseaburg /spotless-customer-service-training Spotless Bin Co Customer Service Training Data Training data for a customer service AI model for Spotless Bin Co, a residential trash can cleaning service. Dataset Description This dataset contains 8,776 conversational examples across 5 categories: Category Count Description FAQs 1,951 Frequently asked questions Service 1,925 Service explanation dialogues Objections 1,925 Objection handling examples Booking 1,975 Booking flow conversations Brand 1,000… See the full description on the dataset page: https://huggingface.co/datasets/rileyseaburg/spotless-customer-service-training.texttext-generation1K<n<10K0 likes21 downloads9mo agoHugging Face14Sergey23214 /sud-corpus-v2 SUD Corpus v2: исторический ущерб женщинам как система Этот корпус создан, чтобы отдельные истории женщин перестали выглядеть набором случайных исключений и стали читаемыми как повторяемые механизмы права, труда, семьи, медицины, культуры, религии и производства знания. Русская версия Что это за корпус SUD Corpus v2 — открытый исследовательский корпус проекта СУД, содержащий 514 полнотекстовых материалов о системном историческом ущербе женщинам… See the full description on the dataset page: https://huggingface.co/datasets/Sergey23214/sud-corpus-v2.texttext-generationn<1K0 likes20 downloads2mo agoHugging Face15Ephraimmm /customer_service_dataset GTBank Customer Service – Synthetic Dataset Overview This dataset contains 500 synthetic user_query / assistant_reply pairs modeled on common GTBank (Guaranty Trust Bank) retail banking customer support interactions, such as password resets, debit card activation, failed transactions, and KYC/transfer limits. It is entirely synthetic — no real customer data, transcripts, or GTBank systems were used — and is intended for prototyping and testing customer-support NLP… See the full description on the dataset page: https://huggingface.co/datasets/Ephraimmm/customer_service_dataset.texttext-generationn<1K0 likes13 downloads3mo agoHugging Face16kshitizgajurel /Multilingual-Nepali-Customer-Care-Services-Datasetgatedtexttext-generation10K<n<100K0 likes11 downloads2y agoHugging Face17kshitizgajurel /Customer-Care-Services-Dataset-in-Nepalitexttext-generation10K<n<100K1 likes9 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.