CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ServiceNow /drbench DRBench: A Realistic Benchmark for Enterprise Deep Research 📄 Paper | 💻 GitHub | 💬 Discord DRBench is the first of its kind benchmark designed to evaluate deep research agents on complex, open-ended enterprise deep research tasks. It tests an agent's ability to conduct multi-hop, insight-driven research across public and private data sources, just like a real enterprise analyst. ✨ Key Features 🔎 Real Deep Research Tasks: Not simple fact lookups. Tasks… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/drbench.documentquestion-answeringn<1K5 likes824 downloads6mo agoHugging Face02ServiceNow-AI /AgentJudgeBench AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling A benchmark for systematically evaluating how reliably LLM judges assess agentic tool-calling workflows across structured, dependency-driven tasks. Why this benchmark? AgentJudgeBench measures how reliably LLM judges assess agentic tool-calling outputs. It provides 3,808 benchmark records spanning six DAG topologies and three difficulty… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/AgentJudgeBench.tabularquestion-answering100K<n<1M0 likes493 downloads24d agoHugging Face03serval-uni-lu /orc-bench ORC-bench Task 1: Topological Path Finding Task 2: Topological Connectivity Task 3: Linear Power Flow Task 4: Contingency Analysis Task 5: Power Grid ControlTask 6: Power Flow Optimization Task 1: Topological Path Finding Problem Formulation This task assesses the spatial reasoning ability of the model by asking it to determine the shortest path between two specific buses in a given power grid state. The grid state… See the full description on the dataset page: https://huggingface.co/datasets/serval-uni-lu/orc-bench.textquestion-answering10K<n<100K0 likes205 downloads5mo agoHugging Face04serdarsrts /turkish-court-decisions-duplicate Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-court-decisions-duplicate.tabulartext-generation10M<n<100M1 likes201 downloads27d agoHugging Face05AngieYYF /SPADE-customer-service-dialogue SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection Paper | Code SPADE contains a repository of customer service line synthetic user dialogues with goals, augmented from MultiWOZ 2.1 using GPT-3.5 and Llama 70B. The datasets are intended for training and evaluating machine generated text detectors in dialogue settings. There are 15 English datasets generated using 5 different augmentation methods and 2 large language models… See the full description on the dataset page: https://huggingface.co/datasets/AngieYYF/SPADE-customer-service-dialogue.tabulartext-generation10K<n<100K3 likes187 downloads1y agoHugging Face06sermonindex /bible The Bible in 1,004 Languages 14,497,397 verses across 1,253 translations in 1,004 languages, every verse keyed to the same chapter-and-verse address so that any two languages can be aligned by joining on book, chapter and verse. The Bible is the most widely translated text in existence, and for several hundred of the languages here it is the largest — sometimes the only — substantial digitised text. That makes this corpus unusually useful for low-resource machine translation… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible.tabulartext-generation10M<n<100M2 likes113 downloads15d agoHugging Face07datatab /ultrafeedback_binarized_serbian Dataset Card for UltraFeedback Binarized Serbian Dataset Description This dataset is a Serbian-translated version of the UltraFeedback dataset, utilized for training Zephyr-7Β-β. The original dataset comprises 64k English-language prompts, each paired with four completions from various models. In this Serbian version, the prompts and completions have been translated into Serbian. The dataset creation process remains the same: selecting the completion with the highest… See the full description on the dataset page: https://huggingface.co/datasets/datatab/ultrafeedback_binarized_serbian.tabulartext-generation100K<n<1M0 likes85 downloads3y agoHugging Face08sermonindex /early-church-fathers Early Church Fathers — Scripture Citation Index 68,240 passages from 349 Church Fathers, each keyed to the Bible verse it comments on. Drawn from 20,253 distinct works and covering all 66 books. This is a patristic catena in machine-readable form: given a verse, it returns what the Fathers said about it. Nothing comparable exists as an open dataset — the underlying translations are freely available, but the verse-level alignment is the work, and that is what this releases.… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/early-church-fathers.tabulartext-retrieval10K<n<100K0 likes66 downloads15d agoHugging Face09sermonindex /bible-parallel-english Parallel Bible — English Translations and Ancient Versions A verse-aligned parallel corpus of the Protestant Bible in seventeen English translations, spanning 1599 to 2022, plus the Latin Vulgate and Syriac Peshitta for the New Testament. Looking for every language? This repository is a curated English set, chosen for spread across translation families and small enough to load whole. For the full corpus — 1,253 translations in 1,004 languages, 14.4M verses — see… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-parallel-english.tabulartext-generation100K<n<1M0 likes52 downloads15d agoHugging Face10thientrangngv /SERA-KimiK3-Django-SWEAgent-Cliff32k-T1 SERA Kimi-K3 Django SWE-Agent — Cliff-chunked T1 (first rollout) 572 training records built from 210 Kimi-K3 SWE-agent trajectories on Django, split to fit a 32,768-token context with CliffCompaction instead of being truncated. Why chunked A 100+ step agent rollout does not fit a 32k training window — 76% of the source T1 trajectories exceed it. Truncating them throws away most of the supervision, and trains the model on a context format it never sees at… See the full description on the dataset page: https://huggingface.co/datasets/thientrangngv/SERA-KimiK3-Django-SWEAgent-Cliff32k-T1.tabulartext-generationn<1K1 likes50 downloads1mo agoHugging Face11dzcorpora /algerian-darija-customer-service-samplegated Algerian Darija customer messages — stratified sample 500 spontaneous Algerian Darija messages, written by real customers, drawn from a first-party corpus of 869,166 customer messages. Every message here is unique after normalization, de-identified, and typed by a human — nothing elicited, translated, scraped or generated. Algerian Darija (ISO 639-3 arq) is spoken by around 45 million people and is one of the worst-covered varieties in current language models. For scale: PADIC… See the full description on the dataset page: https://huggingface.co/datasets/dzcorpora/algerian-darija-customer-service-sample.tabulartext-generation1K<n<10K1 likes47 downloads11d agoHugging Face12AngieYYF /Frames-synthetic-customer-service-dialogue Frames Synthetic Customer Service Dialogues This contains a repository of customer service line synthetic user dialogues with goals, augmented from Frames using Qwen2.5-32B. The datasets are intended for training and evaluating machine generated text detectors in dialogue settings. Dataset Structure The datasets are of parquet file format and contain the following columns: Column Description dia_no Unique ID for each dialogue. Dialogues with the same ID… See the full description on the dataset page: https://huggingface.co/datasets/AngieYYF/Frames-synthetic-customer-service-dialogue.tabulartext-generation1K<n<10K3 likes42 downloads1y agoHugging Face13beatsprom /autonomous-cloud-gpu-slurm-serving-suite ⚡ Autonomous Cloud GPU Infrastructure, Slurm Orchestration & Distributed Serving Suite (2026) A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous AI Supercomputing & LLM Serving Agents ⚡ Overview & Industry Problem Operating massive AI supercomputers (thousands of NVIDIA H100/H200 and Blackwell GPUs) requires coordinating Slurm cluster schedules, topology-aware NVLink cliques, NCCL AllReduce rings, RoCE v2 lossless fabrics… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-cloud-gpu-slurm-serving-suite.tabulartext-generation1K<n<10K0 likes42 downloads8d agoHugging Face14serenalyoko /HiCUPID 💖 HiCUPID Dataset 📌 Dataset Summary We introduce 💖 HiCUPID, a benchmark designed to train and evaluate Large Language Models (LLMs) for personalized AI assistant applications. Why HiCUPID? Most open-source conversational datasets lack personalization, making it hard to develop AI assistants that adapt to users. HiCUPID fills this gap by providing: ✅ A tailored dataset with structured dialogues and QA pairs. ✅ An automated evaluation model (based… See the full description on the dataset page: https://huggingface.co/datasets/serenalyoko/HiCUPID.tabularquestion-answering100K<n<1M0 likes40 downloads3mo agoHugging Face15thientrangngv /SERA-KimiK3-Django-SWEAgent-Cliff32k-T2 SERA Kimi-K3 Django SWE-Agent — Cliff-chunked T2 (second rollout) 227 training records built from 137 Kimi-K3 SWE-agent trajectories on Django, split to fit a 32,768-token context with CliffCompaction instead of being truncated. Why chunked A 100+ step agent rollout does not fit a 32k training window — 27% of the source T2 trajectories exceed it. Truncating them throws away most of the supervision, and trains the model on a context format it never sees at… See the full description on the dataset page: https://huggingface.co/datasets/thientrangngv/SERA-KimiK3-Django-SWEAgent-Cliff32k-T2.tabulartext-generationn<1K1 likes37 downloads1mo agoHugging Face16Ghanashyaam /CallAgentAI-Hinglish-Customer-Service CallAgent AI: Hinglish Business Conversations Dataset This dataset contains synthetic, high-quality "Hinglish" (Hindi + English code-switching) customer service interactions. It was generated by CallAgent AI (callagentai.in) — India's leading AI voice receptionist platform designed specifically for Indian SMBs. Why this dataset exists Global voice AI models often fail to capture the unique nuances of Indian business calls, which heavily rely on fluid language… See the full description on the dataset page: https://huggingface.co/datasets/Ghanashyaam/CallAgentAI-Hinglish-Customer-Service.tabulartext-generationn<1K0 likes37 downloads25d agoHugging Face17louisbrulenaudet /code-service-national Code du service national, non-instruct (2025-07-11) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-service-national.tabulartext-generationn<1K0 likes35 downloads1y agoHugging Face18louisbrulenaudet /code-impositions-biens-services Code des impositions sur les biens et services, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-impositions-biens-services.tabulartext-generation1K<n<10K0 likes31 downloads1y agoHugging Face19fineset-io /time-series-foundation-models-papers Time Series Foundation Models Papers — FineSet A research-paper dataset on Time Series Foundation Models Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Time Series Foundation Models Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/time-series-foundation-models-papers.tabulartext-classificationn<1K0 likes24 downloads3mo agoHugging Face20PhillyMac /Servant_Leadership_Practical Servant Leadership — Practical This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Servant_Leadership_Practical.tabulartext-generationn<1K0 likes11 downloads5mo agoHugging Face21aiagentkarl /mcp-server-catalog MCP Server Catalog A comprehensive catalog of 38 Model Context Protocol (MCP) servers for AI agents, covering data access, agent infrastructure, business-to-agent interfaces, compliance, and more. Overview This dataset provides a structured catalog of MCP servers that give AI agents access to real-world data and capabilities. Each server follows the MCP standard and can be used with Claude, GPT, and other LLMs that support tool use. Categories Category… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/mcp-server-catalog.tabulartext-generationn<1K1 likes9 downloads6mo agoHugging Face22PhillyMac /Servant_Leadership_Theory Servant Leadership — Theory This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Servant_Leadership_Theory.tabulartext-generationn<1K0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.