CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ServiceNow /repliqa RepLiQA - Repository of Likely Question-Answer for benchmarking NeurIPS Datasets presentation Dataset Summary RepLiQA is an evaluation dataset that contains Context-Question-Answer triplets, where contexts are non-factual but natural-looking documents about made up entities such as people or places that do not exist in reality. RepLiQA is human-created, and designed to test for the ability of Large Language Models (LLMs) to find and use contextual information in provided… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/repliqa.documentquestion-answering10K<n<100K24 likes8.5k downloads1y agoHugging Face02datatab /serbian-llm-benchmark Serbian LLM Evaluation Dataset Welcome to the Serbian LLM Evaluation Dataset, your one-stop solution for evaluating Serbian Language Models (LLMs) like never before! This comprehensive toolkit empowers you to measure model performance across diverse domains in Serbian, ensuring your models are smarter, faster, and more intuitive. Whether you're a researcher, developer, or just an enthusiast—this dataset is tailor-made to help your LLM thrive. 🔍 What's Inside? This… See the full description on the dataset page: https://huggingface.co/datasets/datatab/serbian-llm-benchmark.textquestion-answering10K<n<100K8 likes1.3k downloads2y agoHugging Face03ServiceNow /drbench DRBench: A Realistic Benchmark for Enterprise Deep Research 📄 Paper | 💻 GitHub | 💬 Discord DRBench is the first of its kind benchmark designed to evaluate deep research agents on complex, open-ended enterprise deep research tasks. It tests an agent's ability to conduct multi-hop, insight-driven research across public and private data sources, just like a real enterprise analyst. ✨ Key Features 🔎 Real Deep Research Tasks: Not simple fact lookups. Tasks… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/drbench.documentquestion-answeringn<1K5 likes831 downloads6mo agoHugging Face04ServiceNow-AI /AgentJudgeBench AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling A benchmark for systematically evaluating how reliably LLM judges assess agentic tool-calling workflows across structured, dependency-driven tasks. Why this benchmark? AgentJudgeBench measures how reliably LLM judges assess agentic tool-calling outputs. It provides 3,808 benchmark records spanning six DAG topologies and three difficulty… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/AgentJudgeBench.tabularquestion-answering100K<n<1M0 likes461 downloads25d agoHugging Face05hulk10 /service_public_pro-full-documentstextquestion-answering1K<n<10K0 likes393 downloads19h agoHugging Face06hulk10 /service_public_part-full-documents 🇫🇷 Dataset Service-Public.fr – Fiches administratives structurées Ce dataset est constitué à partir des contenus officiels publiés sur la plateformeService-Public.fr.Il regroupe des fiches pratiques et ressources administratives à destination des particuliers et des professionnels, couvrant un large éventail de démarches et de thématiques de l’administration française. La structure et la méthodologie de ce dataset sont fortement inspirées du dataset Service-Public.fr practical… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/service_public_part-full-documents.textquestion-answering1K<n<10K0 likes374 downloads19h agoHugging Face07ServiceNow /Dr-CiK Dr-CiK: A Testbed for Foresight-Driven Agents Dr-CiK is a benchmark for evaluating whether agents can retrieve forecasting-relevant context from a noisy document corpus, filter out distractors, distill the retrieved context into forecast-useful evidence, and produce forecasts grounded in that evidence. Real-world time-series forecasting often depends not only on historical observations but also on external context that must be actively discovered from heterogeneous, noisy… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/Dr-CiK.tabulartime-series-forecasting10K<n<100K3 likes333 downloads3mo agoHugging Face08serval-uni-lu /orc-bench ORC-bench Task 1: Topological Path Finding Task 2: Topological Connectivity Task 3: Linear Power Flow Task 4: Contingency Analysis Task 5: Power Grid ControlTask 6: Power Flow Optimization Task 1: Topological Path Finding Problem Formulation This task assesses the spatial reasoning ability of the model by asking it to determine the shortest path between two specific buses in a given power grid state. The grid state… See the full description on the dataset page: https://huggingface.co/datasets/serval-uni-lu/orc-bench.textquestion-answering10K<n<100K0 likes205 downloads5mo agoHugging Face09serdarsrts /turkish-court-decisions-duplicate Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet), 1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden yerel/istinaf mahkemeleri. Kapsam Kaynak Karar sayısı Yıl aralığı Metin Dosya Yargıtay (yargitay) 9.820.145 1997–2026 19.5 milyar karakter 17 Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-court-decisions-duplicate.tabulartext-generation10M<n<100M1 likes203 downloads28d agoHugging Face10RichardSakaguchiMS /brazilian-customer-service-conversations Brazilian Customer Service Conversations Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR). De um like me apoie em manter esse dataset! Descricao Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de: Chatbots de atendimento Classificacao de intencao (intent classification) Analise de sentimento em conversas Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.texttext-classificationn<1K5 likes141 downloads10mo agoHugging Face11sermonindex /bible-reference Bible Reference Corpus Thirteen aligned reference datasets for study of the biblical text: Greek and Hebrew lexicons keyed to Strong's numbers, an interlinear word map, the critical apparatus of eight Greek editions, cross-reference and topical indexes, and geolocated places. Published by SermonIndex. Everything in this repository is public domain or CC BY 4.0. Sources with share-alike terms are kept in a separate repository, sermonindex/bible-reference-sa, so that a share-alike… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-reference.tabulartext-retrieval100K<n<1M0 likes135 downloads16d agoHugging Face12datatab /ultrafeedback_binarized_serbian Dataset Card for UltraFeedback Binarized Serbian Dataset Description This dataset is a Serbian-translated version of the UltraFeedback dataset, utilized for training Zephyr-7Β-β. The original dataset comprises 64k English-language prompts, each paired with four completions from various models. In this Serbian version, the prompts and completions have been translated into Serbian. The dataset creation process remains the same: selecting the completion with the highest… See the full description on the dataset page: https://huggingface.co/datasets/datatab/ultrafeedback_binarized_serbian.tabulartext-generation100K<n<1M0 likes90 downloads3y agoHugging Face13serhanayberkkilic /physiotherapy-evidence-qa 🏥 Physiotherapy Evidence QA: A Bilingual Clinical Corpus Physiotherapy Evidence QA is a large-scale, expert-curated bilingual dataset comprising 143,711 aligned question-answer pairs. It focuses on evidence-based physiotherapy, musculoskeletal rehabilitation, outcome measures, and clinical research methodology. This corpus is designed to facilitate the development of Medical Large Language Models (Med-LLMs), Clinical Decision Support Systems (CDSS), and Cross-Lingual Information… See the full description on the dataset page: https://huggingface.co/datasets/serhanayberkkilic/physiotherapy-evidence-qa.textquestion-answering100K<n<1M4 likes87 downloads10mo agoHugging Face14sermonindex /early-church-fathers Early Church Fathers — Scripture Citation Index 68,240 passages from 349 Church Fathers, each keyed to the Bible verse it comments on. Drawn from 20,253 distinct works and covering all 66 books. This is a patristic catena in machine-readable form: given a verse, it returns what the Fathers said about it. Nothing comparable exists as an open dataset — the underlying translations are freely available, but the verse-level alignment is the work, and that is what this releases.… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/early-church-fathers.tabulartext-retrieval10K<n<100K0 likes67 downloads16d agoHugging Face15emgena /omnimcp_nextjs_server_actions_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_nextjs_server_actions_teaser.texttext-generationn<1K0 likes67 downloads9d agoHugging Face16FunDialogues /customer-service-robot-support This Dialogue Comprised of fictitious examples of dialogues between a customer encountering problems with a robotic arm and a technical support agent. Check out the example below: "id": 1, "description": "Robotic arm calibration issue", "dialogue": "Customer: My robotic arm seems to be misaligned. It's not picking objects accurately. What can I do? Agent: It appears that the arm may need recalibration. Please follow the instructions in the user manual to reset the calibration… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/customer-service-robot-support.tabularquestion-answeringn<1K2 likes63 downloads3y agoHugging Face17imran-siddique /context-as-a-service CaaS Benchmark Corpus v1 A diverse collection of synthetic enterprise documents for benchmarking context extraction and RAG systems. Dataset Description This dataset contains 16 representative enterprise documents spanning multiple formats and domains, designed to evaluate: Structure-aware indexing - Can the system identify high-value vs. low-value content? Time decay relevance - Does the system properly weight recent vs. old information? Pragmatic truth detection - Can… See the full description on the dataset page: https://huggingface.co/datasets/imran-siddique/context-as-a-service.tabulartext-retrievaln<1K0 likes63 downloads8mo agoHugging Face18serdarsrts /turkish-competition-authority-decisions Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026 The complete published decision history of the Turkish Competition Authority (Rekabet Kurumu) — every Competition Board decision the regulator has made public, in full text, with derived structural metadata. 10,367 decisions · 113,297 pages · 323 million characters · 29 years Every decision carries its outcome, the articles of Law 4054 it turns on, the panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-competition-authority-decisions.tabulartext-classification10K<n<100K0 likes61 downloads28d agoHugging Face19leeroy-jankins /DoD-Instruction-8130-01-Installation-of-Geospatial-Information-And-Services 🗺️ DoD Installation Geospatial Information and Services Question-Answer Dataset Source: DoD Instruction 8130.01 Source Effective Date: April 9, 2015 Change Incorporated: Change 3, effective August 4, 2020 Source Organization: Office of the Under Secretary of Defense for Acquisition and Sustainment Source Ownership: United States Department of Defense 📋 Overview Dataset Summary The DoD Installation Geospatial Information and Services… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8130-01-Installation-of-Geospatial-Information-And-Services.documentquestion-answering0 likes55 downloads2mo agoHugging Face20Sergey23214 /coverture-103k-gender-history 🏛 COVERTURE: Institutional Gender History Corpus (103,270 Evidentiary Dossiers) "Culture is not a neutral mirror of reality. It is a disciplinary machine that normalizes domination through humor, law, romance, and erasure." The Coverture Corpus is a large-scale, evidentiary research dataset comprising 103,270 structured analytical dossiers documenting the institutional, legal, economic, domestic, and cultural technologies of patriarchal control over women from Antiquity to… See the full description on the dataset page: https://huggingface.co/datasets/Sergey23214/coverture-103k-gender-history.texttext-classification100K<n<1M0 likes52 downloads18d agoHugging Face21freococo /imam_albani_weaknfab_series_dataset Imam al-Albani Weak & Fabricated Hadith Dataset (AR–EN–MY) This dataset contains weak, rejected, or fabricated hadiths classified byImam Muhammad Nasir al-Din al-Albani, presented in Arabic, English, and Myanmar (Burmese). Translated with Gemini Pro 3.0. Dataset Structure Each row represents one hadith with a global unique ID and multilingual fields. CSV Column Order global_id – Unique sequential ID (primary key) hadith_arabic_text – Original Arabic text… See the full description on the dataset page: https://huggingface.co/datasets/freococo/imam_albani_weaknfab_series_dataset.texttranslation1K<n<10K0 likes50 downloads9mo agoHugging Face22leeroy-jankins /DOD-Enterprise-DevSecOps-Reference-Design-AWS-Managed-Services DoD Enterprise DevSecOps AWS Managed Services Reference Design Question-Answer Dataset Maintainer: Terry Eppler Owner: US Federal Government Dataset Summary This dataset contains document-grounded question-and-answer records based on DoD Enterprise DevSecOps Reference Design: AWS Managed Services (DoD IaC Baseline), Version 0.2, September 2021. The source presents a draft Department of Defense reference design for implementing a DevSecOps software factory… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Enterprise-DevSecOps-Reference-Design-AWS-Managed-Services.documentquestion-answering0 likes47 downloads2mo agoHugging Face23FunDialogues /customer-service-grocery-cashier This Dialogue Comprised of fictitious examples of dialogues between a customer at a grocery store and the cashier. Check out the example below: "id": 1, "description": "Price inquiry", "dialogue": "Customer: Excuse me, could you tell me the price of the apples per pound? Cashier: Certainly! The price for the apples is $1.99 per pound." How to Load Dialogues Loading dialogues can be accomplished using the fun dialogues library or Hugging Face datasets library.… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/customer-service-grocery-cashier.tabularquestion-answeringn<1K4 likes45 downloads3y agoHugging Face24SerFabio89 /italian-open-sft-chat-dataset Italian Open SFT Chat Dataset An Italian-first, model-neutral synthetic SFT and chat dataset for fine-tuning Italian-capable LLMs. It targets instruction tuning, Italian chat behavior, structured output generation, JSON/YAML/CSV format following, coding assistance, safety refusals, multi-turn dialogue and reasoning-style final answers. This v0.1.0 package does not include long-context QA records. This dataset is intended for users searching for an Italian instruction tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/SerFabio89/italian-open-sft-chat-dataset.texttext-generation10K<n<100K0 likes43 downloads5mo agoHugging Face25beatsprom /autonomous-cloud-gpu-slurm-serving-suite ⚡ Autonomous Cloud GPU Infrastructure, Slurm Orchestration & Distributed Serving Suite (2026) A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous AI Supercomputing & LLM Serving Agents ⚡ Overview & Industry Problem Operating massive AI supercomputers (thousands of NVIDIA H100/H200 and Blackwell GPUs) requires coordinating Slurm cluster schedules, topology-aware NVLink cliques, NCCL AllReduce rings, RoCE v2 lossless fabrics… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-cloud-gpu-slurm-serving-suite.tabulartext-generation1K<n<10K0 likes42 downloads9d agoHugging Face26serenalyoko /HiCUPID 💖 HiCUPID Dataset 📌 Dataset Summary We introduce 💖 HiCUPID, a benchmark designed to train and evaluate Large Language Models (LLMs) for personalized AI assistant applications. Why HiCUPID? Most open-source conversational datasets lack personalization, making it hard to develop AI assistants that adapt to users. HiCUPID fills this gap by providing: ✅ A tailored dataset with structured dialogues and QA pairs. ✅ An automated evaluation model (based… See the full description on the dataset page: https://huggingface.co/datasets/serenalyoko/HiCUPID.tabularquestion-answering100K<n<1M0 likes40 downloads3mo agoHugging Face27louisbrulenaudet /code-service-national Code du service national, non-instruct (2025-07-11) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-service-national.tabulartext-generationn<1K0 likes35 downloads1y agoHugging Face28Chamaka8 /Serendip-sft-sinhala Serendip-SFT-Sinhala Dataset 🇱🇰 📊 Dataset Summary Serendip-SFT-Sinhala is a large-scale Sinhala instruction-tuning dataset with 293,613 high-quality examples for supervised fine-tuning (SFT) of large language models. Created to train SerendipLLM, a Sinhala language model designed to excel at instruction-following, question-answering, summarization, and text classification. 🌟 Highlights 🇱🇰 293,613 Sinhala examples (largest Sinhala SFT dataset) 📚 4 task… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/Serendip-sft-sinhala.texttext-generation100K<n<1M0 likes34 downloads7mo agoHugging Face29louisbrulenaudet /code-impositions-biens-services Code des impositions sur les biens et services, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-impositions-biens-services.tabulartext-generation1K<n<10K0 likes32 downloads1y agoHugging Face30W-L /Customer-service-tickets-qwen-qa Customer Support Tickets QA (English) — Qwen SFT Dataset This dataset is formatted for supervised fine-tuning (SFT) of Qwen-style chat models on customer support email tasks. source dataset: Tobi-Bueck/customer-support-tickets It is designed for training models to read a customer ticket, understand its context, and generate an appropriate support response. Depending on the prompt design, the same data can also support auxiliary tasks such as queue prediction, priority prediction… See the full description on the dataset page: https://huggingface.co/datasets/W-L/Customer-service-tickets-qwen-qa.texttext-generation10K<n<100K1 likes32 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.