datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
repliqa
RepLiQA - Repository of Likely Question-Answer for benchmarking
NeurIPS Datasets presentation
Dataset Summary
RepLiQA is an evaluation dataset that contains Context-Question-Answer triplets, where contexts are non-factual but natural-looking documents about made up entities such as people or places that do not exist in reality. RepLiQA is human-created, and designed to test for the ability of Large Language Models (LLMs) to find and use contextual information in provided… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/repliqa.serbian-llm-benchmark
Serbian LLM Evaluation Dataset
Welcome to the Serbian LLM Evaluation Dataset, your one-stop solution for evaluating Serbian Language Models (LLMs) like never before! This comprehensive toolkit empowers you to measure model performance across diverse domains in Serbian, ensuring your models are smarter, faster, and more intuitive. Whether you're a researcher, developer, or just an enthusiast—this dataset is tailor-made to help your LLM thrive.
🔍 What's Inside?
This… See the full description on the dataset page: https://huggingface.co/datasets/datatab/serbian-llm-benchmark.drbench
DRBench: A Realistic Benchmark for Enterprise Deep Research
📄 Paper | 💻 GitHub | 💬 Discord
DRBench is the first of its kind benchmark designed to evaluate deep research agents on complex, open-ended enterprise deep research tasks. It tests an agent's ability to conduct multi-hop, insight-driven research across public and private data sources, just like a real enterprise analyst.
✨ Key Features
🔎 Real Deep Research Tasks: Not simple fact lookups. Tasks… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/drbench.AgentJudgeBench
AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling
A benchmark for systematically evaluating how reliably LLM judges assess
agentic tool-calling workflows across structured, dependency-driven tasks.
Why this benchmark?
AgentJudgeBench measures how reliably LLM judges assess agentic tool-calling outputs. It provides 3,808 benchmark records spanning six DAG topologies and three difficulty… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/AgentJudgeBench.service_public_pro-full-documentsservice_public_part-full-documents
🇫🇷 Dataset Service-Public.fr – Fiches administratives structurées
Ce dataset est constitué à partir des contenus officiels publiés sur la plateformeService-Public.fr.Il regroupe des fiches pratiques et ressources administratives à destination des particuliers et des professionnels, couvrant un large éventail de démarches et de thématiques de l’administration française.
La structure et la méthodologie de ce dataset sont fortement inspirées du dataset Service-Public.fr practical… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/service_public_part-full-documents.Dr-CiK
Dr-CiK: A Testbed for Foresight-Driven Agents
Dr-CiK is a benchmark for evaluating whether agents can retrieve
forecasting-relevant context from a noisy document corpus, filter out
distractors, distill the retrieved context into forecast-useful evidence, and
produce forecasts grounded in that evidence.
Real-world time-series forecasting often depends not only on historical
observations but also on external context that must be actively discovered
from heterogeneous, noisy… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/Dr-CiK.orc-bench
ORC-bench
Task 1: Topological Path Finding
Task 2: Topological Connectivity
Task 3: Linear Power Flow
Task 4: Contingency Analysis
Task 5: Power Grid ControlTask 6: Power Flow Optimization
Task 1: Topological Path Finding
Problem Formulation
This task assesses the spatial reasoning ability of the model by asking it to determine the shortest path between two specific buses in a given power grid state. The grid state… See the full description on the dataset page: https://huggingface.co/datasets/serval-uni-lu/orc-bench.turkish-court-decisions-duplicate
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-court-decisions-duplicate.brazilian-customer-service-conversations
Brazilian Customer Service Conversations
Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR).
De um like me apoie em manter esse dataset!
Descricao
Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de:
Chatbots de atendimento
Classificacao de intencao (intent classification)
Analise de sentimento em conversas
Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.bible-reference
Bible Reference Corpus
Thirteen aligned reference datasets for study of the biblical text: Greek and
Hebrew lexicons keyed to Strong's numbers, an interlinear word map, the critical
apparatus of eight Greek editions, cross-reference and topical indexes, and
geolocated places.
Published by SermonIndex.
Everything in this repository is public domain or CC BY 4.0. Sources with
share-alike terms are kept in a separate repository,
sermonindex/bible-reference-sa,
so that a share-alike… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-reference.ultrafeedback_binarized_serbian
Dataset Card for UltraFeedback Binarized Serbian
Dataset Description
This dataset is a Serbian-translated version of the UltraFeedback dataset, utilized for training Zephyr-7Β-β. The original dataset comprises 64k English-language prompts, each paired with four completions from various models. In this Serbian version, the prompts and completions have been translated into Serbian. The dataset creation process remains the same: selecting the completion with the highest… See the full description on the dataset page: https://huggingface.co/datasets/datatab/ultrafeedback_binarized_serbian.physiotherapy-evidence-qa
🏥 Physiotherapy Evidence QA: A Bilingual Clinical Corpus
Physiotherapy Evidence QA is a large-scale, expert-curated bilingual dataset comprising 143,711 aligned question-answer pairs. It focuses on evidence-based physiotherapy, musculoskeletal rehabilitation, outcome measures, and clinical research methodology.
This corpus is designed to facilitate the development of Medical Large Language Models (Med-LLMs), Clinical Decision Support Systems (CDSS), and Cross-Lingual Information… See the full description on the dataset page: https://huggingface.co/datasets/serhanayberkkilic/physiotherapy-evidence-qa.early-church-fathers
Early Church Fathers — Scripture Citation Index
68,240 passages from 349 Church Fathers, each keyed to the Bible verse it
comments on. Drawn from 20,253 distinct works and covering all 66 books.
This is a patristic catena in machine-readable form: given a verse, it returns
what the Fathers said about it. Nothing comparable exists as an open dataset —
the underlying translations are freely available, but the verse-level alignment
is the work, and that is what this releases.… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/early-church-fathers.omnimcp_nextjs_server_actions_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_nextjs_server_actions_teaser.customer-service-robot-support
This Dialogue
Comprised of fictitious examples of dialogues between a customer encountering problems with a robotic arm and a technical support agent. Check out the example below:
"id": 1,
"description": "Robotic arm calibration issue",
"dialogue": "Customer: My robotic arm seems to be misaligned. It's not picking objects accurately. What can I do? Agent: It appears that the arm may need recalibration. Please follow the instructions in the user manual to reset the calibration… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/customer-service-robot-support.context-as-a-service
CaaS Benchmark Corpus v1
A diverse collection of synthetic enterprise documents for benchmarking context extraction and RAG systems.
Dataset Description
This dataset contains 16 representative enterprise documents spanning multiple formats and domains, designed to evaluate:
Structure-aware indexing - Can the system identify high-value vs. low-value content?
Time decay relevance - Does the system properly weight recent vs. old information?
Pragmatic truth detection - Can… See the full description on the dataset page: https://huggingface.co/datasets/imran-siddique/context-as-a-service.turkish-competition-authority-decisions
Turkish Competition Authority Decisions (Rekabet Kurulu Kararları), 1997–2026
The complete published decision history of the Turkish Competition Authority
(Rekabet Kurumu) — every Competition Board decision the regulator has made public,
in full text, with derived structural metadata.
10,367 decisions · 113,297 pages · 323 million characters · 29 years
Every decision carries its outcome, the articles of Law 4054 it turns on, the
panel that decided it (as stable pseudonymous ids… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-competition-authority-decisions.DoD-Instruction-8130-01-Installation-of-Geospatial-Information-And-Services
🗺️ DoD Installation Geospatial Information and Services Question-Answer Dataset
Source: DoD Instruction 8130.01
Source Effective Date: April 9, 2015
Change Incorporated: Change 3, effective August 4, 2020
Source Organization: Office of the Under Secretary of Defense for Acquisition and Sustainment
Source Ownership: United States Department of Defense
📋 Overview
Dataset Summary
The DoD Installation Geospatial Information and Services… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8130-01-Installation-of-Geospatial-Information-And-Services.coverture-103k-gender-history
🏛 COVERTURE: Institutional Gender History Corpus (103,270 Evidentiary Dossiers)
"Culture is not a neutral mirror of reality. It is a disciplinary machine that normalizes domination through humor, law, romance, and erasure."
The Coverture Corpus is a large-scale, evidentiary research dataset comprising 103,270 structured analytical dossiers documenting the institutional, legal, economic, domestic, and cultural technologies of patriarchal control over women from Antiquity to… See the full description on the dataset page: https://huggingface.co/datasets/Sergey23214/coverture-103k-gender-history.imam_albani_weaknfab_series_dataset
Imam al-Albani Weak & Fabricated Hadith Dataset (AR–EN–MY)
This dataset contains weak, rejected, or fabricated hadiths classified byImam Muhammad Nasir al-Din al-Albani, presented in Arabic, English, and Myanmar (Burmese). Translated with Gemini Pro 3.0.
Dataset Structure
Each row represents one hadith with a global unique ID and multilingual fields.
CSV Column Order
global_id – Unique sequential ID (primary key)
hadith_arabic_text – Original Arabic text… See the full description on the dataset page: https://huggingface.co/datasets/freococo/imam_albani_weaknfab_series_dataset.DOD-Enterprise-DevSecOps-Reference-Design-AWS-Managed-Services
DoD Enterprise DevSecOps AWS Managed Services Reference Design Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on DoD Enterprise DevSecOps Reference Design: AWS Managed Services (DoD IaC Baseline), Version 0.2, September 2021.
The source presents a draft Department of Defense reference design for implementing a DevSecOps software factory… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Enterprise-DevSecOps-Reference-Design-AWS-Managed-Services.customer-service-grocery-cashier
This Dialogue
Comprised of fictitious examples of dialogues between a customer at a grocery store and the cashier. Check out the example below:
"id": 1,
"description": "Price inquiry",
"dialogue": "Customer: Excuse me, could you tell me the price of the apples per pound? Cashier: Certainly! The price for the apples is $1.99 per pound."
How to Load Dialogues
Loading dialogues can be accomplished using the fun dialogues library or Hugging Face datasets library.… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/customer-service-grocery-cashier.italian-open-sft-chat-dataset
Italian Open SFT Chat Dataset
An Italian-first, model-neutral synthetic SFT and chat dataset for fine-tuning Italian-capable LLMs. It targets instruction tuning, Italian chat behavior, structured output generation, JSON/YAML/CSV format following, coding assistance, safety refusals, multi-turn dialogue and reasoning-style final answers. This v0.1.0 package does not include long-context QA records.
This dataset is intended for users searching for an Italian instruction tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/SerFabio89/italian-open-sft-chat-dataset.autonomous-cloud-gpu-slurm-serving-suite
⚡ Autonomous Cloud GPU Infrastructure, Slurm Orchestration & Distributed Serving Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous AI Supercomputing & LLM Serving Agents
⚡ Overview & Industry Problem
Operating massive AI supercomputers (thousands of NVIDIA H100/H200 and Blackwell GPUs) requires coordinating Slurm cluster schedules, topology-aware NVLink cliques, NCCL AllReduce rings, RoCE v2 lossless fabrics… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-cloud-gpu-slurm-serving-suite.HiCUPID
💖 HiCUPID Dataset
📌 Dataset Summary
We introduce 💖 HiCUPID, a benchmark designed to train and evaluate Large Language Models (LLMs) for personalized AI assistant applications.
Why HiCUPID?
Most open-source conversational datasets lack personalization, making it hard to develop AI assistants that adapt to users. HiCUPID fills this gap by providing:
✅ A tailored dataset with structured dialogues and QA pairs.
✅ An automated evaluation model (based… See the full description on the dataset page: https://huggingface.co/datasets/serenalyoko/HiCUPID.code-service-national
Code du service national, non-instruct (2025-07-11)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-service-national.Serendip-sft-sinhala
Serendip-SFT-Sinhala Dataset 🇱🇰
📊 Dataset Summary
Serendip-SFT-Sinhala is a large-scale Sinhala instruction-tuning dataset with 293,613 high-quality examples for supervised fine-tuning (SFT) of large language models.
Created to train SerendipLLM, a Sinhala language model designed to excel at instruction-following, question-answering, summarization, and text classification.
🌟 Highlights
🇱🇰 293,613 Sinhala examples (largest Sinhala SFT dataset)
📚 4 task… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/Serendip-sft-sinhala.code-impositions-biens-services
Code des impositions sur les biens et services, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-impositions-biens-services.Customer-service-tickets-qwen-qa
Customer Support Tickets QA (English) — Qwen SFT Dataset
This dataset is formatted for supervised fine-tuning (SFT) of Qwen-style chat models on customer support email tasks. source dataset: Tobi-Bueck/customer-support-tickets
It is designed for training models to read a customer ticket, understand its context, and generate an appropriate support response. Depending on the prompt design, the same data can also support auxiliary tasks such as queue prediction, priority prediction… See the full description on the dataset page: https://huggingface.co/datasets/W-L/Customer-service-tickets-qwen-qa.
