datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
drbench
DRBench: A Realistic Benchmark for Enterprise Deep Research
📄 Paper | 💻 GitHub | 💬 Discord
DRBench is the first of its kind benchmark designed to evaluate deep research agents on complex, open-ended enterprise deep research tasks. It tests an agent's ability to conduct multi-hop, insight-driven research across public and private data sources, just like a real enterprise analyst.
✨ Key Features
🔎 Real Deep Research Tasks: Not simple fact lookups. Tasks… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/drbench.AgentJudgeBench
AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling
A benchmark for systematically evaluating how reliably LLM judges assess
agentic tool-calling workflows across structured, dependency-driven tasks.
Why this benchmark?
AgentJudgeBench measures how reliably LLM judges assess agentic tool-calling outputs. It provides 3,808 benchmark records spanning six DAG topologies and three difficulty… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/AgentJudgeBench.orc-bench
ORC-bench
Task 1: Topological Path Finding
Task 2: Topological Connectivity
Task 3: Linear Power Flow
Task 4: Contingency Analysis
Task 5: Power Grid ControlTask 6: Power Flow Optimization
Task 1: Topological Path Finding
Problem Formulation
This task assesses the spatial reasoning ability of the model by asking it to determine the shortest path between two specific buses in a given power grid state. The grid state… See the full description on the dataset page: https://huggingface.co/datasets/serval-uni-lu/orc-bench.turkish-court-decisions-duplicate
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-court-decisions-duplicate.SPADE-customer-service-dialogue
SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection
Paper | Code
SPADE contains a repository of customer service line synthetic user dialogues with goals, augmented from MultiWOZ 2.1 using GPT-3.5 and Llama 70B.
The datasets are intended for training and evaluating machine generated text detectors in dialogue settings.
There are 15 English datasets generated using 5 different augmentation methods and 2 large language models… See the full description on the dataset page: https://huggingface.co/datasets/AngieYYF/SPADE-customer-service-dialogue.bible
The Bible in 1,004 Languages
14,497,397 verses across 1,253 translations in 1,004 languages, every verse
keyed to the same chapter-and-verse address so that any two languages can be
aligned by joining on book, chapter and verse.
The Bible is the most widely translated text in existence, and for several
hundred of the languages here it is the largest — sometimes the only —
substantial digitised text. That makes this corpus unusually useful for
low-resource machine translation… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible.ultrafeedback_binarized_serbian
Dataset Card for UltraFeedback Binarized Serbian
Dataset Description
This dataset is a Serbian-translated version of the UltraFeedback dataset, utilized for training Zephyr-7Β-β. The original dataset comprises 64k English-language prompts, each paired with four completions from various models. In this Serbian version, the prompts and completions have been translated into Serbian. The dataset creation process remains the same: selecting the completion with the highest… See the full description on the dataset page: https://huggingface.co/datasets/datatab/ultrafeedback_binarized_serbian.early-church-fathers
Early Church Fathers — Scripture Citation Index
68,240 passages from 349 Church Fathers, each keyed to the Bible verse it
comments on. Drawn from 20,253 distinct works and covering all 66 books.
This is a patristic catena in machine-readable form: given a verse, it returns
what the Fathers said about it. Nothing comparable exists as an open dataset —
the underlying translations are freely available, but the verse-level alignment
is the work, and that is what this releases.… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/early-church-fathers.bible-parallel-english
Parallel Bible — English Translations and Ancient Versions
A verse-aligned parallel corpus of the Protestant Bible in seventeen English
translations, spanning 1599 to 2022, plus the Latin Vulgate and Syriac Peshitta
for the New Testament.
Looking for every language? This repository is a curated English set,
chosen for spread across translation families and small enough to load whole.
For the full corpus — 1,253 translations in 1,004 languages, 14.4M verses —
see… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-parallel-english.SERA-KimiK3-Django-SWEAgent-Cliff32k-T1
SERA Kimi-K3 Django SWE-Agent — Cliff-chunked T1 (first rollout)
572 training records built from 210 Kimi-K3 SWE-agent trajectories on
Django, split to fit a 32,768-token context with
CliffCompaction instead of being truncated.
Why chunked
A 100+ step agent rollout does not fit a 32k training window — 76% of the
source T1 trajectories exceed it. Truncating them throws away most of the
supervision, and trains the model on a context format it never sees at… See the full description on the dataset page: https://huggingface.co/datasets/thientrangngv/SERA-KimiK3-Django-SWEAgent-Cliff32k-T1.algerian-darija-customer-service-sample
Algerian Darija customer messages — stratified sample
500 spontaneous Algerian Darija messages, written by real customers, drawn from a
first-party corpus of 869,166 customer messages. Every message here is unique
after normalization, de-identified, and typed by a human — nothing elicited, translated, scraped or
generated.
Algerian Darija (ISO 639-3 arq) is spoken by around 45 million people and is one of the worst-covered
varieties in current language models. For scale: PADIC… See the full description on the dataset page: https://huggingface.co/datasets/dzcorpora/algerian-darija-customer-service-sample.Frames-synthetic-customer-service-dialogue
Frames Synthetic Customer Service Dialogues
This contains a repository of customer service line synthetic user dialogues with goals, augmented from Frames using Qwen2.5-32B.
The datasets are intended for training and evaluating machine generated text detectors in dialogue settings.
Dataset Structure
The datasets are of parquet file format and contain the following columns:
Column
Description
dia_no
Unique ID for each dialogue. Dialogues with the same ID… See the full description on the dataset page: https://huggingface.co/datasets/AngieYYF/Frames-synthetic-customer-service-dialogue.autonomous-cloud-gpu-slurm-serving-suite
⚡ Autonomous Cloud GPU Infrastructure, Slurm Orchestration & Distributed Serving Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous AI Supercomputing & LLM Serving Agents
⚡ Overview & Industry Problem
Operating massive AI supercomputers (thousands of NVIDIA H100/H200 and Blackwell GPUs) requires coordinating Slurm cluster schedules, topology-aware NVLink cliques, NCCL AllReduce rings, RoCE v2 lossless fabrics… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-cloud-gpu-slurm-serving-suite.HiCUPID
💖 HiCUPID Dataset
📌 Dataset Summary
We introduce 💖 HiCUPID, a benchmark designed to train and evaluate Large Language Models (LLMs) for personalized AI assistant applications.
Why HiCUPID?
Most open-source conversational datasets lack personalization, making it hard to develop AI assistants that adapt to users. HiCUPID fills this gap by providing:
✅ A tailored dataset with structured dialogues and QA pairs.
✅ An automated evaluation model (based… See the full description on the dataset page: https://huggingface.co/datasets/serenalyoko/HiCUPID.SERA-KimiK3-Django-SWEAgent-Cliff32k-T2
SERA Kimi-K3 Django SWE-Agent — Cliff-chunked T2 (second rollout)
227 training records built from 137 Kimi-K3 SWE-agent trajectories on
Django, split to fit a 32,768-token context with
CliffCompaction instead of being truncated.
Why chunked
A 100+ step agent rollout does not fit a 32k training window — 27% of the
source T2 trajectories exceed it. Truncating them throws away most of the
supervision, and trains the model on a context format it never sees at… See the full description on the dataset page: https://huggingface.co/datasets/thientrangngv/SERA-KimiK3-Django-SWEAgent-Cliff32k-T2.CallAgentAI-Hinglish-Customer-Service
CallAgent AI: Hinglish Business Conversations Dataset
This dataset contains synthetic, high-quality "Hinglish" (Hindi + English code-switching) customer service interactions. It was generated by CallAgent AI (callagentai.in) — India's leading AI voice receptionist platform designed specifically for Indian SMBs.
Why this dataset exists
Global voice AI models often fail to capture the unique nuances of Indian business calls, which heavily rely on fluid language… See the full description on the dataset page: https://huggingface.co/datasets/Ghanashyaam/CallAgentAI-Hinglish-Customer-Service.code-service-national
Code du service national, non-instruct (2025-07-11)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-service-national.code-impositions-biens-services
Code des impositions sur les biens et services, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-impositions-biens-services.time-series-foundation-models-papers
Time Series Foundation Models Papers — FineSet
A research-paper dataset on Time Series Foundation Models Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Time Series Foundation Models Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/time-series-foundation-models-papers.Servant_Leadership_Practical
Servant Leadership — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Servant_Leadership_Practical.mcp-server-catalog
MCP Server Catalog
A comprehensive catalog of 38 Model Context Protocol (MCP) servers for AI agents, covering data access, agent infrastructure, business-to-agent interfaces, compliance, and more.
Overview
This dataset provides a structured catalog of MCP servers that give AI agents access to real-world data and capabilities. Each server follows the MCP standard and can be used with Claude, GPT, and other LLMs that support tool use.
Categories
Category… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/mcp-server-catalog.Servant_Leadership_Theory
Servant Leadership — Theory
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Servant_Leadership_Theory.
