datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
drbench
DRBench: A Realistic Benchmark for Enterprise Deep Research
📄 Paper | 💻 GitHub | 💬 Discord
DRBench is the first of its kind benchmark designed to evaluate deep research agents on complex, open-ended enterprise deep research tasks. It tests an agent's ability to conduct multi-hop, insight-driven research across public and private data sources, just like a real enterprise analyst.
✨ Key Features
🔎 Real Deep Research Tasks: Not simple fact lookups. Tasks… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/drbench.AgentJudgeBench
AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling
A benchmark for systematically evaluating how reliably LLM judges assess
agentic tool-calling workflows across structured, dependency-driven tasks.
Why this benchmark?
AgentJudgeBench measures how reliably LLM judges assess agentic tool-calling outputs. It provides 3,808 benchmark records spanning six DAG topologies and three difficulty… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/AgentJudgeBench.Kubu-hai
Dataset Card for kubu-hai.model 🙅♂️🤖
^|D Look Ma, an instruction dataset that wasn't generated by GPTs!
Dataset Summary
kubu-hai is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better. No Robots was modelled after the instruction dataset described in OpenAI's InstructGPT paper, and is comprised mostly of… See the full description on the dataset page: https://huggingface.co/datasets/Seriki/Kubu-hai.orc-bench
ORC-bench
Task 1: Topological Path Finding
Task 2: Topological Connectivity
Task 3: Linear Power Flow
Task 4: Contingency Analysis
Task 5: Power Grid ControlTask 6: Power Flow Optimization
Task 1: Topological Path Finding
Problem Formulation
This task assesses the spatial reasoning ability of the model by asking it to determine the shortest path between two specific buses in a given power grid state. The grid state… See the full description on the dataset page: https://huggingface.co/datasets/serval-uni-lu/orc-bench.turkish-court-decisions-duplicate
Türk İçtihat Korpusu — 11.045.085 Mahkeme Kararı
Türkiye'nin kamuya açık mahkeme kararlarından derlenmiş, bilinen en büyük Türkçe
hukuk metni veri seti. 11.045.085 karar, 31.5 milyar karakter düz metin (5.50 GB Parquet),
1962'den 2026'ya. Yargıtay, Danıştay, Anayasa Mahkemesi ve UYAP Emsal üzerinden
yerel/istinaf mahkemeleri.
Kapsam
Kaynak
Karar sayısı
Yıl aralığı
Metin
Dosya
Yargıtay (yargitay)
9.820.145
1997–2026
19.5 milyar karakter
17
Danıştay… See the full description on the dataset page: https://huggingface.co/datasets/serdarsrts/turkish-court-decisions-duplicate.servicenow-tasks
ServiceNow Tasks
An evaluation dataset for web agents operating on ServiceNow instances. Each row contains a task goal, a configuration dict, and a standalone Python validation function that scores agent performance by querying the ServiceNow API.
Derived from the WorkArena L1 benchmark.
Dataset overview
330 samples across 33 task types grouped into 8 categories
Each task has 10 seeded variants
Validators are self-contained Python functions that call the… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/servicenow-tasks.SPADE-customer-service-dialogue
SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection
Paper | Code
SPADE contains a repository of customer service line synthetic user dialogues with goals, augmented from MultiWOZ 2.1 using GPT-3.5 and Llama 70B.
The datasets are intended for training and evaluating machine generated text detectors in dialogue settings.
There are 15 English datasets generated using 5 different augmentation methods and 2 large language models… See the full description on the dataset page: https://huggingface.co/datasets/AngieYYF/SPADE-customer-service-dialogue.retail-bank-servicing-alignment-sft
Retail Bank Servicing Alignment SFT
The training corpus for the Granite retail-bank servicing agent. It is the
released tool-use SFT corpus merged with a servicing-alignment continuation
curriculum that teaches multi-turn behaviours the base corpus does not: what to
do when the customer says "that one", when a policy question interrupts a
transfer, when the agent's own previous turn was wrong, and when the honest
answer is that the agent cannot see what it was asked about.
Every… See the full description on the dataset page: https://huggingface.co/datasets/spkc83/retail-bank-servicing-alignment-sft.SerendibLLM-PoemSong-Dataset
SerendibLLM Poem and Song Dataset
Sinhala poem and song lyrics dataset for fine-tuning Sinhala LLMs on creative generation.
Built as part of the Serendib LLM Honours project (UCLan 2025-2026).
19,184 instruction-response entries
Sources: kawmuthu.blogspot.com (poems) + lyrics-lk.com (songs)
6 instruction variants per entry (Sinhala + English prompts)
8 themes: love, nature, sadness, joy, spring, country, religion, general
brazilian-customer-service-conversations
Brazilian Customer Service Conversations
Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR).
De um like me apoie em manter esse dataset!
Descricao
Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de:
Chatbots de atendimento
Classificacao de intencao (intent classification)
Analise de sentimento em conversas
Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.tool-reasoning-sft-CODING-allenai-SERA-data-cleaned-rectified
SERA — Consolidated & Rectified
211,360 multi-turn SWE-agent coding trajectories from the SERA (Soft-Verified Efficient Repository Agents) project, consolidated from 4 source datasets into a single file with strict reasoning + tool-call format and validated FSM transitions.
Origin
Derived from Allen AI's Open Coding Agents release:
Source Dataset
Rows
Teacher
Scale
Rollout
allenai/Sera-4.5A-Full-T1
72,118
GLM-4.5-Air
full
T1
allenai/Sera-4.5A-Full-T2
66,337… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-CODING-allenai-SERA-data-cleaned-rectified.eva-bench
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
EVA-Bench is an end-to-end evaluation framework for conversational voice agents that orchestrates bot-to-bot audio conversations and scores them on both task accuracy and interaction experience.
About
No existing benchmark jointly addresses the two core evaluation challenges for voice agents: generating realistic simulated conversations, and measuring quality across the full scope of… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/eva-bench.bible
The Bible in 1,004 Languages
14,497,397 verses across 1,253 translations in 1,004 languages, every verse
keyed to the same chapter-and-verse address so that any two languages can be
aligned by joining on book, chapter and verse.
The Bible is the most widely translated text in existence, and for several
hundred of the languages here it is the largest — sometimes the only —
substantial digitised text. That makes this corpus unusually useful for
low-resource machine translation… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible.eva
A New Framework for Evaluating Voice Agents (EVA)
Most voice agent benchmarks evaluate either what the agent does or how it sounds. EVA evaluates both.
EVA is an open-source evaluation framework for conversational voice agents that scores complete, multi-turn spoken conversations across two fundamental dimensions:
EVA-A (Accuracy): Did the agent complete the task correctly and faithfully?
EVA-X (Experience): Was the interaction natural, concise, and appropriate for spoken… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/eva.physiotherapy-evidence-qa
🏥 Physiotherapy Evidence QA: A Bilingual Clinical Corpus
Physiotherapy Evidence QA is a large-scale, expert-curated bilingual dataset comprising 143,711 aligned question-answer pairs. It focuses on evidence-based physiotherapy, musculoskeletal rehabilitation, outcome measures, and clinical research methodology.
This corpus is designed to facilitate the development of Medical Large Language Models (Med-LLMs), Clinical Decision Support Systems (CDSS), and Cross-Lingual Information… See the full description on the dataset page: https://huggingface.co/datasets/serhanayberkkilic/physiotherapy-evidence-qa.ultrafeedback_binarized_serbian
Dataset Card for UltraFeedback Binarized Serbian
Dataset Description
This dataset is a Serbian-translated version of the UltraFeedback dataset, utilized for training Zephyr-7Β-β. The original dataset comprises 64k English-language prompts, each paired with four completions from various models. In this Serbian version, the prompts and completions have been translated into Serbian. The dataset creation process remains the same: selecting the completion with the highest… See the full description on the dataset page: https://huggingface.co/datasets/datatab/ultrafeedback_binarized_serbian.serendip-cpt-sinhala
Serendib LLM CPT Sinhala Corpus
A large-scale, deduplicated, quality-filtered Sinhala plain-text corpus built for
Continual Pre-Training (CPT) of large language models. This dataset was used to adapt
Meta-LLaMA-3-8B to the Sinhala language domain as part of the
Serendib LLM Honours Degree Research Project
at the University of Central Lancashire (UCLan), 2025–2026.
This is one of the largest openly published Sinhala NLP corpora available, containing
23,449,223 training documents… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/serendip-cpt-sinhala.customer-service-sft-50k
Customer Service SFT (50K)
50,000 ShareGPT-format customer service conversations across 8 industries and 18 issue types. Each conversation includes a system prompt establishing the agent's role, authority limits, and policy constraints — training models to operate within defined boundaries while resolving issues empathetically and effectively.
Motivation
Customer service is one of the highest-volume LLM deployment contexts. Models need to balance:
Empathy with… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/customer-service-sft-50k.SERA-GLM5.2-Django-SWEAgent-T1
SERA GLM-5.2 Django SWE-Agent — T1 (first-rollout trajectories)
167 software-engineering agent trajectories generated with the SERA SVG pipeline (paper).
Teacher: GLM-5.2 (temperature 0.6), reasoning traces preserved in <think> blocks
Harness: SWE-agent (str_replace_editor, bash, submit tools), 75-step cap, SWE-Bench Django container (django__django-7530, base commit f8fab6f9)
Stage: rollout one — vague bug prompt over 200 randomly sampled Django functions; kept submitted… See the full description on the dataset page: https://huggingface.co/datasets/thientrangngv/SERA-GLM5.2-Django-SWEAgent-T1.SERA-GLM5.2-Django-SWEAgent-Raw-T1
SERA GLM-5.2 Django SWE-Agent - RAW T1 (first rollout, thinking enabled)
258 raw, pre-postprocess first-rollout trajectories with native GLM-5.2 reasoning traces, from the SERA SVG pipeline (paper).
Released raw so you can choose your own filtering, verification threshold and reasoning-trace handling. Companion: SERA-GLM5.2-Django-SWEAgent-Raw-T2.
Schema
Mirrors allenai/Sera-*-T1/T2:
column
notes
messages
JSON string - apply json.loads(). Raw SWE-agent… See the full description on the dataset page: https://huggingface.co/datasets/thientrangngv/SERA-GLM5.2-Django-SWEAgent-Raw-T1.omnimcp_nextjs_server_actions_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_nextjs_server_actions_teaser.early-church-fathers
Early Church Fathers — Scripture Citation Index
68,240 passages from 349 Church Fathers, each keyed to the Bible verse it
comments on. Drawn from 20,253 distinct works and covering all 66 books.
This is a patristic catena in machine-readable form: given a verse, it returns
what the Fathers said about it. Nothing comparable exists as an open dataset —
the underlying translations are freely available, but the verse-level alignment
is the work, and that is what this releases.… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/early-church-fathers.SERA-KimiK3-Django-SWEAgent-Raw-T1
SERA Kimi-K3 Django SWE-Agent - RAW T1 (first rollout)
300 raw, pre-postprocess first-rollout agent trajectories generated with the
SERA SVG pipeline (paper),
using Kimi K3 as the teacher.
Released raw so you can choose your own filtering, verification threshold and reasoning-trace handling.
Companion: SERA-KimiK3-Django-SWEAgent-Raw-T2.
Schema
Mirrors allenai/Sera-*-T1/T2:
column
notes
messages
JSON string - apply json.loads(). Raw SWE-agent history:… See the full description on the dataset page: https://huggingface.co/datasets/thientrangngv/SERA-KimiK3-Django-SWEAgent-Raw-T1.SERA-KimiK3-Django-SWEAgent-Cliff32k-T1
SERA Kimi-K3 Django SWE-Agent — Cliff-chunked T1 (first rollout)
572 training records built from 210 Kimi-K3 SWE-agent trajectories on
Django, split to fit a 32,768-token context with
CliffCompaction instead of being truncated.
Why chunked
A 100+ step agent rollout does not fit a 32k training window — 76% of the
source T1 trajectories exceed it. Truncating them throws away most of the
supervision, and trains the model on a context format it never sees at… See the full description on the dataset page: https://huggingface.co/datasets/thientrangngv/SERA-KimiK3-Django-SWEAgent-Cliff32k-T1.Sera-4.5A-Full-T1-v3
laion/Sera-4.5A-Full-T1-v3
Subset of allenai/Sera-4.5A-Full-T1.
Size: 72,118 rows (full dataset: 72,118 rows).
Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages
field (as JSON string), instance_id, rollout_patch, func_name, func_path,
problem_statement, target_patch, docker_image. Adds a source field pointing
back to the parent dataset.
Each assistant message carries a native tool_calls array (OpenAI tool-calling format)
and a train: bool flag for… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3.DoD-Instruction-8130-01-Installation-of-Geospatial-Information-And-Services
🗺️ DoD Installation Geospatial Information and Services Question-Answer Dataset
Source: DoD Instruction 8130.01
Source Effective Date: April 9, 2015
Change Incorporated: Change 3, effective August 4, 2020
Source Organization: Office of the Under Secretary of Defense for Acquisition and Sustainment
Source Ownership: United States Department of Defense
📋 Overview
Dataset Summary
The DoD Installation Geospatial Information and Services… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8130-01-Installation-of-Geospatial-Information-And-Services.kubectl-mcp-server-tool-call-reasoning-6k
kubectl-mcp-server-tool-call-reasoning-6k
MCP tool-calling SFT 資料集,由 Agent Tools Fine-Tuning Platform 以「反向生成 + teacher solver 驗證」流程產生。
語言:繁體中文
工具(來自 MCP server):install_helm_chart, upgrade_helm_chart, uninstall_helm_chart, helm_list, helm_status, helm_history, helm_get_values, helm_get_manifest, helm_get_notes, helm_get_hooks, helm_get_all, helm_show_chart, helm_show_values, helm_show_readme, helm_show_crds, helm_show_all, helm_search_repo, helm_search_hub, helm_repo_list… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/kubectl-mcp-server-tool-call-reasoning-6k.bible-parallel-english
Parallel Bible — English Translations and Ancient Versions
A verse-aligned parallel corpus of the Protestant Bible in seventeen English
translations, spanning 1599 to 2022, plus the Latin Vulgate and Syriac Peshitta
for the New Testament.
Looking for every language? This repository is a curated English set,
chosen for spread across translation families and small enough to load whole.
For the full corpus — 1,253 translations in 1,004 languages, 14.4M verses —
see… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-parallel-english.SERA-KimiK3-Django-SWEAgent-Raw-T2
SERA Kimi-K3 Django SWE-Agent - RAW T2 (second rollout)
160 raw, pre-postprocess second-rollout agent trajectories generated with the
SERA SVG pipeline (paper),
using Kimi K3 as the teacher.
Each row is an independent attempt at the synthetic PR derived from a first rollout; target_patch
holds that first-rollout patch so soft verification can be recomputed at any r.
Companion: SERA-KimiK3-Django-SWEAgent-Raw-T1.
Schema
Mirrors allenai/Sera-*-T1/T2:
column… See the full description on the dataset page: https://huggingface.co/datasets/thientrangngv/SERA-KimiK3-Django-SWEAgent-Raw-T2.algerian-darija-customer-service-sample
Algerian Darija customer messages — stratified sample
500 spontaneous Algerian Darija messages, written by real customers, drawn from a
first-party corpus of 869,166 customer messages. Every message here is unique
after normalization, de-identified, and typed by a human — nothing elicited, translated, scraped or
generated.
Algerian Darija (ISO 639-3 arq) is spoken by around 45 million people and is one of the worst-covered
varieties in current language models. For scale: PADIC… See the full description on the dataset page: https://huggingface.co/datasets/dzcorpora/algerian-darija-customer-service-sample.
