datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dr-CiK
Dr-CiK: A Testbed for Foresight-Driven Agents
Dr-CiK is a benchmark for evaluating whether agents can retrieve
forecasting-relevant context from a noisy document corpus, filter out
distractors, distill the retrieved context into forecast-useful evidence, and
produce forecasts grounded in that evidence.
Real-world time-series forecasting often depends not only on historical
observations but also on external context that must be actively discovered
from heterogeneous, noisy… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/Dr-CiK.asr-ser-quechua-collao-embeddings
ASR-SER embeddings for Quechua Collao
This repository contains embeddings only. It does not contain raw audio.
These embeddings were extracted for ASR-to-SER transfer experiments on the Quechua Collao emotional speech corpus. The release is intended for reproducible feature-extraction experiments and downstream analysis.
Dataset contents
One PyTorch tensor per utterance stored as an embedding file under embeddings/
A sanitized metadata table describing utterance… See the full description on the dataset page: https://huggingface.co/datasets/QuechuaBase/asr-ser-quechua-collao-embeddings.drbench
DRBench: A Realistic Benchmark for Enterprise Deep Research
📄 Paper | 💻 GitHub | 💬 Discord
DRBench is the first of its kind benchmark designed to evaluate deep research agents on complex, open-ended enterprise deep research tasks. It tests an agent's ability to conduct multi-hop, insight-driven research across public and private data sources, just like a real enterprise analyst.
✨ Key Features
🔎 Real Deep Research Tasks: Not simple fact lookups. Tasks… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/drbench.AgentJudgeBench
AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling
A benchmark for systematically evaluating how reliably LLM judges assess
agentic tool-calling workflows across structured, dependency-driven tasks.
Why this benchmark?
AgentJudgeBench measures how reliably LLM judges assess agentic tool-calling outputs. It provides 3,808 benchmark records spanning six DAG topologies and three difficulty… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/AgentJudgeBench.pelican-svg-drawings
Pelican SVG: frontier model drawings, scored
139 SVGs produced by seven frontier models answering Simon Willison's prompt,
"generate an SVG of a pelican riding a bicycle", each with the score the
pelican_svg_env
OpenEnv environment gave it. Four of the seven are open weights and three are closed.
Simon has run that prompt against nearly every model release since early 2025, but the
results live as embedded images across 129 blog posts and scattered gists. This dataset
exists… See the full description on the dataset page: https://huggingface.co/datasets/sergiopaniego/pelican-svg-drawings.orc-bench
ORC-bench
Task 1: Topological Path Finding
Task 2: Topological Connectivity
Task 3: Linear Power Flow
Task 4: Contingency Analysis
Task 5: Power Grid ControlTask 6: Power Flow Optimization
Task 1: Topological Path Finding
Problem Formulation
This task assesses the spatial reasoning ability of the model by asking it to determine the shortest path between two specific buses in a given power grid state. The grid state… See the full description on the dataset page: https://huggingface.co/datasets/serval-uni-lu/orc-bench.bible-reference
Bible Reference Corpus
Thirteen aligned reference datasets for study of the biblical text: Greek and
Hebrew lexicons keyed to Strong's numbers, an interlinear word map, the critical
apparatus of eight Greek editions, cross-reference and topical indexes, and
geolocated places.
Published by SermonIndex.
Everything in this repository is public domain or CC BY 4.0. Sources with
share-alike terms are kept in a separate repository,
sermonindex/bible-reference-sa,
so that a share-alike… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-reference.ConflictGUI
ConflictGUI
ConflictGUI is a benchmark for evaluating conflict awareness in GUI agents.
Dataset Splits
Split
Feasible
Conflict1
Conflict2
Total
Calibration
564
300
300
1,164
Test
1,800
822
874
3,496
Sources
ConflictGUI is constructed from AMEX, AndroidControl, and AITZ. The conflict instructions and labels are newly annotated, while screenshots and original instructions remain subject to their respective upstream terms.… See the full description on the dataset page: https://huggingface.co/datasets/serein356/ConflictGUI.bible
The Bible in 1,004 Languages
14,497,397 verses across 1,253 translations in 1,004 languages, every verse
keyed to the same chapter-and-verse address so that any two languages can be
aligned by joining on book, chapter and verse.
The Bible is the most widely translated text in existence, and for several
hundred of the languages here it is the largest — sometimes the only —
substantial digitised text. That makes this corpus unusually useful for
low-resource machine translation… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible.PrivacyAlign
PrivacyAlign
PrivacyAlign is a human-annotated preference dataset for training and evaluating privacy-aligned tool-use agents. Each row pairs two candidate final actions from different models for the same agentic scenario, along with human preference labels and per-response privacy annotations (leaks and omissions).
The scenarios are synthetic. The user names, emails, memories, and tool trajectories are all generated, and no real user data is included.
Splits… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/PrivacyAlign.bible-reference-sa
Bible Reference Corpus — Share-Alike Tier
The part of the SermonIndex Bible reference corpus that derives from
CC BY-SA 4.0 sources, kept in its own repository so that the share-alike
obligation is explicit and does not spread to the permissively licensed
material.
The main corpus is
sermonindex/bible-reference
(CC BY 4.0). Join on strongs / key.
If you use this repository, your derivative work must also be share-alike.
If that does not suit you, use the CC BY 4.0 repository… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-reference-sa.context-as-a-service
CaaS Benchmark Corpus v1
A diverse collection of synthetic enterprise documents for benchmarking context extraction and RAG systems.
Dataset Description
This dataset contains 16 representative enterprise documents spanning multiple formats and domains, designed to evaluate:
Structure-aware indexing - Can the system identify high-value vs. low-value content?
Time decay relevance - Does the system properly weight recent vs. old information?
Pragmatic truth detection - Can… See the full description on the dataset page: https://huggingface.co/datasets/imran-siddique/context-as-a-service.repro-time-series-saliency-maps-explaining-models-across-multiple-domains-traces
Agent traces
Agent sessions published from a Trackio Logbook.
early-church-fathers
Early Church Fathers — Scripture Citation Index
68,240 passages from 349 Church Fathers, each keyed to the Bible verse it
comments on. Drawn from 20,253 distinct works and covering all 66 books.
This is a patristic catena in machine-readable form: given a verse, it returns
what the Fathers said about it. Nothing comparable exists as an open dataset —
the underlying translations are freely available, but the verse-level alignment
is the work, and that is what this releases.… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/early-church-fathers.uk-pods
uk-pods - speech datasets of Ukrainian podcasts.
Preparation
Clone the dataset repository and extract the content of clips.tar.gz archive.
git clone https://huggingface.co/datasets/taras-sereda/uk-pods
cd uk-pods && tar -zxvf clips.tar.gz
To use these manifests for training/inference with NeMo [1] modify audio_filepath to absolute locations of audio files extracted in previous step.
# data_root=<clonned_repo_dir> # /home/ubuntu/uk-pods
data_root=$(realpath .)
sed -i… See the full description on the dataset page: https://huggingface.co/datasets/taras-sereda/uk-pods.SERA-KimiK3-Django-SWEAgent-Cliff32k-T1
SERA Kimi-K3 Django SWE-Agent — Cliff-chunked T1 (first rollout)
572 training records built from 210 Kimi-K3 SWE-agent trajectories on
Django, split to fit a 32,768-token context with
CliffCompaction instead of being truncated.
Why chunked
A 100+ step agent rollout does not fit a 32k training window — 76% of the
source T1 trajectories exceed it. Truncating them throws away most of the
supervision, and trains the model on a context format it never sees at… See the full description on the dataset page: https://huggingface.co/datasets/thientrangngv/SERA-KimiK3-Django-SWEAgent-Cliff32k-T1.sermon-index
SermonIndex — Sermon Metadata and Scripture Index
Metadata for the sermon archive at SermonIndex.net,
including the mapping between scripture passages and the sermons that expound them.
Table
Rows
What it holds
sermon
64,106
Title, speaker, summary, description, media links, URLs
sermon_scripture
281,754
Scripture references, one row per sermon–passage pair
sermon_topic
115,526
Topic assignments, one row per sermon–topic pair
speaker
2,188
Preachers and authors… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/sermon-index.bible-parallel-english
Parallel Bible — English Translations and Ancient Versions
A verse-aligned parallel corpus of the Protestant Bible in seventeen English
translations, spanning 1599 to 2022, plus the Latin Vulgate and Syriac Peshitta
for the New Testament.
Looking for every language? This repository is a curated English set,
chosen for spread across translation families and small enough to load whole.
For the full corpus — 1,253 translations in 1,004 languages, 14.4M verses —
see… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-parallel-english.x402-service-index
x402 Service Index
A free, daily-refreshed snapshot of live x402 pay-per-call API services discovered across
the public Coinbase CDP bazaar and PayAI discovery feeds, each one live-probed with one unpaid
request per crawl so alive reflects a real answer, not a directory listing.
400 services tracked, 395 alive as of the latest crawl (generated_at field in index.json).
Fields: resource, method, network, price_usd, alive. network is a CAIP-2 chain id
(e.g. eip155:8453 = Base).… See the full description on the dataset page: https://huggingface.co/datasets/nimapro1381/x402-service-index.algerian-darija-customer-service-sample
Algerian Darija customer messages — stratified sample
500 spontaneous Algerian Darija messages, written by real customers, drawn from a
first-party corpus of 869,166 customer messages. Every message here is unique
after normalization, de-identified, and typed by a human — nothing elicited, translated, scraped or
generated.
Algerian Darija (ISO 639-3 arq) is spoken by around 45 million people and is one of the worst-covered
varieties in current language models. For scale: PADIC… See the full description on the dataset page: https://huggingface.co/datasets/dzcorpora/algerian-darija-customer-service-sample.massive_serve_dpr_wiki_contriever_ivfpqSERA-KimiK3-Django-SWEAgent-Cliff32k-T2
SERA Kimi-K3 Django SWE-Agent — Cliff-chunked T2 (second rollout)
227 training records built from 137 Kimi-K3 SWE-agent trajectories on
Django, split to fit a 32,768-token context with
CliffCompaction instead of being truncated.
Why chunked
A 100+ step agent rollout does not fit a 32k training window — 27% of the
source T2 trajectories exceed it. Truncating them throws away most of the
supervision, and trains the model on a context format it never sees at… See the full description on the dataset page: https://huggingface.co/datasets/thientrangngv/SERA-KimiK3-Django-SWEAgent-Cliff32k-T2.clickhouse-server-imagerepro-learning-fingerprints-for-medical-time-series-with-redundancy-constrained-info-traces
Agent traces
Agent sessions published from a Trackio Logbook.
CallAgentAI-Hinglish-Customer-Service
CallAgent AI: Hinglish Business Conversations Dataset
This dataset contains synthetic, high-quality "Hinglish" (Hindi + English code-switching) customer service interactions. It was generated by CallAgent AI (callagentai.in) — India's leading AI voice receptionist platform designed specifically for Indian SMBs.
Why this dataset exists
Global voice AI models often fail to capture the unique nuances of Indian business calls, which heavily rely on fluid language… See the full description on the dataset page: https://huggingface.co/datasets/Ghanashyaam/CallAgentAI-Hinglish-Customer-Service.blackwell-sm120-serving-matrix
Blackwell SM120 Serving Matrix
Measured serving results for LLMs on workstation and server Blackwell (sm_120, RTX PRO 6000, 96 GB), across vLLM, SGLang and llama.cpp, with FP8 and NVFP4 passes. Each row is one exact cell: a model, a runtime and version, a quantization, a topology and a load point, with the startup result, throughput or latency where measured, and a link to the raw artifact.
The point of the matrix is the cells that are known-broken as much as the ones that are… See the full description on the dataset page: https://huggingface.co/datasets/jahnclawdmonet/blackwell-sm120-serving-matrix.massive_serve_dpr_wiki_qwen3_0.6b_ivfpqallura-org__Mistral-Small-24b-Sertraline-0304-details
Dataset Card for Evaluation run of allura-org/Mistral-Small-24b-Sertraline-0304
Dataset automatically created during the evaluation run of model allura-org/Mistral-Small-24b-Sertraline-0304
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allura-org__Mistral-Small-24b-Sertraline-0304-details.articles-metadata
Psychology Articles Metadata (FR-EN Bilingual)
A bilingual (French / English) metadata catalog of clinical psychology articles published on psychologieetserenite.com, authored by Gildas Garrec (CBT psychopractitioner). Each article is paired across the two languages with canonical URLs, themes, keywords, word counts and timestamps.
The dataset is designed for:
Translation alignment research (FR↔EN parallel article metadata)
Multilingual text classification (psychology themes)… See the full description on the dataset page: https://huggingface.co/datasets/psychologie-et-serenite/articles-metadata.custom-service-ticket-zh-tw
Dataset Card:客服工單分類與結構化抽取(合成資料)
資料集摘要
繁體中文(台灣)電商客服工單合成資料集,用於六分類意圖分類 + 結構化欄位抽取任務。全部資料由 teacher LLM(Claude / GPT,經 OpenRouter 呼叫)依 schema-first structured output 生成,不含任何真實客服工單,無個資/版權疑慮。共 3,148 筆,切分為 train / validation / test 三份。
使用方式
from datasets import load_dataset
ds = load_dataset("albertkingdom/custom-service-ticket-zh-tw", data_files={
"train": "train.jsonl",
"validation": "validation.jsonl",
"test": "test.jsonl",
})… See the full description on the dataset page: https://huggingface.co/datasets/albertkingdom/custom-service-ticket-zh-tw.
