datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
counselbench-100
CounselBench-100
CounselBench-100 v3.2.5 is a synthetic legal-work benchmark with 100
authored matters across ten practice workflows. Every task has a natural employee
request, a 97-asset evidence room, twelve portfolio decisions, 5–9 supported
actions, 3–7 evidence holds, and a distinct deep multi-provider MCP trajectory.
The answer is not preclassified in the evidence. Each portfolio item requires an
immutable identity join, an operative-authority and revision lookup, a… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/counselbench-100.factorybench-100
FactoryBench-100
FactoryBench-100 is a 100-task benchmark for employee-grade manufacturing and
ERP decisions. Each public prompt is a short, high-level employee request; it
does not name the systems, files, API calls, answer schema, or execution order.
The isolated SQLite world exposes documented Oracle Fusion Cloud 26a REST
operations alongside Gmail v1, Drive v3, Sheets v4, and Slack Web API operations
over synthetic state.
Harbor runs the authoritative SQLite state and trace… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/factorybench-100.salesbench-100
SalesBench-100
SalesBench-100 is a synthetic long-horizon sales-agent benchmark with 100 original workflows across Salesforce, HubSpot, Gong, and a seeded evidence room. Each task begins with a high-level employee request and has its own authored causal rule and provider transition. Identity, operating facts, authority, governed policy, live-system indexes, and exceptions are separated so no mounted business asset publishes a selected option or precomputed change. Every task has… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/salesbench-100.blobfish-domainbench-24
Blobfish DomainBench-24 v3.3.2
DomainBench-24 v3.3.2 is the public release record for 24 realistic stateful agent tasks across six professional domains. Every task runs in a checked-in SQLite-backed MCP world with typed read/write tools and deterministic state, exact argument-aware trace, containment, persisted causal workpaper, stakeholder handoff, provider-native exact-record readback for every changed domain row, and persisted-result readback verification.
The release… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/blobfish-domainbench-24.hubbench
HubBench 1.4.0
One Blobfish-authored, oracle-proven benchmark family per Harbor Hub professional-domain cluster. Every task is an employee decision worked over a dependent chain of evidence — never a lookup — against mock stateful tools over an isolated SQLite world. The agent reaches the world only through its public surfaces (MCP over streamable HTTP, a terminal tool CLI, a REST API, and a web console); a deterministic verifier (HubScore) grades the finished world from… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/hubbench.reviewarena
ReviewArena
ReviewArena accompanies the NeurIPS Evaluations & Datasets submission ReviewArena: A Large-Scale Cross-Conference Dataset and Benchmark for LLM Peer Review.
This release is a large, multi-conference corpus of peer-reviewed papers + their reviews + author rebuttals + acceptance decisions, harvested from OpenReview and aligned with OCR'd full-text markdown of each paper PDF where available.
51,529 papers
196,099 reviews
558,785 OCR'd PDF pages (markdown inlined… See the full description on the dataset page: https://huggingface.co/datasets/Samarth0710/reviewarena.reviewbench
ReviewBench
A large, multi-conference corpus of peer-reviewed papers + their reviews + author rebuttals + acceptance decisions, harvested from OpenReview and aligned with OCR'd full-text markdown of every paper.
51,529 papers
196,099 reviews
558,785 OCR'd PDF pages (markdown inlined per row)
7 conferences, 22 venue/year combinations, 2020 – 2026
from datasets import load_dataset
ds = load_dataset("/reviewbench")
print(ds)
# DatasetDict({
# neurips: Dataset(num_rows=...)# iclr:… See the full description on the dataset page: https://huggingface.co/datasets/Samarth0710/reviewbench.FrontierFinance
FrontierFinance: A benchmark for measuring the frontier intelligence of finance AI agents.
arXiv report | website | grading code
1. Overview
Investors use AI agents across their entire workflow — idea screening and discovery, company and market research, financial data collection and modeling, portfolio tracking, and catalyst monitoring. Measuring how well an AI system performs across this range is both important and hard: a benchmark must be broad enough to span… See the full description on the dataset page: https://huggingface.co/datasets/samaya-ai/FrontierFinance.function_calling_v3_SAMPLE
Trelis Function Calling Dataset - VERSION 3 - SAMPLE
This is a SAMPLE of the v3 dataset available for purchase here.
Features:
Allows models to be fine-tuned for function-calling.
The dataset is human generated and does not make use of Llama 2 or OpenAI!
The dataset includes 66 training rows, 19 validation rows and 5 test rows (for manual evaluation).
Based on eight functions: search_bing, search_arxiv, save_chat, read_json_file, list_files, get_current_weather, delete_file… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/function_calling_v3_SAMPLE.that-backpacker-article-corpus
That Backpacker Article Corpus
This dataset contains a structured corpus of long-form travel articles published on ThatBackpacker.com, authored primarily by Audrey Bergner as part of the Samuel & Audrey Media Network.
The corpus includes 323 article records covering destination guides, multi-day itineraries, hiking, food travel, cultural experiences, city guides, transportation, accommodations, and practical travel planning.
It is intended for non-commercial research, retrieval… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/that-backpacker-article-corpus.agentvidbench-sample
AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above.
Sample selection
The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench-sample.Neuro-sama-QnAThis dataset was manually created, line by line, by my tiny hand!
Why? Because I was just bored during my summer.
wmt26-mist-sample
Update Log
22 June 2026 (latest) - we updated our data mix because some BELEBELE samples did not have the context. If you downloaded data before 22 June, please download the new version.
16 June 2026 - first version
Summary
The wmt26-mist-sample is a multilingual mix provided by the WMT26 MIST shared task organizers as a starting point for fine-tuning multilingual LLMs. It contains three types of tasks, to cover same-language and cross-lingual comprehension and… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/wmt26-mist-sample.groundtruth-hallucination-bench-sample
Groundtruth Data Hallucination Benchmark Sample
This public teaser contains 180 representative, source-backed examples from Groundtruth Data products.
Groundtruth Data builds verified evaluation, remediation, and held-out validation datasets for AI models using authoritative source data. The commercial workflow is:
Find where a model fails.
Prove the failure with a larger verified evaluation.
Provide targeted remediation/training data.
Validate improvement on untouched held-out… See the full description on the dataset page: https://huggingface.co/datasets/Groundtruth-Data/groundtruth-hallucination-bench-sample.Rail_Freight_Logistics_Company_Email_Archive_Sample
Ukrainian Rail-Freight Correspondence Corpus (Sample)
Real operational correspondence from a working freight forwarding business, and the
documents attached to it — consignment notes, service acts, invoices, wagon
manifests. Not scraped, not synthetic, and never published anywhere before.
This is a de-identified sample released for evaluation. It is drawn from a larger
private archive; see Full archive below.
Published by Akuma London · akumalondon.com
Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.synthea-ncd-instructions
Synthea NCD Instructions
Synthetic EHR-based instruction-tuning dataset for training LLMs to predict non-communicable disease (NCD) risk, specifically Type 2 Diabetes and Hypertension.
Quick Start
from datasets import load_dataset
dataset = load_dataset("samwell/synthea-ncd-instructions")
# View a sample
print(dataset["train"][0])
Dataset Description
This dataset contains instruction-tuning examples derived from synthetic patient records generated using… See the full description on the dataset page: https://huggingface.co/datasets/samwell/synthea-ncd-instructions.agentvidbench-sample
AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above.
Sample selection
The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/agentvidbench/agentvidbench-sample.project-23-argentina-travel-archive
🇦🇷 Project 23 Argentina Travel Archive
Dataset Description
This dataset contains a structured archive of Argentina-focused travel, media reference, article, video transcript, and photography metadata records from the Samuel & Audrey Media Network.
The archive is part of Project 23, a long-term effort to document Argentina’s 23 provinces through travel guides, videos, photography, regional logistics, cultural coverage, and public source records. The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/project-23-argentina-travel-archive.picture-perfect-portfolios-article-corpus
Picture Perfect Portfolios Article Corpus
This dataset contains a structured corpus of long-form finance and investing articles published on PicturePerfectPortfolios.com.
The corpus includes 448 article records covering portfolio construction, asset allocation, capital efficiency, return stacking, managed futures, trend following, risk management, ETFs, systematic investing, alternative strategies, and DIY investor education.
It is intended for non-commercial research, retrieval… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/picture-perfect-portfolios-article-corpus.che-argentina-travel-article-corpus
Che Argentina Travel Article Corpus
This dataset contains a structured corpus of long-form Argentina travel articles published on CheArgentinaTravel.com by the Samuel & Audrey Media Network.
The corpus includes 88 article records covering Argentina travel guides, itineraries, cultural experiences, regional food, transportation, accommodations, local logistics, and destination planning. It includes coverage of major areas such as Buenos Aires and Patagonia, along with regional… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/che-argentina-travel-article-corpus.BMGQ-MultiHop-Sample
🧩 BMGQ (Sample Release) – Bottom-up Multi-hop Question Generation Dataset
A Sampled Subset of BMGQ: Complex, Retrieval-Resistant, Multi-hop Reasoning Questions
👥 Authors
Bingsen Qiu, Zijian Liu, Xiao Liu, Bingjie Wang, Feier Zhang, Yixuan Qin, Chunyan Li, Haoshen Yang, Zeren Gao
📘 Dataset Summary
BMGQ is a dataset of complex, hard-to-search, multi-hop reasoning questions automatically generated using our proposed framework:
BMGQ: A Bottom-up… See the full description on the dataset page: https://huggingface.co/datasets/Fayer/BMGQ-MultiHop-Sample.samuel-and-audrey-youtube-transcripts-en
Samuel & Audrey YouTube Transcripts EN Corpus, 2012–2026
This dataset contains the English transcript archive from the Samuel and Audrey - Travel and Food Videos YouTube channel.
The corpus covers travel and food videos published between 2012 and 2026. It includes full transcript records, cue-level transcript segments, YouTube video identifiers, publication dates, titles, view counts captured at export time, tags, source URLs, transcript text, and subtitle-style payloads where… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-and-audrey-youtube-transcripts-en.cvalues_samplesThis dataset contains samples of the cvalues english dataset used for training domain invariant reward models by few-shot generalization.
Each just_sample contains the sample (10 examples) by itself, while the sampler files contain the sample repeated 1000 times for a total of 10000 examples.
References:
Cvalues: https://github.com/X-PLUG/CValues
Cvalues english dataset: https://huggingface.co/datasets/david9dragon9/cvalues-english
samuel-y-audrey-youtube-transcripts-es-en
Samuel y Audrey Bilingual YouTube Transcript Corpus ES/EN
This dataset contains a structured bilingual transcript corpus from the Samuel y Audrey Spanish-language travel channel.
The corpus includes 643 video records with Spanish and English transcript material, video-level metadata, subtitle-style text, and cleaned transcript fields. It is intended for non-commercial research, translation analysis, retrieval workflows, language study, and media archive organization.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-y-audrey-youtube-transcripts-es-en.MIRA-MATH
MIRA-Math
MIRA-Math is a synthetic benchmark for minimal information requesting and mathematical reasoning. It evaluates a narrow diagnostic capability: when a mathematical problem is underdetermined from the solver's view, can a model identify the exact missing atomic fact, ask for it precisely, and then use it to compute the correct final answer?
Each instance is generated from a complete latent mathematical state with a unique answer. The solver, called Agent A in the… See the full description on the dataset page: https://huggingface.co/datasets/samersaabjr/MIRA-MATH.academic-citations-and-media-references
Academic Citations and Media References Dataset
This dataset contains structured citation and reference records connected to the Samuel & Audrey Media Network.
It includes normalized records for academic citations, research references, media mentions, tourism-sector references, awards, public profiles, podcast/interview references, and finance-media references connected to projects such as Nomadic Samuel, That Backpacker, Che Argentina Travel, Picture Perfect Portfolios, and the… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/academic-citations-and-media-references.nomadic-samuel-article-corpus
Nomadic Samuel Article Corpus
This dataset contains a structured corpus of long-form travel articles published on NomadicSamuel.com by the Samuel & Audrey Media Network.
The corpus includes 422 article records covering global travel, destination guides, overland logistics, food, culture, road trips, itineraries, and practical travel planning. It is intended for non-commercial research, retrieval workflows, text analysis, archive search, travel writing study, and media organization.… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/nomadic-samuel-article-corpus.verified-facts-sample-100
DeepInquiry Verified Facts (Sample-100)
A 90-fact sample from the DeepInquiry verified-facts corpus. Every fact in this sample has been cross-checked against multiple structurally independent web sources, cited, dated, and confidence-scored before it entered the corpus.
This is a preview sample. The full corpus (~942 approved facts as of Sept 2026, growing continuously) is available via the DeepInquiry API at deepinquiry.ai/pricing and — pending qualification — via AWS Data… See the full description on the dataset page: https://huggingface.co/datasets/deepinquiry/verified-facts-sample-100.Insurance-ChatBot-TestBench-Sample
Insurance ChatBot TestBench Dataset (Sample)
Dataset Description:
The dataset presented here includes 80 example prompts from the Insurance ChatBot TestBench, a specialized test set developed to evaluate the performance of generative AI chatbots in the insurance industry. These prompts are used in the analysis described in the blog post "Gen AI Chatbots in the Insurance Industry: Are they Trustworthy?". The test bench assesses chatbot performance across three critical dimensions:… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Insurance-ChatBot-TestBench-Sample.KnowDoBench
KnowDoBench
Cannot, Should Not, Did Anyway: Benchmarking Metacognitive Control Failure in Frontier LLMs
Samir Haq, MD, MS · Shehni Nadeem, MD — Michael E. DeBakey VA Medical Center · Baylor College of Medicine
KnowDoBench is a multi-domain, expert-validated dataset for evaluating whether LLMs correctly answer or correctly refuse tasks that require recognizing and enforcing knowledge boundaries.
Each case has deterministic ground truth: the model must either produce a correct… See the full description on the dataset page: https://huggingface.co/datasets/sammydman/KnowDoBench.
