CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SamuelChien821 /counselbench-100 CounselBench-100 CounselBench-100 v3.2.5 is a synthetic legal-work benchmark with 100 authored matters across ten practice workflows. Every task has a natural employee request, a 97-asset evidence room, twelve portfolio decisions, 5–9 supported actions, 3–7 evidence holds, and a distinct deep multi-provider MCP trajectory. The answer is not preclassified in the evidence. Each portfolio item requires an immutable identity join, an operative-authority and revision lookup, a… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/counselbench-100.documentquestion-answeringn<1K0 likes2.5k downloads26d agoHugging Face02SamuelChien821 /factorybench-100 FactoryBench-100 FactoryBench-100 is a 100-task benchmark for employee-grade manufacturing and ERP decisions. Each public prompt is a short, high-level employee request; it does not name the systems, files, API calls, answer schema, or execution order. The isolated SQLite world exposes documented Oracle Fusion Cloud 26a REST operations alongside Gmail v1, Drive v3, Sheets v4, and Slack Web API operations over synthetic state. Harbor runs the authoritative SQLite state and trace… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/factorybench-100.documentquestion-answeringn<1K0 likes1.9k downloads24d agoHugging Face03SamuelChien821 /salesbench-100 SalesBench-100 SalesBench-100 is a synthetic long-horizon sales-agent benchmark with 100 original workflows across Salesforce, HubSpot, Gong, and a seeded evidence room. Each task begins with a high-level employee request and has its own authored causal rule and provider transition. Identity, operating facts, authority, governed policy, live-system indexes, and exceptions are separated so no mounted business asset publishes a selected option or precomputed change. Every task has… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/salesbench-100.documenttext-generationn<1K1 likes1.8k downloads26d agoHugging Face04SamuelChien821 /blobfish-domainbench-24 Blobfish DomainBench-24 v3.3.2 DomainBench-24 v3.3.2 is the public release record for 24 realistic stateful agent tasks across six professional domains. Every task runs in a checked-in SQLite-backed MCP world with typed read/write tools and deterministic state, exact argument-aware trace, containment, persisted causal workpaper, stakeholder handoff, provider-native exact-record readback for every changed domain row, and persisted-result readback verification. The release… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/blobfish-domainbench-24.documentquestion-answeringn<1K1 likes612 downloads26d agoHugging Face05SamuelChien821 /hubbench HubBench 1.4.0 One Blobfish-authored, oracle-proven benchmark family per Harbor Hub professional-domain cluster. Every task is an employee decision worked over a dependent chain of evidence — never a lookup — against mock stateful tools over an isolated SQLite world. The agent reaches the world only through its public surfaces (MCP over streamable HTTP, a terminal tool CLI, a REST API, and a web console); a deterministic verifier (HubScore) grades the finished world from… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/hubbench.documentquestion-answeringn<1K0 likes545 downloads21d agoHugging Face06Samarth0710 /reviewarena ReviewArena ReviewArena accompanies the NeurIPS Evaluations & Datasets submission ReviewArena: A Large-Scale Cross-Conference Dataset and Benchmark for LLM Peer Review. This release is a large, multi-conference corpus of peer-reviewed papers + their reviews + author rebuttals + acceptance decisions, harvested from OpenReview and aligned with OCR'd full-text markdown of each paper PDF where available. 51,529 papers 196,099 reviews 558,785 OCR'd PDF pages (markdown inlined… See the full description on the dataset page: https://huggingface.co/datasets/Samarth0710/reviewarena.tabulartext-generation10K<n<100K2 likes516 downloads3mo agoHugging Face07Samarth0710 /reviewbench ReviewBench A large, multi-conference corpus of peer-reviewed papers + their reviews + author rebuttals + acceptance decisions, harvested from OpenReview and aligned with OCR'd full-text markdown of every paper. 51,529 papers 196,099 reviews 558,785 OCR'd PDF pages (markdown inlined per row) 7 conferences, 22 venue/year combinations, 2020 – 2026 from datasets import load_dataset ds = load_dataset("/reviewbench") print(ds) # DatasetDict({ # neurips: Dataset(num_rows=...)# iclr:… See the full description on the dataset page: https://huggingface.co/datasets/Samarth0710/reviewbench.tabulartext-generation10K<n<100K0 likes312 downloads5mo agoHugging Face08samaya-ai /FrontierFinancegated FrontierFinance: A benchmark for measuring the frontier intelligence of finance AI agents. arXiv report | website | grading code 1. Overview Investors use AI agents across their entire workflow — idea screening and discovery, company and market research, financial data collection and modeling, portfolio tracking, and catalyst monitoring. Measuring how well an AI system performs across this range is both important and hard: a benchmark must be broad enough to span… See the full description on the dataset page: https://huggingface.co/datasets/samaya-ai/FrontierFinance.texttext-generationn<1K11 likes230 downloads1mo agoHugging Face09Trelis /function_calling_v3_SAMPLE Trelis Function Calling Dataset - VERSION 3 - SAMPLE This is a SAMPLE of the v3 dataset available for purchase here. Features: Allows models to be fine-tuned for function-calling. The dataset is human generated and does not make use of Llama 2 or OpenAI! The dataset includes 66 training rows, 19 validation rows and 5 test rows (for manual evaluation). Based on eight functions: search_bing, search_arxiv, save_chat, read_json_file, list_files, get_current_weather, delete_file… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/function_calling_v3_SAMPLE.textquestion-answeringn<1K1 likes191 downloads3y agoHugging Face10samuelandaudreymedianetwork /that-backpacker-article-corpus That Backpacker Article Corpus This dataset contains a structured corpus of long-form travel articles published on ThatBackpacker.com, authored primarily by Audrey Bergner as part of the Samuel & Audrey Media Network. The corpus includes 323 article records covering destination guides, multi-day itineraries, hiking, food travel, cultural experiences, city guides, transportation, accommodations, and practical travel planning. It is intended for non-commercial research, retrieval… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/that-backpacker-article-corpus.texttext-generation100K<n<1M1 likes157 downloads4mo agoHugging Face11KamiKrafton /agentvidbench-sample AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above. Sample selection The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench-sample.textvideo-text-to-textn<1K0 likes129 downloads2mo agoHugging Face12neifuisan /Neuro-sama-QnAThis dataset was manually created, line by line, by my tiny hand! Why? Because I was just bored during my summer. textquestion-answeringn<1K50 likes125 downloads2y agoHugging Face13pinzhenchen /wmt26-mist-sample Update Log 22 June 2026 (latest) - we updated our data mix because some BELEBELE samples did not have the context. If you downloaded data before 22 June, please download the new version. 16 June 2026 - first version Summary The wmt26-mist-sample is a multilingual mix provided by the WMT26 MIST shared task organizers as a starting point for fine-tuning multilingual LLMs. It contains three types of tasks, to cover same-language and cross-lingual comprehension and… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/wmt26-mist-sample.textquestion-answering10K<n<100K1 likes99 downloads3mo agoHugging Face14Groundtruth-Data /groundtruth-hallucination-bench-sample Groundtruth Data Hallucination Benchmark Sample This public teaser contains 180 representative, source-backed examples from Groundtruth Data products. Groundtruth Data builds verified evaluation, remediation, and held-out validation datasets for AI models using authoritative source data. The commercial workflow is: Find where a model fails. Prove the failure with a larger verified evaluation. Provide targeted remediation/training data. Validate improvement on untouched held-out… See the full description on the dataset page: https://huggingface.co/datasets/Groundtruth-Data/groundtruth-hallucination-bench-sample.textquestion-answeringn<1K1 likes91 downloads2d agoHugging Face15akumalondon /Rail_Freight_Logistics_Company_Email_Archive_Sample Ukrainian Rail-Freight Correspondence Corpus (Sample) Real operational correspondence from a working freight forwarding business, and the documents attached to it — consignment notes, service acts, invoices, wagon manifests. Not scraped, not synthetic, and never published anywhere before. This is a de-identified sample released for evaluation. It is drawn from a larger private archive; see Full archive below. Published by Akuma London · akumalondon.com Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.tabulartext-generation1K<n<10K0 likes90 downloads13d agoHugging Face16samwell /synthea-ncd-instructions Synthea NCD Instructions Synthetic EHR-based instruction-tuning dataset for training LLMs to predict non-communicable disease (NCD) risk, specifically Type 2 Diabetes and Hypertension. Quick Start from datasets import load_dataset dataset = load_dataset("samwell/synthea-ncd-instructions") # View a sample print(dataset["train"][0]) Dataset Description This dataset contains instruction-tuning examples derived from synthetic patient records generated using… See the full description on the dataset page: https://huggingface.co/datasets/samwell/synthea-ncd-instructions.texttext-generation10K<n<100K0 likes84 downloads6mo agoHugging Face17agentvidbench /agentvidbench-sample AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above. Sample selection The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/agentvidbench/agentvidbench-sample.textvideo-text-to-textn<1K0 likes76 downloads2mo agoHugging Face18samuelandaudreymedianetwork /project-23-argentina-travel-archive 🇦🇷 Project 23 Argentina Travel Archive Dataset Description This dataset contains a structured archive of Argentina-focused travel, media reference, article, video transcript, and photography metadata records from the Samuel & Audrey Media Network. The archive is part of Project 23, a long-term effort to document Argentina’s 23 provinces through travel guides, videos, photography, regional logistics, cultural coverage, and public source records. The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/project-23-argentina-travel-archive.texttext-retrieval10K<n<100K1 likes74 downloads4mo agoHugging Face19samuelandaudreymedianetwork /picture-perfect-portfolios-article-corpus Picture Perfect Portfolios Article Corpus This dataset contains a structured corpus of long-form finance and investing articles published on PicturePerfectPortfolios.com. The corpus includes 448 article records covering portfolio construction, asset allocation, capital efficiency, return stacking, managed futures, trend following, risk management, ETFs, systematic investing, alternative strategies, and DIY investor education. It is intended for non-commercial research, retrieval… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/picture-perfect-portfolios-article-corpus.texttext-generation1K<n<10K1 likes73 downloads4mo agoHugging Face20samuelandaudreymedianetwork /che-argentina-travel-article-corpus Che Argentina Travel Article Corpus This dataset contains a structured corpus of long-form Argentina travel articles published on CheArgentinaTravel.com by the Samuel & Audrey Media Network. The corpus includes 88 article records covering Argentina travel guides, itineraries, cultural experiences, regional food, transportation, accommodations, local logistics, and destination planning. It includes coverage of major areas such as Buenos Aires and Patagonia, along with regional… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/che-argentina-travel-article-corpus.texttext-generationn<1K1 likes72 downloads4mo agoHugging Face21Fayer /BMGQ-MultiHop-Sample 🧩 BMGQ (Sample Release) – Bottom-up Multi-hop Question Generation Dataset A Sampled Subset of BMGQ: Complex, Retrieval-Resistant, Multi-hop Reasoning Questions 👥 Authors Bingsen Qiu, Zijian Liu, Xiao Liu, Bingjie Wang, Feier Zhang, Yixuan Qin, Chunyan Li, Haoshen Yang, Zeren Gao 📘 Dataset Summary BMGQ is a dataset of complex, hard-to-search, multi-hop reasoning questions automatically generated using our proposed framework: BMGQ: A Bottom-up… See the full description on the dataset page: https://huggingface.co/datasets/Fayer/BMGQ-MultiHop-Sample.textquestion-answeringn<1K0 likes71 downloads10mo agoHugging Face22samuelandaudreymedianetwork /samuel-and-audrey-youtube-transcripts-en Samuel & Audrey YouTube Transcripts EN Corpus, 2012–2026 This dataset contains the English transcript archive from the Samuel and Audrey - Travel and Food Videos YouTube channel. The corpus covers travel and food videos published between 2012 and 2026. It includes full transcript records, cue-level transcript segments, YouTube video identifiers, publication dates, titles, view counts captured at export time, tags, source URLs, transcript text, and subtitle-style payloads where… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-and-audrey-youtube-transcripts-en.texttext-generation1M<n<10M1 likes71 downloads4mo agoHugging Face23david9dragon9 /cvalues_samplesThis dataset contains samples of the cvalues english dataset used for training domain invariant reward models by few-shot generalization. Each just_sample contains the sample (10 examples) by itself, while the sampler files contain the sample repeated 1000 times for a total of 10000 examples. References: Cvalues: https://github.com/X-PLUG/CValues Cvalues english dataset: https://huggingface.co/datasets/david9dragon9/cvalues-english textquestion-answering10K<n<100K0 likes69 downloads2y agoHugging Face24samuelandaudreymedianetwork /samuel-y-audrey-youtube-transcripts-es-en Samuel y Audrey Bilingual YouTube Transcript Corpus ES/EN This dataset contains a structured bilingual transcript corpus from the Samuel y Audrey Spanish-language travel channel. The corpus includes 643 video records with Spanish and English transcript material, video-level metadata, subtitle-style text, and cleaned transcript fields. It is intended for non-commercial research, translation analysis, retrieval workflows, language study, and media archive organization. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-y-audrey-youtube-transcripts-es-en.texttranslation1K<n<10K1 likes67 downloads4mo agoHugging Face25samersaabjr /MIRA-MATH MIRA-Math MIRA-Math is a synthetic benchmark for minimal information requesting and mathematical reasoning. It evaluates a narrow diagnostic capability: when a mathematical problem is underdetermined from the solver's view, can a model identify the exact missing atomic fact, ask for it precisely, and then use it to compute the correct final answer? Each instance is generated from a complete latent mathematical state with a unique answer. The solver, called Agent A in the… See the full description on the dataset page: https://huggingface.co/datasets/samersaabjr/MIRA-MATH.textquestion-answering1K<n<10K0 likes67 downloads3mo agoHugging Face26samuelandaudreymedianetwork /academic-citations-and-media-references Academic Citations and Media References Dataset This dataset contains structured citation and reference records connected to the Samuel & Audrey Media Network. It includes normalized records for academic citations, research references, media mentions, tourism-sector references, awards, public profiles, podcast/interview references, and finance-media references connected to projects such as Nomadic Samuel, That Backpacker, Che Argentina Travel, Picture Perfect Portfolios, and the… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/academic-citations-and-media-references.texttext-retrievaln<1K2 likes66 downloads4mo agoHugging Face27samuelandaudreymedianetwork /nomadic-samuel-article-corpus Nomadic Samuel Article Corpus This dataset contains a structured corpus of long-form travel articles published on NomadicSamuel.com by the Samuel & Audrey Media Network. The corpus includes 422 article records covering global travel, destination guides, overland logistics, food, culture, road trips, itineraries, and practical travel planning. It is intended for non-commercial research, retrieval workflows, text analysis, archive search, travel writing study, and media organization.… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/nomadic-samuel-article-corpus.texttext-generationn<1K1 likes64 downloads4mo agoHugging Face28deepinquiry /verified-facts-sample-100 DeepInquiry Verified Facts (Sample-100) A 90-fact sample from the DeepInquiry verified-facts corpus. Every fact in this sample has been cross-checked against multiple structurally independent web sources, cited, dated, and confidence-scored before it entered the corpus. This is a preview sample. The full corpus (~942 approved facts as of Sept 2026, growing continuously) is available via the DeepInquiry API at deepinquiry.ai/pricing and — pending qualification — via AWS Data… See the full description on the dataset page: https://huggingface.co/datasets/deepinquiry/verified-facts-sample-100.tabularquestion-answeringn<1K0 likes62 downloads24d agoHugging Face29rhesis /Insurance-ChatBot-TestBench-Sample Insurance ChatBot TestBench Dataset (Sample) Dataset Description: The dataset presented here includes 80 example prompts from the Insurance ChatBot TestBench, a specialized test set developed to evaluate the performance of generative AI chatbots in the insurance industry. These prompts are used in the analysis described in the blog post "Gen AI Chatbots in the Insurance Industry: Are they Trustworthy?". The test bench assesses chatbot performance across three critical dimensions:… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Insurance-ChatBot-TestBench-Sample.textquestion-answeringn<1K0 likes61 downloads2y agoHugging Face30sammydman /KnowDoBench KnowDoBench Cannot, Should Not, Did Anyway: Benchmarking Metacognitive Control Failure in Frontier LLMs Samir Haq, MD, MS · Shehni Nadeem, MD — Michael E. DeBakey VA Medical Center · Baylor College of Medicine KnowDoBench is a multi-domain, expert-validated dataset for evaluating whether LLMs correctly answer or correctly refuse tasks that require recognizing and enforcing knowledge boundaries. Each case has deterministic ground truth: the model must either produce a correct… See the full description on the dataset page: https://huggingface.co/datasets/sammydman/KnowDoBench.tabulartext-classificationn<1K0 likes60 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.