CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Samarth0710 /reviewarena ReviewArena ReviewArena accompanies the NeurIPS Evaluations & Datasets submission ReviewArena: A Large-Scale Cross-Conference Dataset and Benchmark for LLM Peer Review. This release is a large, multi-conference corpus of peer-reviewed papers + their reviews + author rebuttals + acceptance decisions, harvested from OpenReview and aligned with OCR'd full-text markdown of each paper PDF where available. 51,529 papers 196,099 reviews 558,785 OCR'd PDF pages (markdown inlined… See the full description on the dataset page: https://huggingface.co/datasets/Samarth0710/reviewarena.tabulartext-generation10K<n<100K2 likes550 downloads3mo agoHugging Face02SamuelChien821 /hubbench HubBench 1.4.0 One Blobfish-authored, oracle-proven benchmark family per Harbor Hub professional-domain cluster. Every task is an employee decision worked over a dependent chain of evidence — never a lookup — against mock stateful tools over an isolated SQLite world. The agent reaches the world only through its public surfaces (MCP over streamable HTTP, a terminal tool CLI, a REST API, and a web console); a deterministic verifier (HubScore) grades the finished world from… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/hubbench.documentquestion-answeringn<1K0 likes538 downloads20d agoHugging Face03Samarth0710 /reviewbench ReviewBench A large, multi-conference corpus of peer-reviewed papers + their reviews + author rebuttals + acceptance decisions, harvested from OpenReview and aligned with OCR'd full-text markdown of every paper. 51,529 papers 196,099 reviews 558,785 OCR'd PDF pages (markdown inlined per row) 7 conferences, 22 venue/year combinations, 2020 – 2026 from datasets import load_dataset ds = load_dataset("/reviewbench") print(ds) # DatasetDict({ # neurips: Dataset(num_rows=...)# iclr:… See the full description on the dataset page: https://huggingface.co/datasets/Samarth0710/reviewbench.tabulartext-generation10K<n<100K0 likes314 downloads5mo agoHugging Face04akumalondon /Rail_Freight_Logistics_Company_Email_Archive_Sample Ukrainian Rail-Freight Correspondence Corpus (Sample) Real operational correspondence from a working freight forwarding business, and the documents attached to it — consignment notes, service acts, invoices, wagon manifests. Not scraped, not synthetic, and never published anywhere before. This is a de-identified sample released for evaluation. It is drawn from a larger private archive; see Full archive below. Published by Akuma London · akumalondon.com Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.tabulartext-generation1K<n<10K0 likes90 downloads12d agoHugging Face05sammydman /KnowDoBench KnowDoBench Cannot, Should Not, Did Anyway: Benchmarking Metacognitive Control Failure in Frontier LLMs Samir Haq, MD, MS · Shehni Nadeem, MD — Michael E. DeBakey VA Medical Center · Baylor College of Medicine KnowDoBench is a multi-domain, expert-validated dataset for evaluating whether LLMs correctly answer or correctly refuse tasks that require recognizing and enforcing knowledge boundaries. Each case has deterministic ground truth: the model must either produce a correct… See the full description on the dataset page: https://huggingface.co/datasets/sammydman/KnowDoBench.tabulartext-classificationn<1K0 likes63 downloads5mo agoHugging Face06deepinquiry /verified-facts-sample-100 DeepInquiry Verified Facts (Sample-100) A 90-fact sample from the DeepInquiry verified-facts corpus. Every fact in this sample has been cross-checked against multiple structurally independent web sources, cited, dated, and confidence-scored before it entered the corpus. This is a preview sample. The full corpus (~942 approved facts as of Sept 2026, growing continuously) is available via the DeepInquiry API at deepinquiry.ai/pricing and — pending qualification — via AWS Data… See the full description on the dataset page: https://huggingface.co/datasets/deepinquiry/verified-facts-sample-100.tabularquestion-answeringn<1K0 likes62 downloads23d agoHugging Face07deepinquiry /sample-90 DeepInquiry Verified Facts (Sample-90) A 90-fact sample from the DeepInquiry verified-facts corpus. Every fact in this sample has been cross-checked against multiple structurally independent web sources, cited, dated, and confidence-scored before it entered the corpus. This is a preview sample. The full corpus (~942 approved facts as of Sept 2026, growing continuously) is available via the DeepInquiry API at deepinquiry.ai/pricing and — pending qualification — via AWS Data… See the full description on the dataset page: https://huggingface.co/datasets/deepinquiry/sample-90.tabularquestion-answeringn<1K0 likes53 downloads17d agoHugging Face08eshangj /stackoverflow_q_and_a_sample Description GitHub repository: https://github.com/EshanJayasundara/Stackoverflow-Python-Q-and-A-Extractor. GitHub repository contains the automated workflow for extracting the question and answer pairs from Stackoverflow. This dataset contains the question-answer pairs extracted from Stackoverflow using Stack Exchange API v2.3 and used following endpoints, /answers/{ids} GET /questions GET From 2020 January 1 to Today 1. Dataset description, Contains only python… See the full description on the dataset page: https://huggingface.co/datasets/eshangj/stackoverflow_q_and_a_sample.tabularquestion-answering10K<n<100K1 likes47 downloads1y agoHugging Face09BRlkl /samantha-r01-recursive-reasoning-corpus Samantha R01 Recursive Reasoning Corpus Answer-only corpus for the first isolated Samantha silent-tick / recursive-latent-reasoning validation. It is normalized for pre_train_recursive_reasoning.py and intentionally contains no visible chain-of-thought or source rationale fields. The repository is private because it combines sources with mixed or unspecified redistribution terms. Access does not supersede any upstream license. Split policy train: 50,000… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/samantha-r01-recursive-reasoning-corpus.tabularquestion-answering10K<n<100K0 likes46 downloads1mo agoHugging Face10YMinglai /GSM-DC-Dataset-Sample GSM-DC Test Dataset This dataset contains the test set for GSM-DC (Grade School Math with Distractor Chains), a synthetic math reasoning dataset with controlled complexity. Dataset Details Total Problems: 6300 Operation Counts (OP): 16-22 (out-of-distribution test set) Problem Types: Graph-based mathematical reasoning problems Noise Levels: Light, Medium, Hard (distractor difficulty) Dataset Structure Each problem in all_problems.json contains: problem_text:… See the full description on the dataset page: https://huggingface.co/datasets/YMinglai/GSM-DC-Dataset-Sample.tabularquestion-answering1K<n<10K0 likes45 downloads9mo agoHugging Face11dFusionAILabs /sample-fusion-intelligence-traces Sample Fusion Intelligence Traces Structured AI reasoning traces from dFusion's Fusion Intelligence system. Each record captures a complete agentic workflow: a real user query on a domain-specific topic, the full message chain including system prompts, tool calls, search results, intermediate reasoning steps, and a final synthesized answer — along with human feedback. These are not synthetic benchmarks. They are traces from real queries submitted by real users on live financial… See the full description on the dataset page: https://huggingface.co/datasets/dFusionAILabs/sample-fusion-intelligence-traces.tabularquestion-answeringn<1K0 likes40 downloads6mo agoHugging Face12samanthadies /representational_stability Dataset Card for Representational Stability Fictional Data Dataset Summary The Representational Stability fictional dataset is made to supplement the Trilemma of Truth dataset (here). The Trilemma of Truth data contains three types of statements: Factually true statements Factually false statements Synthetic, neither-valued statements generated to mimic statements unseen during LLM training The Representational Stability fictional dataset adds new types of statements:… See the full description on the dataset page: https://huggingface.co/datasets/samanthadies/representational_stability.tabulartext-classification1K<n<10K1 likes37 downloads10mo agoHugging Face13HerrHruby /synthetic-science-v2-sample Synthetic Scientific Research Threads — v2 (sample) A synthetic continual-learning benchmark: each episode is a coherent sequence of short fictional scientific research documents about a single made-up entity, with per-document QA anchors. Later documents build on, revise, or supersede earlier ones. Designed to stress test-time / meta-learning approaches where a model must adapt to a stream of documents and answer questions grounded in what it has just seen. This is a sample… See the full description on the dataset page: https://huggingface.co/datasets/HerrHruby/synthetic-science-v2-sample.tabularquestion-answeringn<1K0 likes33 downloads3mo agoHugging Face14solsticestudioai /synthetic-enterprise-ops-pack-sample Solstice Synthetic Enterprise Operations Pack (Sample) A multi-system graph dataset for agent evaluation and RAG benchmarking. This dataset simulates the interconnected operations of a modern technology company, linking sales activities, engineering workflows, IT support, and internal communications. Built by Solstice AI Studio as a free sample of a larger commercial pack. 100% synthetic — no real company or employee data. What's in the box This dataset consists of 32… See the full description on the dataset page: https://huggingface.co/datasets/solsticestudioai/synthetic-enterprise-ops-pack-sample.tabulargraph-mln<1K0 likes30 downloads5mo agoHugging Face15samahadhoud /idea-first-code-later-cp Idea First, Code Later: CP Benchmark Benchmark dataset for the paper: "Idea First, Code Later: Disentangling Problem Solving from Code Generation in Evaluating LLMs for Competitive Programming" A curated benchmark of 83 competitive programming problems designed for evaluating LLMs on algorithmic problem-solving separately from code generation. Motivation We curate problems from seven contests that are not hosted on major public CP platforms (e.g., Codeforces, AtCoder).… See the full description on the dataset page: https://huggingface.co/datasets/samahadhoud/idea-first-code-later-cp.tabulartext-generationn<1K0 likes29 downloads8mo agoHugging Face16Jackrong /DeepSeek-v3.1-reasoner-Distilled-math-samples DeepSeek-V3.1 Distillation with NVIDIA Nemotron-Post-Training-Dataset-v2 (Math Subset) The release of DeepSeek-V3.1 has attracted wide attention in the AI community. Its significant improvements in reasoning ability provide a new opportunity to explore optimization of domain-specific models. To investigate the potential of this model in complex mathematical reasoning tasks, I selected the math subset from NVIDIA’s newly released Nemotron-Post-Training-Dataset-v2 as seed problems and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-v3.1-reasoner-Distilled-math-samples.tabularquestion-answeringn<1K1 likes28 downloads1y agoHugging Face17williamTLmiller /bva-decisions-structured-sample2019Present BVA Structured Decisions (2019–2025) Structured, issue-level records extracted from U.S. Board of Veterans' Appeals (BVA) decisions — each decision parsed into its issues, conditions, outcomes, citations, and reasoning, with per-document provenance and completeness flags. Built for training and evaluating legal-AI models on veterans' disability adjudication. This is a 2900-decision sample, balanced across seven years (2019–2025, decisions/year), so it's representative of the… See the full description on the dataset page: https://huggingface.co/datasets/williamTLmiller/bva-decisions-structured-sample2019Present.tabulartext-classification1K<n<10K0 likes25 downloads3mo agoHugging Face18samuelb98 /ds_final_p CourseMate Synthetic QA 10k Dataset of synthetic question–answer pairs grounded in short context chunks extracted from course PDFs (Intro to Data Science).Each row contains: question, answer, context + metadata (topic, doc_name, page, chunk_id, difficulty, generation info). Quick stats Rows: 10000 Topics (lectures): 11 Documents: 11 Distinct pages: 33 Distinct chunks (doc_name, chunk_id): 374 Question unique rate: 0.0655 Distributions Topics… See the full description on the dataset page: https://huggingface.co/datasets/samuelb98/ds_final_p.tabularquestion-answeringn<1K0 likes23 downloads8mo agoHugging Face19codelucas /ceo-quotes-verified-sample 🎙️ CEO Transcripts — Verified Executive Interviews The World's Largest Database of Verified C-Suite Transcripts 20,000+ Executives · 100,000+ Transcripts · 400,000+ Quotes · S&P 500 + NASDAQ + Global Leaders 🔥 What's In This Sample? This is a free evaluation sample from CEOInterviews.ai featuring 9 of the most market-moving voices in finance, tech, and policy. Executive Role Why They Matter Jensen Huang CEO, NVIDIA Every AI… See the full description on the dataset page: https://huggingface.co/datasets/codelucas/ceo-quotes-verified-sample.imagetext-generation1K<n<10K2 likes19 downloads10mo agoHugging Face20s-nlp /popqa-olmo-3-7b-instruct-temp0.9-samples99-logprobs OLMo-3-7B-Instruct self-consistency generations with logprobs on PopQA This dataset contains 99 self-consistency generations per question for the PopQA benchmark, produced with allenai/OLMo-3-7B-Instruct at temperature 0.9, together with token-level log probabilities for each completion. The file is intended for post-hoc analysis, self-consistency curves, adaptive stopping, and related aggregation methods. Source Base benchmark: PopQA Model: allenai/OLMo-3-7B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/popqa-olmo-3-7b-instruct-temp0.9-samples99-logprobs.tabularquestion-answering1K<n<10K0 likes18 downloads6mo agoHugging Face21aioil-ai /polish-court-rulings-sample Polish Court Rulings — Sample (korpus-pl) A production-grade, PII-hardened corpus of Polish court rulings — free evaluation sample. Full corpus: 505,611 rulings · ~3.18B tokens, licensed commercially. Contact: licensing@aioil.ai · aioil.ai What this is This sample contains 500 Polish court rulings drawn from the full korpus-pl dataset — a cleaned, deduplicated and PII-audited corpus of Polish jurisprudence built for AI training, evaluation and legal RAG… See the full description on the dataset page: https://huggingface.co/datasets/aioil-ai/polish-court-rulings-sample.tabulartext-generationn<1K0 likes17 downloads2mo agoHugging Face22Among-us-X /AmongUs-X-sample AmongUs-X — Representative Sample This is a representative ~385 MB sample of the full AmongUs-X dataset (18 GB, ~8,720 games), provided so reviewers can quickly inspect data quality without downloading the full corpus. Full dataset (8,720 games, 18 GB): https://huggingface.co/datasets/Among-us-X/AmongUs-X Full dataset DOI: https://doi.org/10.57967/hf/8698 Companion code: https://github.com/among-us-X/Among-Us-X What's in this sample section size contents… See the full description on the dataset page: https://huggingface.co/datasets/Among-us-X/AmongUs-X-sample.documenttext-generation10K<n<100K0 likes11 downloads5mo agoHugging Face23TwinDoc /math-qa-sample_addsub-kogated 데이터 출처 AI-HUB 에서 다운로드 받은 숫자연산 기계독해 데이터 를 사용해서 만든 데이터입니다. 경제 > Train > json 파일을 DataFrame 형태로 변형하여 전처리 및 답변 생성을 하였습니다. Raw 데이터의 answer 정보를 참고하여 답변을 생성하였습니다. 답변 생성 시 gpt-4o 를 활용했습니다. 저작권에 의해 본 데이터는 외부 반출 및 타인의 acess 승낙은 불허합니다. 데이터 설명 본 데이터의 Type 은 '가산/감산' 로만 구성되어 있습니다. 데이터 예시 ### context ### 2분기 순이익만 떼서 보면 증가세가 더욱 뚜렷하다. 신한금융은 9961억원, KB금융은 9911억원으로 1분기보다 각각 8.5%, 17.2% 늘었다. 하나금융은 6584억원, 우리금융은 6103억원으로 증가율은 각각 20.6%, 7.3%이다. 특히 KB금융은 분기 기준 사상 최대 실적을 올렸다. 수출 부진에 미·중… See the full description on the dataset page: https://huggingface.co/datasets/TwinDoc/math-qa-sample_addsub-ko.tabularquestion-answering1K<n<10K0 likes4 downloads2y agoHugging Face24TwinDoc /math-qa-sample_ext-kogated 데이터 출처 AI-HUB 에서 다운로드 받은 숫자연산 기계독해 데이터 를 사용해서 만든 데이터입니다. 경제 > Train > json 파일을 DataFrame 형태로 변형하여 전처리 및 답변 생성을 하였습니다. Raw 데이터의 answer 정보를 참고하여 답변을 생성하였습니다. 답변 생성 시 gpt-4o 를 활용했습니다. 저작권에 의해 본 데이터는 외부 반출 및 타인의 acess 승낙은 불허합니다. 데이터 설명 본 데이터의 Type 은 '단서추출' 로만 구성되어 있습니다. 데이터 예시 ### context ### 서울시가 민속 대명절인 추석을 맞아 내달 1일부터 20일까지 상생상회(매장), 네이버(온라인), 롯데백화점(매장)과 함께 팔도특산물로 구성된 명절 직거래장터를 진행한다고 31일 밝혔다. 팔도특산물을 구매할 수 있는 지역상생 거점공간인 '상생상회' 매장에서는 상주, 제주 등 14개 시도의 117개 농가에서 생산한 총… See the full description on the dataset page: https://huggingface.co/datasets/TwinDoc/math-qa-sample_ext-ko.tabularquestion-answering1K<n<10K0 likes3 downloads2y agoHugging Face25nlp-brin-id /QA_hukum_samplesgatedThis dataset sample was constructed by generating QA pairs from Law No. 17 of 2008 on Shipping (Undang-Undang No 17 Tahun 2008 Tentang Pelayaran). It is then manually verified by human validators and reviewers, resulting in QA pairs with fine-grained label categories. tabularquestion-answeringn<1K1 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.