CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Anonymous-Team-HC-RAG /Multi-doc-2025 Dataset Card for Multi-Doc-2025 Dataset Summary Multi-Doc-2025 is a financial question-answering benchmark built from SEC Form 10-K annual reports of S&P 500 companies. It is designed to evaluate retrieval-augmented generation (RAG) and financial QA systems under three reasoning settings that are not jointly covered by existing financial QA benchmarks: cross-company reasoning, cross-year reasoning, and hybrid-modal reasoning over both text and tables. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-Team-HC-RAG/Multi-doc-2025.textquestion-answering1K<n<10K2 likes287 downloads4mo agoHugging Face02anonymous-nips2026 /Agent-ValueBench Agent-ValueBench Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions). This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts. Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.tabularquestion-answering1K<n<10K0 likes205 downloads5mo agoHugging Face03memoreason-anonymous /MemoReason MemoReason MemoReason evaluates how entity familiarity affects document-grounded reasoning in language models. It pairs factual passages with fictitious variants that preserve task structure and specified reasoning operations. The benchmark contains 100 document templates, 12 questions per template, and 109,200 question-answer examples across ten settings. Questions cover extraction, arithmetic, temporal reasoning, and inference. Subsets Split Description… See the full description on the dataset page: https://huggingface.co/datasets/memoreason-anonymous/MemoReason.textquestion-answering100K<n<1M0 likes149 downloads4h agoHugging Face04AnonymousNu /SciRec SciReC: Diagnostic Evaluation of Relational Reasoning in Multimodal Scientific Conversations with Adaptive Interaction This dataset contains multimodal question-answering examples grounded in textbook figures. Records in the figure-grounded configurations are filtered to include only examples whose referenced image files are present in this release. Configurations visual: 13791 figure-grounded visual questions with resolved images. knowledge: 13501 caption/text-grounded… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousNu/SciRec.imagevisual-question-answering10K<n<100K0 likes128 downloads5mo agoHugging Face05anonymous-insightladder-2026 /insight-ladder-imo2024 Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission). Overview A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with: 4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.tabulartext-generationn<1K0 likes81 downloads5mo agoHugging Face06anonymous-release-username /OmniMemBench Contents data/ — 103 benchmark samples, each with 128K/256K/512K/1M token tiers Usage Use the download script in the code repo: python download_data.py Data Format Each data/{run_id}/benchmark_{tier}.json contains: character_profile: persona and conversation style multi_session_dialogues: multi-session conversation history with multimodal references QAs: evaluation questions with ground truth answers, evidence chains, and clues… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-release-username/OmniMemBench.textquestion-answering1K<n<10K0 likes71 downloads2mo agoHugging Face07anonymousaaai123 /DocHopQA_Dataset DocHopQA Dataset Paper: https://arxiv.org/abs/2508.15851 Overview We introduce DocHop-QA, a large-scale benchmark comprising 11,379 QA instances for multimodal, multi-document, multi-hop question answering. Constructed from publicly available scientific documents sourced from PubMed Central, DocHop-QA is domain-agnostic and incorporates diverse information formats, including textual passages, tables, and structural layout cues. Unlike existing datasets, DocHop-QA does not… See the full description on the dataset page: https://huggingface.co/datasets/anonymousaaai123/DocHopQA_Dataset.textquestion-answering10K<n<100K0 likes62 downloads9mo agoHugging Face08anonymous-Data-Preparation-Bench /Data-Prep-Bench Data-Prep-Bench Dataset Overview This dataset is a comprehensive resource built for Supervised Fine-Tuning (SFT) and evaluation of Large Language Models (LLMs), covering six domains: Finance, Medicine, Law, Mathematics, Science, and General. A key feature of this dataset is that we employed 12 different data generation methods (including Agent-based methods, DataFlow series, pure LLM-based generation, and a SKILL method) using multiple cutting-edge models (such as GPT-5… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-Data-Preparation-Bench/Data-Prep-Bench.texttext-generation1M<n<10M0 likes56 downloads5mo agoHugging Face09Anonymous-zxcvbnm /IO-Bench IO-Bench IO-Bench is a 155-example evaluation dataset for mathematical economics reasoning. Each record contains a standalone economics question, a reference answer, a machine-comparable answer field, symbolic answer metadata where applicable, and review status metadata. Unless otherwise noted, the dataset materials in this repository are licensed under the Creative Commons Attribution-NoDerivatives 4.0 International License (CC BY-ND 4.0). See LICENSE for details.… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-zxcvbnm/IO-Bench.tabularquestion-answeringn<1K0 likes47 downloads5mo agoHugging Face10cogbench-anonymous /cogbench-passages CogBench Passages A curated collection of 120 short academic passages across 8 subjects, used as the source text for the CogBench benchmark — an evaluation of whether LLMs can generate questions that satisfy specified Bloom's-Taxonomy cognitive levels under deterministic, code-checkable constraints. Anonymized for double-blind review (NeurIPS 2026 Evaluations & Datasets Track). Author and institutional metadata will be added after acceptance. Evaluative role This… See the full description on the dataset page: https://huggingface.co/datasets/cogbench-anonymous/cogbench-passages.textquestion-answeringn<1K0 likes38 downloads5mo agoHugging Face11anonymous-antiguessing-2026 /anti-guessing-olympiad Algorithmic Anti-Guessing Olympiad Math Benchmark A memorization-robust olympiad math benchmark constructed via a three-power-tier algorithmic anti-guessing pipeline. The pipeline operationalizes the per-problem verification gap (answer-only accuracy minus solution-correctness accuracy across a fixed target-model set) as both a construction criterion and an evaluation metric. Three released subsets This dataset ships three complementary subsets to support headline… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-antiguessing-2026/anti-guessing-olympiad.texttext-generation1K<n<10K1 likes35 downloads5mo agoHugging Face12anonymousneuripssubmission /BRUNCH BRUNCH: A Utility-Centric Benchmark for Evaluating Deep Research Dataset Description BRUNCH is a benchmark dataset for evaluating whether deep research agents can produce surveys that are genuinely useful for downstream reasoning. The motivation behind BRUNCH is that many existing evaluations of deep research systems focus on surface-level qualities, such as whether a report cites relevant papers, whether it looks well structured, or whether it matches a reference answer.… See the full description on the dataset page: https://huggingface.co/datasets/anonymousneuripssubmission/BRUNCH.texttext-classificationn<1K0 likes33 downloads5mo agoHugging Face13anonymous-penguin /DECADE DECADE: Dataset for Evolving Context And Dialogue Evaluation DECADE is a benchmark for evaluating long-term memory reasoning in personalized conversational AI. It simulates a decade (2016–2026) of user interactions across 500 QA instances, each paired with a personal conversation history of up to 1,047 sessions. Task Given a user's long conversation history (haystack) and a question posed from a future date, a system must retrieve the relevant sessions and synthesize an… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-penguin/DECADE.textquestion-answeringn<1K0 likes29 downloads5mo agoHugging Face14anonymousAIresearcher /horizonmath HorizonMath HorizonMath is a benchmark of research-level mathematical problems for measuring progress in reasoning toward mathematical discovery with automatic verification. Files data/problems_full.json data/problems_full.jsonl data/baselines.json data/baselines.jsonl croissant.json Loading from datasets import load_dataset problems = load_dataset( "anonymousAIresearcher/horizonmath", name="problems", split="train", ) baselines = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/anonymousAIresearcher/horizonmath.tabularquestion-answeringn<1K0 likes18 downloads5mo agoHugging Face15anonymousSub10 /mobiplant Dataset Card for MoBiPlant Dataset Summary MoBiPlant is a multiple-choice question-answering dataset curated by plant molecular biologists worldwide. It comprises two merged versions: Expert MoBiPlant: 565 expert-level questions authored by leading researchers. Synthetic MoBiPlant: 1,075 questions generated by large language models from papers in top plant science journals. Each example consists of a question about plant molecular biology, a set of answer options, and… See the full description on the dataset page: https://huggingface.co/datasets/anonymousSub10/mobiplant.tabularquestion-answering1K<n<10K0 likes11 downloads1y agoHugging Face16anonymous-submission-678 /backtrader-mcq-base-pool-all-strategies Backtrader MCQ Benchmark This dataset contains multiple-choice questions for evaluating whether a model can reason about trading-strategy behavior using the Backtrader backtesting framework. Each question provides a complete backtest configuration and asks for a single answer choice in the format <<< X >>>, where X is one of A, B, C, or D. The primary evaluation file used in the paper is: backtrader_mcq_balanced_30_all_strategies.jsonl The larger supporting pool is:… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-678/backtrader-mcq-base-pool-all-strategies.textquestion-answeringn<1K0 likes10 downloads5mo agoHugging Face17anonymous-789 /TSQA TSQA: Time-Sensitive Question Answering Benchmark TSQA is a benchmark designed to evaluate a model’s ability to handle time-aware factual knowledge. Unlike standard static QA datasets, TSQA tests whether models can identify facts whose correct answers change over time. Dataset Overview Name: TSQA (Time-Sensitive Question Answering) Years Covered: 2013–2024 Number of Questions: 10,063 Choices per Question: 4 (one correct, three distractors) Temporal Context: Each… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-789/TSQA.textmultiple-choice10K<n<100K1 likes9 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.