CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anonymous-md /EDGAR_FILINGS_DATASET SFD: SEC Filings Dataset (v1) SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation. This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in: The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.tabulartext-generation1M<n<10M2 likes872 downloads5mo agoHugging Face02Anonymous8976 /M4Benchimagequestion-answering1K<n<10K3 likes307 downloads1y agoHugging Face03anonymous-submission-dataset-1 /CaReCoS CaReCoS A medical acoustic question-answering dataset for reasoning over mel spectrograms of heart, lung, and cough sounds. Each record provides a clinical question, the mel-spectrogram image of a recording, a ground-truth answer, and the recording's clinical metadata. The task is purely visual: a model receives the spectrogram image together with the question and must reason over the spectrogram to produce the answer. The raw audio is not used as model input - the original .wav… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-dataset-1/CaReCoS.audioquestion-answeringn<1K0 likes293 downloads15d agoHugging Face04Anonymous-Team-HC-RAG /Multi-doc-2025 Dataset Card for Multi-Doc-2025 Dataset Summary Multi-Doc-2025 is a financial question-answering benchmark built from SEC Form 10-K annual reports of S&P 500 companies. It is designed to evaluate retrieval-augmented generation (RAG) and financial QA systems under three reasoning settings that are not jointly covered by existing financial QA benchmarks: cross-company reasoning, cross-year reasoning, and hybrid-modal reasoning over both text and tables. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-Team-HC-RAG/Multi-doc-2025.textquestion-answering1K<n<10K2 likes287 downloads4mo agoHugging Face05anonymous-nips2026 /Agent-ValueBench Agent-ValueBench Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions). This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts. Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.tabularquestion-answering1K<n<10K0 likes205 downloads5mo agoHugging Face06jurisbenchomni-anonymous /jurisbenchomni JurisBenchOmni — Model Predictions This repository hosts the per-sample model predictions and the analysis-ready aggregated CSVs that accompany our paper JurisBenchOmni: A Judicial-Practice-Oriented Benchmark for Legal Omni-Modal Evaluation. What you get here lets you read off, slice, and re-aggregate the per-task / per-dimension / pipeline-level numbers we cite in the paper, without having to re-run inference. Companion code that produces the aggregated tables from the per-sample… See the full description on the dataset page: https://huggingface.co/datasets/jurisbenchomni-anonymous/jurisbenchomni.tabularquestion-answering1K<n<10K0 likes178 downloads5mo agoHugging Face07Time-HD-Anonymous /ST-Bench ST-Bench: Spatial-Temporal Reasoning Benchmark ST-Bench is a comprehensive benchmark dataset for training and evaluating spatial-temporal reasoning capabilities in large language models. It includes data with raw time series, text descriptions, and image visualizations. 📊 Dataset Overview Default Data (with time_series key) Subset Description Files Total Size ST-Align Alignment data for initial training 3 files ~3.2GB ST-Causal Causal reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Time-HD-Anonymous/ST-Bench.textquestion-answering10K<n<100K1 likes172 downloads9mo agoHugging Face08anonymousNeurIPS2026submission4281 /reviewarena ReviewArena ReviewArena accompanies the NeurIPS Evaluations & Datasets submission ReviewArena: A Large-Scale Cross-Conference Dataset and Benchmark for LLM Peer Review. This release is a large, multi-conference corpus of peer-reviewed papers + their reviews + author rebuttals + acceptance decisions, harvested from OpenReview and aligned with OCR'd full-text markdown of each paper PDF where available. 51,529 papers 196,099 reviews 558,785 OCR'd PDF pages (markdown inlined per… See the full description on the dataset page: https://huggingface.co/datasets/anonymousNeurIPS2026submission4281/reviewarena.tabulartext-generation10K<n<100K0 likes163 downloads5mo agoHugging Face09memoreason-anonymous /MemoReason MemoReason MemoReason evaluates how entity familiarity affects document-grounded reasoning in language models. It pairs factual passages with fictitious variants that preserve task structure and specified reasoning operations. The benchmark contains 100 document templates, 12 questions per template, and 109,200 question-answer examples across ten settings. Questions cover extraction, arithmetic, temporal reasoning, and inference. Subsets Split Description… See the full description on the dataset page: https://huggingface.co/datasets/memoreason-anonymous/MemoReason.textquestion-answering100K<n<1M0 likes149 downloads1h agoHugging Face10AnonymousNu /SciRec SciReC: Diagnostic Evaluation of Relational Reasoning in Multimodal Scientific Conversations with Adaptive Interaction This dataset contains multimodal question-answering examples grounded in textbook figures. Records in the figure-grounded configurations are filtered to include only examples whose referenced image files are present in this release. Configurations visual: 13791 figure-grounded visual questions with resolved images. knowledge: 13501 caption/text-grounded… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousNu/SciRec.imagevisual-question-answering10K<n<100K0 likes128 downloads5mo agoHugging Face11physics-r1-anonymous /physcorp-a PhysCorp-A — Audited Physics Training Corpus (6,432 records) The audit-clean training corpus released alongside the Physics-R1 paper (NeurIPS 2026 D&B Track submission). Subset of the 14,294-record pre-audit pool that survives the joint two-stage contamination audit against all six paper-canonical eval splits. Lineage 14,294 pre-audit pool (PhysCorp-pre-audit) aggregating nine source families. Construction audit (Stage-1 5-gram Jaccard ≥ 0.4 union Stage-2… See the full description on the dataset page: https://huggingface.co/datasets/physics-r1-anonymous/physcorp-a.textquestion-answering1K<n<10K0 likes126 downloads5mo agoHugging Face12anonymous-insightladder-2026 /insight-ladder-imo2024 Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission). Overview A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with: 4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.tabulartext-generationn<1K0 likes81 downloads5mo agoHugging Face13anonymous-release-username /OmniMemBench Contents data/ — 103 benchmark samples, each with 128K/256K/512K/1M token tiers Usage Use the download script in the code repo: python download_data.py Data Format Each data/{run_id}/benchmark_{tier}.json contains: character_profile: persona and conversation style multi_session_dialogues: multi-session conversation history with multimodal references QAs: evaluation questions with ground truth answers, evidence chains, and clues… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-release-username/OmniMemBench.textquestion-answering1K<n<10K0 likes71 downloads2mo agoHugging Face14Anonymous0515 /GRE30Kimagequestion-answering1K<n<10K1 likes68 downloads1y agoHugging Face15anonymousaaai123 /DocHopQA_Dataset DocHopQA Dataset Paper: https://arxiv.org/abs/2508.15851 Overview We introduce DocHop-QA, a large-scale benchmark comprising 11,379 QA instances for multimodal, multi-document, multi-hop question answering. Constructed from publicly available scientific documents sourced from PubMed Central, DocHop-QA is domain-agnostic and incorporates diverse information formats, including textual passages, tables, and structural layout cues. Unlike existing datasets, DocHop-QA does not… See the full description on the dataset page: https://huggingface.co/datasets/anonymousaaai123/DocHopQA_Dataset.textquestion-answering10K<n<100K0 likes62 downloads9mo agoHugging Face16anonymous111111111 /MUSE-Bench MUSE-Bench: Memory Utilization Evaluation Benchmark Official dataset for the paper "Beyond Memorization: Benchmarking Memory Utilization in Conversational LLM Agents." Anonymous release. This repository is an anonymized copy provided for double-blind peer review. Author and affiliation information is withheld until the review process concludes. Motivation LLM agents increasingly rely on persistent cross-session memory to support long-horizon and personalized… See the full description on the dataset page: https://huggingface.co/datasets/anonymous111111111/MUSE-Bench.textquestion-answeringn<1K0 likes58 downloads24d agoHugging Face17anonymous-Data-Preparation-Bench /Data-Prep-Bench Data-Prep-Bench Dataset Overview This dataset is a comprehensive resource built for Supervised Fine-Tuning (SFT) and evaluation of Large Language Models (LLMs), covering six domains: Finance, Medicine, Law, Mathematics, Science, and General. A key feature of this dataset is that we employed 12 different data generation methods (including Agent-based methods, DataFlow series, pure LLM-based generation, and a SKILL method) using multiple cutting-edge models (such as GPT-5… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-Data-Preparation-Bench/Data-Prep-Bench.texttext-generation1M<n<10M0 likes56 downloads5mo agoHugging Face18anonymous-aardvark /submission14717_fictionalqa_reformatted_triviaqa Reformatted TriviaQA for use alongside FictionalQA Repository: omitted Paper: omitted Dataset Description This dataset is a simple derived view of the validation data from the original TriviaQA dataset hosted by the original creators at hf.co/datasets/mandarjoshi/trivia_qa. To create this view, we extract the wikipedia articles associated with each question, as well as a simplified answer list, and then we create a few versions of the resulting data for use as… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_reformatted_triviaqa.texttext-generation10K<n<100K0 likes49 downloads1y agoHugging Face19Anonymous-zxcvbnm /IO-Bench IO-Bench IO-Bench is a 155-example evaluation dataset for mathematical economics reasoning. Each record contains a standalone economics question, a reference answer, a machine-comparable answer field, symbolic answer metadata where applicable, and review status metadata. Unless otherwise noted, the dataset materials in this repository are licensed under the Creative Commons Attribution-NoDerivatives 4.0 International License (CC BY-ND 4.0). See LICENSE for details.… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-zxcvbnm/IO-Bench.tabularquestion-answeringn<1K0 likes47 downloads5mo agoHugging Face20anonymous-somebody /CHIMERA CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning CHIMERA is a compact, high-difficulty synthetic reasoning dataset with long Chain-of-Thought (CoT) trajectories and broad scientific coverage. It is designed to support reasoning post-training for large language models. All examples are LLM-generated and automatically verified without human annotation. Total: 9,225 problems Subjects: 8 Topics: 1,179 Overview Recent reasoning advances rely heavily on… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-somebody/CHIMERA.texttext-generation1K<n<10K0 likes39 downloads5mo agoHugging Face21cogbench-anonymous /cogbench-passages CogBench Passages A curated collection of 120 short academic passages across 8 subjects, used as the source text for the CogBench benchmark — an evaluation of whether LLMs can generate questions that satisfy specified Bloom's-Taxonomy cognitive levels under deterministic, code-checkable constraints. Anonymized for double-blind review (NeurIPS 2026 Evaluations & Datasets Track). Author and institutional metadata will be added after acceptance. Evaluative role This… See the full description on the dataset page: https://huggingface.co/datasets/cogbench-anonymous/cogbench-passages.textquestion-answeringn<1K0 likes38 downloads5mo agoHugging Face22physics-r1-anonymous /physr1corp PhysR1Corp — Closed-form RL Training Pool (2,268 records) The closed-form (numeric / MCQ-gradable) RL training pool used by Physics-R1 (NeurIPS 2026 D&B Track submission). Carved out of PhysCorp-A (the audited 6,432-record corpus) by dropping open-ended questions, then decontaminated against MMMU-Pro Physics (−87 records) and against PhyX-mini + PhysUniBench-en at cos ≥ 0.85 (−78 additional records: 69 PhyX-mini near-duplicates and 9 PhysUniBench-en template duplicates). One… See the full description on the dataset page: https://huggingface.co/datasets/physics-r1-anonymous/physr1corp.imagequestion-answering1K<n<10K0 likes36 downloads4mo agoHugging Face23anonymous-antiguessing-2026 /anti-guessing-olympiad Algorithmic Anti-Guessing Olympiad Math Benchmark A memorization-robust olympiad math benchmark constructed via a three-power-tier algorithmic anti-guessing pipeline. The pipeline operationalizes the per-problem verification gap (answer-only accuracy minus solution-correctness accuracy across a fixed target-model set) as both a construction criterion and an evaluation metric. Three released subsets This dataset ships three complementary subsets to support headline… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-antiguessing-2026/anti-guessing-olympiad.texttext-generation1K<n<10K1 likes35 downloads5mo agoHugging Face24anonymous-sciconbench /sciconbench SciConBench SciConBench is a benchmark dataset for evaluating large language models (LLMs) on their ability to synthesize scientific conclusions. Dataset Summary Each record corresponds to one published Cochrane (CDSR) review article, identified by its DOI. For each review, the dataset provides: A clinical question derived from the review's Objective section A set of atomic facts decomposed from the review's conclusions Atomic fact pairs mapping each original conclusion… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-sciconbench/sciconbench.textquestion-answering1K<n<10K0 likes35 downloads5mo agoHugging Face25anonymous-submission-1 /SEABED SEABED: SouthEast Asian Benchmark for Evaluating Audio Reasoning SEABED (SouthEast Asian Benchmark for Evaluating Audio Reasoning) covers six audio-reasoning tasks: speech emotion recognition, speech affective interpretation, dialect and language identification, dialectal speech comprehension, prosodic ambiguity resolution, and long-form audio reasoning. This repository releases a stratified 10% sample (541 of 5,404 records) of its QA data for anonymous peer review, so reviewers… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-1/SEABED.textaudio-text-to-textn<1K0 likes35 downloads2mo agoHugging Face26anonymous1ads34 /STRIDE-Bench STRIDE Benchmark Dataset Description STRIDE Benchmark is an evaluation benchmark for assessing the behavioral realism of crowd trajectory generation and simulation models. Rather than comparing trajectories point-by-point, it evaluates whether generated trajectories exhibit behaviors consistent with a given scenario description — measuring trajectory-context consistency through decomposed behavioral questions. Dataset Summary The dataset is distributed as three… See the full description on the dataset page: https://huggingface.co/datasets/anonymous1ads34/STRIDE-Bench.tabulartext-generationn<1K0 likes34 downloads5mo agoHugging Face27anonymousneuripssubmission /BRUNCH BRUNCH: A Utility-Centric Benchmark for Evaluating Deep Research Dataset Description BRUNCH is a benchmark dataset for evaluating whether deep research agents can produce surveys that are genuinely useful for downstream reasoning. The motivation behind BRUNCH is that many existing evaluations of deep research systems focus on surface-level qualities, such as whether a report cites relevant papers, whether it looks well structured, or whether it matches a reference answer.… See the full description on the dataset page: https://huggingface.co/datasets/anonymousneuripssubmission/BRUNCH.texttext-classificationn<1K0 likes33 downloads5mo agoHugging Face28Anonymous-NeurIPS26-VeRA /VeRA VeRA: Reasoning Benchmarks as Executable Specifications NeurIPS 2026 - Evaluations & Datasets Track (Anonymous Submission) What is VeRA? Most reasoning benchmarks are static: the same problems are reused indefinitely, enabling memorization, format exploitation, and saturation. VeRA redefines a benchmark as an executable specification - a triple (template, generator, verifier) that can be sampled post-training to produce unlimited fresh instances with programmatically… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-NeurIPS26-VeRA/VeRA.textquestion-answering10K<n<100K0 likes33 downloads5mo agoHugging Face29anonymous-aardvark /submission14717_fictionalqa_training_splits Training splits view of the FictionalQA dataset The FictionalQA dataset Repository: omitted Paper: omitted Dataset Description This dataset is a derivative of the main dataset. Please see that dataset's README for a detailed description of the assets. The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for the associated paper. The primary purpose of this dataset repository is for transparency and to help… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_training_splits.texttext-generation100K<n<1M0 likes30 downloads1y agoHugging Face30anonymous-aardvark /submission14717_fictionalqa The FictionalQA dataset Repository: omitted Paper: omitted Dataset Summary The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the facts within the fictional documents. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa.tabulartext-generation10K<n<100K0 likes29 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.