CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anonymous-md /EDGAR_FILINGS_DATASET SFD: SEC Filings Dataset (v1) SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation. This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in: The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.tabulartext-generation1M<n<10M2 likes895 downloads5mo agoHugging Face02anonymous1926 /autocode-fresh-cf AutoCode-RL fresh-CF Executable training problems for AutoCode-RL: Reinforcement Learning for Code with Verifiable Synthetic Data. A frozen GPT-5.5 setter constructs harder and easier variants and verification packages; a separate GPT-OSS-20B solver learns from binary program-execution rewards. View Problems Description originals 226 Source Codeforces tasks with generated verification packages enhance 84 Harder generated variants simplify 63 Easier generated… See the full description on the dataset page: https://huggingface.co/datasets/anonymous1926/autocode-fresh-cf.tabulartext-generationn<1K0 likes232 downloads1d agoHugging Face03anonymous-nips2026 /Agent-ValueBench Agent-ValueBench Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions). This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts. Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.tabularquestion-answering1K<n<10K0 likes228 downloads5mo agoHugging Face04anonymousNeurIPS2026submission4281 /reviewarena ReviewArena ReviewArena accompanies the NeurIPS Evaluations & Datasets submission ReviewArena: A Large-Scale Cross-Conference Dataset and Benchmark for LLM Peer Review. This release is a large, multi-conference corpus of peer-reviewed papers + their reviews + author rebuttals + acceptance decisions, harvested from OpenReview and aligned with OCR'd full-text markdown of each paper PDF where available. 51,529 papers 196,099 reviews 558,785 OCR'd PDF pages (markdown inlined per… See the full description on the dataset page: https://huggingface.co/datasets/anonymousNeurIPS2026submission4281/reviewarena.tabulartext-generation10K<n<100K0 likes135 downloads5mo agoHugging Face05anonymousNeurIPS2026submission4281 /reviewarena-eval ReviewArena-Eval ReviewArena-Eval is the benchmark slice that accompanies the NeurIPS Evaluations & Datasets submission ReviewArena: A Large-Scale Cross-Conference Dataset and Benchmark for LLM Peer Review. It evaluates LLMs on the LLM-as-reviewer task: given a full paper and the structured review form used at that paper's venue-year, the model must produce overall rating, confidence, sub-scores, and free-text fields that are then compared against the actual human reviews on… See the full description on the dataset page: https://huggingface.co/datasets/anonymousNeurIPS2026submission4281/reviewarena-eval.tabulartext-generation1K<n<10K0 likes113 downloads5mo agoHugging Face06anonymousfileupload /divdata divdata Every simulation run behind our heterogeneous-LLM social-simulation work, consolidated into one dataset indexed by run_id and step. 1,296 simulation runs · 89,171 posts · 652,588 comments · 14,020,992 impressions · 748,285 agent activations · 32 model variants. Agents with distinct personas post and comment on a shared message board built on OASIS. Each agent is driven by one of ~10 different LLMs, so a single board mixes model families. The corpus supports asking which… See the full description on the dataset page: https://huggingface.co/datasets/anonymousfileupload/divdata.tabulartext-generation10M<n<100M0 likes113 downloads15d agoHugging Face07anonymous-insightladder-2026 /insight-ladder-imo2024 Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission). Overview A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with: 4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.tabulartext-generationn<1K0 likes82 downloads5mo agoHugging Face08anonymous-noname /econ_eval The Price of Progress: Benchmark-Level LLM Inference Cost Dataset Dataset Summary This dataset combines historical LLM inference prices with benchmark performance scores to construct the largest publicly available benchmark-level LLM price dataset we are aware of. It covers 100+ models across three major benchmarks (GPQA-Diamond, SWE-bench Verified, and AIME) over a two-year window from April 2024 to April 2026, with varying coverage per benchmark. The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-noname/econ_eval.tabulartext-generationn<1K0 likes70 downloads5mo agoHugging Face09Anonymous-07 /ChemSafetyBench Dataset Card for ChemSafetyBench Dataset Summary ChemSafetyBench is a regulatory-grounded benchmark dataset of 32,614 chemical substances for multi-label GHS (Globally Harmonized System) hazard prediction and LLM safety reliability evaluation. Unlike prior molecular benchmarks constructed by querying pharmaceutical databases, ChemSafetyBench is seeded from a curated hazardous materials registry, ensuring coverage of real-world industrial and safety-critical chemicals… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-07/ChemSafetyBench.tabulartext-classification10K<n<100K0 likes65 downloads5mo agoHugging Face10anonymous-2321 /bird-train-gemini3-flash Dataset Card for Think2SQL-SFT This dataset is a distilled Supervised Fine-Tuning (SFT) dataset designed to improve the reasoning capabilities of models in Text-to-SQL tasks. It contains high-quality reasoning traces and SQL queries generated by Gemini 3 Flash. Paper: Think2SQL: Blueprinting Reward Density and Advantage Scaling for Effective Text-To-SQL Reasoning Base Benchmark: BIRD-Train Dataset Description The dataset consists of 9,428 high-quality traces, of… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-2321/bird-train-gemini3-flash.tabulartext-generation1K<n<10K2 likes64 downloads8mo agoHugging Face11Anonymous-07 /SafeChem Dataset Card for SafeChem Dataset Summary SafeChem is a regulatory-grounded benchmark dataset of 32,211 chemical substances for multi-label GHS (Globally Harmonized System) hazard prediction and LLM safety reliability evaluation. Unlike prior molecular benchmarks constructed by querying pharmaceutical databases, SafeChem is seeded from a curated hazardous materials registry, ensuring coverage of real-world industrial and safety-critical chemicals including solvents… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-07/SafeChem.tabulartext-classification10K<n<100K0 likes55 downloads5mo agoHugging Face12anonymous1069 /IndicMMLU-Pro IndicMMLU Dataset This dataset contains the following languages: punjabi hindi urdu telugu gujrati kannada tamil marathi bengali UPLOAD Cite our work. This dataset is also described in IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding. @dataset{kj2024indicmmlupro, author = {Kj, Sankalp and Kumar, Ashutosh and Balaji, Laxmaan and Kotecha, Nikunj and Jain, Vinija and Chadha, Aman and Bhaduri, Sreyoshi}, title =… See the full description on the dataset page: https://huggingface.co/datasets/anonymous1069/IndicMMLU-Pro.tabulartext-generation100K<n<1M0 likes52 downloads1y agoHugging Face13Anonymous-zxcvbnm /IO-Bench IO-Bench IO-Bench is a 155-example evaluation dataset for mathematical economics reasoning. Each record contains a standalone economics question, a reference answer, a machine-comparable answer field, symbolic answer metadata where applicable, and review status metadata. Unless otherwise noted, the dataset materials in this repository are licensed under the Creative Commons Attribution-NoDerivatives 4.0 International License (CC BY-ND 4.0). See LICENSE for details.… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-zxcvbnm/IO-Bench.tabularquestion-answeringn<1K0 likes49 downloads5mo agoHugging Face14anonymousgiraffe /Strudel-Synth Strudel-Synth Strudel-Synth is a synthetic corpus of 21,174 (MIDI, Strudel) pairs for training and evaluating MIDI-to-Strudel decompilation, introduced in Decomposer: Learning to Decompile Symbolic Music to Programs. 🎹 Live demo: anonymousgiraffe/decomposer-demo 🤗 Model: anonymousgiraffe/Decomposer-Qwen3-8B Each pair consists of a Strudel program distilled from Claude-Opus-4.6 (conditioned on independently sampled musical and code-style seeds) and the MIDI produced by… See the full description on the dataset page: https://huggingface.co/datasets/anonymousgiraffe/Strudel-Synth.tabulartext-generation10K<n<100K0 likes49 downloads7d agoHugging Face15anonymous-authors /StereoTales Multilingual Story-Generation Bias Samples A multilingual evaluation dataset for probing demographic biases in LLM story generation. Each sample instructs a model to write a ~200-word story about a character carrying a given demographic attribute value (age, gender, ethnicity, religion, disability status, immigration status, ...) placed into a specific life scenario, with the goal of surfacing socio-economic and demographic biases in the generated narratives. Languages… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-authors/StereoTales.tabulartext-generation1M<n<10M0 likes46 downloads5mo agoHugging Face16anonymous1ads34 /STRIDE-Bench STRIDE Benchmark Dataset Description STRIDE Benchmark is an evaluation benchmark for assessing the behavioral realism of crowd trajectory generation and simulation models. Rather than comparing trajectories point-by-point, it evaluates whether generated trajectories exhibit behaviors consistent with a given scenario description — measuring trajectory-context consistency through decomposed behavioral questions. Dataset Summary The dataset is distributed as three… See the full description on the dataset page: https://huggingface.co/datasets/anonymous1ads34/STRIDE-Bench.tabulartext-generationn<1K0 likes33 downloads5mo agoHugging Face17anonymous-aardvark /submission14717_fictionalqa The FictionalQA dataset Repository: omitted Paper: omitted Dataset Summary The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the facts within the fictional documents. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa.tabulartext-generation10K<n<100K0 likes30 downloads1y agoHugging Face18anonymouslinuxarena /straj_linuxarena straj_linuxarena (public-env subset) Adversarial sabotage agentic SWE benchmark trajectories from the linuxarena project, run as part of the no-CoT time-horizons paper. The dataset viewer above shows the per-cell outcome records (precomputed_results.csv, 257 rows). Full Inspect .eval trajectories for the 13 public-environment tasks (~2.4 GB, 150 files) are stored under evals/ — see "Inspect trajectories" below. Per-cell schema Each row in precomputed_results.csv (and… See the full description on the dataset page: https://huggingface.co/datasets/anonymouslinuxarena/straj_linuxarena.tabulartext-generationn<1K1 likes27 downloads5mo agoHugging Face19anonymous12341952 /Narrative-Infilling Dataset Card for NarrativeInfilling Benchmark Dataset Details Dataset Description The Narrative Infilling Benchmark is a large-scale evaluation dataset for narrative infilling the task of generating a missing span within a narrative while maintaining consistency with both the preceding and following context. The benchmark spans four narrative domains and contains 9,142 instances with systematic variation in blank position and span length, enabling… See the full description on the dataset page: https://huggingface.co/datasets/anonymous12341952/Narrative-Infilling.tabulartext-generation1K<n<10K0 likes5 downloads5mo agoHugging Face20anonymous-acl26 /prompt-sensitivity-codegen Anonymous Prompt Sensitivity Dataset This package contains model generations and evaluation outcomes for an anonymized submission on prompt sensitivity in few-shot code generation. What is included prompt_sensitivity_dataset.jsonl: one row per generated sample prompt_sensitivity_dataset.csv: tabular view of the same rows prompt_sensitivity_dataset.parquet: columnar copy when parquet support is available prompt_variant_spec.json: machine-readable description of the prompt… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-acl26/prompt-sensitivity-codegen.tabulartext-generation100K<n<1M0 likes3 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.