CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anonymousasdf /video2mentaltext10K<n<100K0 likes2k downloads5mo agoHugging Face02anonymous-video-benchmark /toc_bench TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models TOC-Bench is a diagnostic benchmark for evaluating whether Video Large Language Models maintain object identity, state, persistence, and temporal relations throughout a video. It focuses on object-centric phenomena including occlusion, disappearance, reappearance, repeated events, event order, temporal location, duration, conditional state, and relative movement. Anonymous-review notice. This… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-video-benchmark/toc_bench.textvideo-text-to-text1K<n<10K0 likes657 downloads2mo agoHugging Face03Anonymous1383 /ship-dataset ShipBench: A Drawing-Grounded VLM Benchmark for Ship Structural Reasoning ShipBench is a metadata-grounded vision-language benchmark on parametrically-generated ship structural drawings. Six commercial ship types × nine drawing-grounded sub-tasks × deterministic ground truth derived directly from the generator's input dictionary (no human annotation, no rule-citation labels). Quick reference Total candidates: 6{,}450 across 6 ship types (Tanker, VLCC, BULKC, CNTR, LNGC… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous1383/ship-dataset.imagevisual-question-answering10K<n<100K0 likes638 downloads5mo agoHugging Face04anonymous1926 /autocode-fresh-cf AutoCode-RL fresh-CF Executable training problems for AutoCode-RL: Reinforcement Learning for Code with Verifiable Synthetic Data. A frozen GPT-5.5 setter constructs harder and easier variants and verification packages; a separate GPT-OSS-20B solver learns from binary program-execution rewards. View Problems Description originals 226 Source Codeforces tasks with generated verification packages enhance 84 Harder generated variants simplify 63 Easier generated… See the full description on the dataset page: https://huggingface.co/datasets/anonymous1926/autocode-fresh-cf.tabulartext-generationn<1K0 likes545 downloads1d agoHugging Face05anonymous2222 /Sympatheia-18k Sympatheia-18k Sympatheia-18k is an emotion-aware spoken dialogue dataset for empathetic speech synthesis research. It contains 18,000 query–response pairs across 12 emotion categories, each accompanied by synthesized audio and text transcripts. Dataset Structure Subset Unique Queries Responses Description Emotional 8,400 train / 3,600 eval 8,400 train / 3,600 eval Emotional queries with emotionally-matched responses Neutral 350 train / 150 eval 4,200 train /… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2222/Sympatheia-18k.audioaudio-to-audio10K<n<100K0 likes477 downloads5mo agoHugging Face06anonymous4486 /repnoise_beavertailtext10K<n<100K0 likes346 downloads2y agoHugging Face07anonymous-8421 /VL-DocIRgated Abstract VL-DocIR is a page-level benchmark for vision-based long-document retrieval built from 29,641 documents rendered into 388,548 page images from Wikipedia, arXiv, PubMed, and SEC proxy statements. The benchmark contains 271,760 questions over 23 domains and six query types, covering single-page, multi-page, and cross-document evidence configurations. Questions are grounded to rendered pages and HTML element identifiers, then filtered with a cleaning pipeline that targets… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-8421/VL-DocIR.textvisual-document-retrieval1M<n<10M0 likes281 downloads5mo agoHugging Face08Anonymous-Team-HC-RAG /Multi-doc-2025 Dataset Card for Multi-Doc-2025 Dataset Summary Multi-Doc-2025 is a financial question-answering benchmark built from SEC Form 10-K annual reports of S&P 500 companies. It is designed to evaluate retrieval-augmented generation (RAG) and financial QA systems under three reasoning settings that are not jointly covered by existing financial QA benchmarks: cross-company reasoning, cross-year reasoning, and hybrid-modal reasoning over both text and tables. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-Team-HC-RAG/Multi-doc-2025.textquestion-answering1K<n<10K2 likes272 downloads4mo agoHugging Face09anonymous-devbench-2026 /devbenchtext1K<n<10K0 likes258 downloads5mo agoHugging Face10Anonymous2535k /agentmorph-bugs-v0.1 AgentMorph AgentMorph is a trajectory-level metamorphic testing benchmark for tool-using LLM agents. Instead of requiring a labeled correct answer for every task, AgentMorph mutates a task in a way that should preserve the user's intent, reruns the agent, and checks whether the original and mutated trajectories preserve a rule-specific invariant. This repository is the anonymous review artifact for the AgentMorph paper. It contains synthetic e-commerce agent trajectories… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous2535k/agentmorph-bugs-v0.1.textn<1K1 likes249 downloads2mo agoHugging Face11anonymous-nips2026 /Agent-ValueBench Agent-ValueBench Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions). This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts. Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.tabularquestion-answering1K<n<10K0 likes228 downloads5mo agoHugging Face12anonymousatom /pushupbench PushUpBench: Video Repetition Counting Benchmark Project Page | Paper | GitHub PushUpBench is a benchmark for evaluating vision-language models (VLMs) on their ability to count exercise repetitions in videos. It was introduced in the paper "PushupBench: Your VLM is not good at counting pushups". The dataset consists of 446 long-form clips (averaging 36.7s) designed to test temporal reasoning and repetition counting beyond simple pattern recognition. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymousatom/pushupbench.textvideo-classificationn<1K0 likes195 downloads5mo agoHugging Face13anonymous-submission-001 /anonymous-datasettabular100M<n<1B0 likes190 downloads7mo agoHugging Face14Agri-LLaVA-Anonymous /Agricultural_pests_and_diseases_instruction_tuning_datatext1K<n<10K2 likes160 downloads2y agoHugging Face15AnonymousNu /SciRec SciReC: Diagnostic Evaluation of Relational Reasoning in Multimodal Scientific Conversations with Adaptive Interaction This dataset contains multimodal question-answering examples grounded in textbook figures. Records in the figure-grounded configurations are filtered to include only examples whose referenced image files are present in this release. Configurations visual: 13791 figure-grounded visual questions with resolved images. knowledge: 13501 caption/text-grounded… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousNu/SciRec.imagevisual-question-answering10K<n<100K0 likes135 downloads5mo agoHugging Face16anonymousxxxy /resource AnchorSIPS 🧠 AnchorSIPS is a synthetic dataset and evaluation resource for evidence-supported psychosis-risk symptom measurement. It is workflow-aligned to an observable structured interview flow and includes transcript-linked evidence targets for benchmark evaluation. This release is intended for public anonymous-review access with research-only methodological use. Plain-Language Summary AnchorSIPS contains fully synthetic interview conversations. No real patient… See the full description on the dataset page: https://huggingface.co/datasets/anonymousxxxy/resource.text10K<n<100K0 likes131 downloads5mo agoHugging Face17Anonymousblind /agent-failure-dynamics AgentHazard: Process-Centric Benchmark for AI Agent Trajectory Analysis AgentHazard is the first benchmark designed specifically for process-level analysis of AI coding agent trajectories. Unlike existing benchmarks that evaluate only final outcomes (pass/fail), AgentHazard provides standardized edit-level annotations, hazard estimation protocols, and stopping-policy evaluation tasks with unified evaluation across 85,050 trajectories from 6+ agent families. Why… See the full description on the dataset page: https://huggingface.co/datasets/Anonymousblind/agent-failure-dynamics.documenttext-classification1K<n<10K0 likes126 downloads3mo agoHugging Face18anonymous-neurips-ED /CTSpinoPelvic1K CTSpinoPelvic1K A fused spine + pelvis 3D CT segmentation dataset built by patient-level crosswalk between three public sources: TCIA CT COLONOGRAPHY — DICOM CT volumes (prone + supine per patient) CTSpine1K (COLONOG subset) — VerSe-convention vertebral label masks CTPelvic1K dataset2 — sacrum + bilateral hip label masks Annotations are placed onto the TCIA CT volume with the highest bone coverage (HU > 200), separately per anatomy. For ~650 patients both annotations land on the… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips-ED/CTSpinoPelvic1K.tabularimage-segmentation1K<n<10K1 likes124 downloads4mo agoHugging Face19contributor-anonymous /Mol2Pro-Binder-DatasetBinder-Dataset as defined in "Generalise or Memorise? Benchmarking Ligand-Conditioned Protein Generation from Sequence-Only Data" Our dataset is curated from the following sources: BindingDB (Liu et al., 2025) Drug Target Commons (DTC) (Tanoli et al., 2018) AlphaFill (Hekkelman et al., 2023) BioLiP (Yang et al., 2012) Citation Anonymised for double-blind review text1M<n<10M0 likes117 downloads5mo agoHugging Face20anony-mouse123 /Instruction_recall_dataset CanaryBench-PII Frequency-aware canary injection benchmark for auditing memorization in finetuned language models, built on the AI4Privacy PII reconstruction task. Dataset Description This dataset is part of CanaryBench, a benchmark for evaluating memorization in finetuned language models across repetition tiers and privacy regimes. Frequency tiers: 1×, 10×, 50× PII types: EMAIL, PHONE Member canaries: 770 Reference canaries: 1000 Tasks: PII detection, secret… See the full description on the dataset page: https://huggingface.co/datasets/anony-mouse123/Instruction_recall_dataset.texttext-generation10K<n<100K0 likes107 downloads2mo agoHugging Face21Anonymousblind /lig-rebuttal-data LIG Rebuttal Data Anonymous data release for COLM 2026 paper: The Latent Intelligence Gap Tiers Tier Contents Size Sufficient for 1 Results Aggregate JSONs ~5 MB Verify every number in rebuttal 2 Matched Per-problem trajectories (K=32) ~700 MB Reproduce all baselines 3 Embeddings Last-layer hidden states ~5 GB Reproduce IWC, verifier Models Model Parameters Benchmarks Qwen2.5-7B-Instruct 7B GSM8K, GPQA, AIME24, AIME25… See the full description on the dataset page: https://huggingface.co/datasets/Anonymousblind/lig-rebuttal-data.tabularn<1K0 likes103 downloads4mo agoHugging Face22memgui-agent-anonymous /MemGUI-3K MemGUI-3K Anonymous Project Page | Anonymous Code | Model MemGUI-3K is a memory-intensive mobile GUI agent trajectory dataset for training and analyzing agents that proactively manage long-horizon context. It contains teacher rollouts from MemGUI-Agent using the ConAct Context-as-Action paradigm, where the agent emits both GUI actions and context actions for history folding and UI memory management. Code, data processing scripts, model training scripts, and evaluation tools are… See the full description on the dataset page: https://huggingface.co/datasets/memgui-agent-anonymous/MemGUI-3K.tabularimage-text-to-text1K<n<10K0 likes103 downloads1mo agoHugging Face23anonymous-release-username /OmniMemBench Contents data/ — 103 benchmark samples, each with 128K/256K/512K/1M token tiers Usage Use the download script in the code repo: python download_data.py Data Format Each data/{run_id}/benchmark_{tier}.json contains: character_profile: persona and conversation style multi_session_dialogues: multi-session conversation history with multimodal references QAs: evaluation questions with ground truth answers, evidence chains, and clues… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-release-username/OmniMemBench.textquestion-answering1K<n<10K0 likes97 downloads2mo agoHugging Face24anonymousatom /arxiv-metadatatext100K<n<1M0 likes96 downloads6mo agoHugging Face25anonymous80934 /HalluCompass HalluCompass A direction-aware diagnostic benchmark and protocol for vision-language model (VLM) hallucination evaluation NeurIPS 2026 Datasets & Benchmarks Track Unified release: 2,200 images across MS-COCO + AMBER + NoCaps + VizWiz, 10,000 queries, two-pass annotation (GPT-4o-mini 1st-pass + 5-human verification, Fleiss' κ = 0.72 inter-annotator agreement on a stratified 250-image validation subset, substantial agreement per Landis & Koch 1977). A balanced POPE-compatible… See the full description on the dataset page: https://huggingface.co/datasets/anonymous80934/HalluCompass.imagevisual-question-answering10K<n<100K0 likes92 downloads5mo agoHugging Face26anonymous-submission-RIG-bench /RIG-Bench RIG-bench Anonymous submission to the NeurIPS 2026 Evaluations & Datasets (E&D) Track. A benchmark for reasoning-driven image generation: given visual context (images + instruction + optional demonstration pairs), the model must produce the answer as a single image. 2,000 samples 4 task families × 11 subtasks ~1.4 GB Files RIG-bench/ ├── README.md ├── samples.jsonl # 2,000 records └── images/<sample_id>/ ├── input_<order>.<ext> ├──… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-RIG-bench/RIG-Bench.imageimage-to-image1K<n<10K0 likes88 downloads5mo agoHugging Face27anonymousresearch123 /deka-retrieval Loading from datasets import load_dataset corpus = load_dataset("anonymousresearch123/deka-retrieval", "corpus", split="train") queries = load_dataset("anonymousresearch123/deka-retrieval", "queries", split="train") labels = load_dataset("anonymousresearch123/deka-retrieval", "labels", split="train") `` texttext-retrieval10K<n<100K0 likes86 downloads4d agoHugging Face28anonymous-insightladder-2026 /insight-ladder-imo2024 Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission). Overview A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with: 4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.tabulartext-generationn<1K0 likes82 downloads5mo agoHugging Face29anonymousmypcbench /mypcbench-baselinetextn<1K0 likes80 downloads5mo agoHugging Face30Anonymous-G /GenixerForShikra-DatasetsPaper: Genixer: Empowering Multimodal Large Language Model as a Powerful Data Generator (ECCV 2024) Arxiv: https://arxiv.org/abs/2312.06731 Description: syn_lcs_filtered60.jsonl and syn_sbu_filtered60.jsonl are two synthetic datasets produced by our Genixer_S model for advancing grounding-based multimodal understanding. tabular100K<n<1M1 likes78 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.