CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pickapic-anonymous /pickapic_v1 Dataset Card for "pickapic_v1" More Information needed tabular100K<n<1M13 likes5.6k downloads3y agoHugging Face02acl-anonymous /CrowdEvaltabular100K<n<1M1 likes2.9k downloads2y agoHugging Face03anonymous-stgnn-aas /TSP_EXECUTION_RUNStabular1K<n<10K1 likes2.7k downloads23d agoHugging Face04anonymous-md /EDGAR_FILINGS_DATASET_2022_2026H1tabular1M<n<10M0 likes1.1k downloads3mo agoHugging Face05anonymous-md /EDGAR_FILINGS_DATASET SFD: SEC Filings Dataset (v1) SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation. This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in: The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.tabulartext-generation1M<n<10M2 likes861 downloads5mo agoHugging Face06anonymousllbench /llbench-dataset LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models via Human Preferences Anonymous release prepared for NeurIPS 2026 review. Please do not redistribute. LL-Bench is a large-scale, human-preference benchmark for evaluating low-level vision restoration in the era of large generative models (LGMs). It compares 10 LGMs with 16 specilist and 5 all-in-one models across 16 low-level vision tasks, paired with dense human annotations:pairwise… See the full description on the dataset page: https://huggingface.co/datasets/anonymousllbench/llbench-dataset.imageimage-to-image100K<n<1M0 likes853 downloads4mo agoHugging Face07anonymous-structured-agent /structured-file-audit-benchmark Paper Data Release This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them. Contents datasets/ Benchmark data and per-task manifests for the three paper-facing splits. datasets/sc_flat/data SC-Flat is derived from DaBench, augmented with a replayable perturbation injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.texttable-question-answering1 likes658 downloads2mo agoHugging Face08anonymous-neurips26-ljasd /LM-SimBench LM-SimBench Dataset Description LM-SimBench is a large-scale training-performance profiling dataset for large language models. The dataset is collected from training runs based on the MindSpeed-LLM framework and the Ascend NPU development stack, covering multiple model families, context lengths, and distributed parallel configurations. Each model is sampled under feasible combinations of data parallelism (DP), tensor parallelism (TP), pipeline parallelism (PP), context… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips26-ljasd/LM-SimBench.tabulartabular-regressionn<1K0 likes498 downloads5mo agoHugging Face09Time-HD-Anonymous /High_Dimensional_Time_Series Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Time-HD-Anonymous/High_Dimensional_Time_Series.tabulartime-series-forecasting100K<n<1M3 likes490 downloads1y agoHugging Face10anonymous-md /EDGAR_FILINGS_DATASET_2016_2021tabular1M<n<10M0 likes455 downloads3mo agoHugging Face11habit-anonymous /HABIT HABIT: Human-Aware Behavior and Interaction Training Dataset for Robot Manipulation ⚠️ Anonymous release. Authors and institutional information are intentionally withheld. This dataset card will be updated when these details become available. TL;DR: HABIT is a large-scale robot demonstration dataset for human-present environments, designed to teach robot policies human-aware behaviors. Keywords: Robot Manipulation Dataset, Human-Robot Interaction, Vision-Language-Action… See the full description on the dataset page: https://huggingface.co/datasets/habit-anonymous/HABIT.tabularrobotics1M<n<10M0 likes397 downloads5mo agoHugging Face12anonymous-nsc-author /Neapolitan-Spoken-Corpus Neapolitan Spoken Corpus (NSC) A corpus of read Neapolitan speech for ASR evaluation, with a validated Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters, metric implementations, per-clip results, and error annotations. This release supersedes the earlier 141-clip single-speaker version of this repository. The earlier release corresponds to Speaker S1 of the present corpus; the old audioData/ and transcripts.csv are replaced by data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.audioautomatic-speech-recognitionn<1K4 likes320 downloads3mo agoHugging Face13anonymous-egomonth /egomonth-dataset EgoMonth Dataset Overview EgoMonth is a month-level egocentric video question-answering benchmark for evaluating long-term spatiotemporal memory in multimodal large language models. The dataset focuses on daily-life first-person videos and QA tasks that require temporal indexing, spatial grounding, multi-video reasoning, and long-horizon memory. This repository provides QA metadata, structured annotations, representative anonymized sample videos, and baseline… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-egomonth/egomonth-dataset.tabularvisual-question-answering1K<n<10K2 likes299 downloads2mo agoHugging Face14TechWolf /anonymous-working-histories Structured Anonymous Career Paths extracted from Resumes Dataset Summary This dataset contains 2164 anonymous career paths across 24 differend industries. Each work experience is tagger with their corresponding ESCO occupation (ESCO v1.1.1). Languages We use the English version of ESCO. All resume data is in English as well. Dataset Structure Each working history contains up to 17 experiences. They appear in order, and each experience has a title… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/anonymous-working-histories.tabulartext-classification1K<n<10K7 likes262 downloads2y agoHugging Face15anonymous-nips2026 /Agent-ValueBench Agent-ValueBench Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions). This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts. Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.tabularquestion-answering1K<n<10K0 likes258 downloads5mo agoHugging Face16anonymous1926 /autocode-fresh-cf AutoCode-RL fresh-CF Executable training problems for AutoCode-RL: Reinforcement Learning for Code with Verifiable Synthetic Data. A frozen GPT-5.5 setter constructs harder and easier variants and verification packages; a separate GPT-OSS-20B solver learns from binary program-execution rewards. View Problems Description originals 226 Source Codeforces tasks with generated verification packages enhance 84 Harder generated variants simplify 63 Easier generated… See the full description on the dataset page: https://huggingface.co/datasets/anonymous1926/autocode-fresh-cf.tabulartext-generationn<1K0 likes232 downloads1d agoHugging Face17AnonymousMouse404 /pnp_20260904_063146_langThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousMouse404/pnp_20260904_063146_lang.tabularrobotics1K<n<10K0 likes229 downloads17d agoHugging Face18anonymousxxx /MoleculeCLA Overview We present MoleculeCLA: a large-scale dataset consisting of approximately 140,000 small molecules derived from computational ligand-target binding analysis, providing nine properties that cover chemical, physical, and biological aspects. Aspect Glide Property (Abbreviation) Description Molecular Characteristics Chemical glide_lipo (lipo) Hydrophobicity Atom type, number glide_hbond (hbond) Hydrogen bond formation propensity Atom type, number Physical… See the full description on the dataset page: https://huggingface.co/datasets/anonymousxxx/MoleculeCLA.tabular1M<n<10M1 likes223 downloads2y agoHugging Face19anonymous321123 /OneOcean_Environment_Dataset oneocean_public_env This folder is prepared for Zenodo upload. Contents Dataset files schema.json: variable/dimension schema hf_sample.parquet: small table sample for Hugging Face dataset schema detection checksums.sha256: SHA256 for all files in this folder Summary Time: 2025-01-01T00:00:00.000000000 -> 2025-01-31T00:00:00.000000000 (n=31) BBox: lat[30.0,40.0], lon[-72.0,-62.0] Resolution (deg): dlat≈0.08264462809917461, dlon≈0.08264462809917461 Vars:… See the full description on the dataset page: https://huggingface.co/datasets/anonymous321123/OneOcean_Environment_Dataset.tabularn<1K0 likes189 downloads7mo agoHugging Face20Anonymous-NeurIPS26-TabularMath /TabularMath TabularMath TL;DR. 114 tabular regression tasks, each compiled from a math word problem into a Python (generator, verifier) pair that is validated against the original seed answer. 2,048 rows per task, integer targets y, zero label noise. Use it to diagnose whether your tabular model can move from fitting to computing under controlled output extrapolation. TabularMath is a program-verified tabular benchmark that probes whether tabular machine-learning models can move from… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-NeurIPS26-TabularMath/TabularMath.tabulartabular-regression100K<n<1M0 likes179 downloads5mo agoHugging Face21AnonymousPaperReview /R2R_Router_Training Anonymous Review Only Note that this dataset is only used for anonymous review. tabulartoken-classification1M<n<10M0 likes167 downloads1y agoHugging Face22anonymous-for-review /AU8image1M<n<10M0 likes166 downloads1y agoHugging Face23AnonymousMouse404 /PawnsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.images.board": { "dtype": "video", "shape": [ 480, 640, 3 ], "names": [ "height", "width", "channel" ], "info": { "video.height":… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousMouse404/Pawns.tabularrobotics10K<n<100K0 likes166 downloads22d agoHugging Face24ano6060 /anonymous-ride-gold-lite RIDE Gold Lite RIDE Gold Lite is the smaller benchmark-ready release of the RIDE dataset. It contains fixed train/test snapshot splits, a canonical evaluation table, and model-ready representations for train delay prediction on Belgian passenger railway operations. This release mirrors the structure and prediction task of RIDE Gold Standard, but uses fewer snapshots and rows for faster inspection, development, and lower-cost experimentation. Links Code repository:… See the full description on the dataset page: https://huggingface.co/datasets/ano6060/anonymous-ride-gold-lite.tabular100K<n<1M0 likes151 downloads5mo agoHugging Face25anonymousML123 /Pexels-Pairs-Masklets-330K Pexels-Pairs-Masklets-330K (review sample) Per-video masklet annotations stored as Parquet shards. This repository is a small sample released for anonymous peer review. It contains 5 shards drawn from 5 different set_* directories of the full collection, which holds roughly 330K shards across 301 sets. Contents are unmodified; only the number of shards is reduced. Layout set_0000/<video_id>.mp4.parquet set_0075/<video_id>.mp4.parquet… See the full description on the dataset page: https://huggingface.co/datasets/anonymousML123/Pexels-Pairs-Masklets-330K.tabularn<1K0 likes146 downloads19d agoHugging Face26double-blind-anonymous /go-mo-dataset GO-MO, a massive Graph agumented Open urban MObility dataset This is the official dataset repository for the GO-MO traffic dataset. The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain). GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024). Additionally, the GO-MO dataset introduces two graph… See the full description on the dataset page: https://huggingface.co/datasets/double-blind-anonymous/go-mo-dataset.tabulartime-series-forecasting1B<n<10B0 likes145 downloads8mo agoHugging Face27anonymous222bit /Ambig-DS-T Ambig-DS-T: Target Ambiguity Benchmark A benchmark for measuring how well data-science agents handle ambiguous prediction targets in tabular Kaggle competitions. Each task is a Kaggle competition derived from DSBench. For every task we provide two prompt variants — one in which the target column is named, and one in which the target is hidden behind two candidate columns. The agent must select and predict the true target; submissions are graded by the original competition metric… See the full description on the dataset page: https://huggingface.co/datasets/anonymous222bit/Ambig-DS-T.tabulartabular-classificationn<1K0 likes135 downloads5mo agoHugging Face28ano6060 /anonymous-ride-gold-standard RIDE Gold Standard RIDE Gold Standard is the full benchmark-ready release of the RIDE dataset. It contains fixed train/test snapshot splits, a canonical evaluation table, and model-ready representations for train delay prediction on Belgian passenger railway operations. This release is intended as the primary benchmark tier for RIDE. It is used for full-scale evaluation and comparison of models under the shared RIDE prediction task and evaluation protocol. Links Code… See the full description on the dataset page: https://huggingface.co/datasets/ano6060/anonymous-ride-gold-standard.tabular1M<n<10M0 likes132 downloads5mo agoHugging Face29anonymous-motif-scaffolding /prosite_functional_motif_scaffolding_benchmark PROSITE-derived Functional Motif Benchmark This archive contains an anonymized dataset artifact for a systematically derived benchmark of structurally conserved functional motif-scaffolding cases from PROSITE-linked experimental protein structures. The benchmark is intended for static motif-scaffolding evaluation with standard MotifBench-style pipelines. Cases are derived from PROSITE motif-pattern entries, mapped to experimentally resolved PDB structures, filtered for recurrent… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-motif-scaffolding/prosite_functional_motif_scaffolding_benchmark.tabularothern<1K0 likes131 downloads5mo agoHugging Face30anonymous-neurips-ED /CTSpinoPelvic1K CTSpinoPelvic1K A fused spine + pelvis 3D CT segmentation dataset built by patient-level crosswalk between three public sources: TCIA CT COLONOGRAPHY — DICOM CT volumes (prone + supine per patient) CTSpine1K (COLONOG subset) — VerSe-convention vertebral label masks CTPelvic1K dataset2 — sacrum + bilateral hip label masks Annotations are placed onto the TCIA CT volume with the highest bone coverage (HU > 200), separately per anatomy. For ~650 patients both annotations land on the… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips-ED/CTSpinoPelvic1K.tabularimage-segmentation1K<n<10K1 likes120 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.