datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FACETpickapic_v1
Dataset Card for "pickapic_v1"
More Information needed
MultiSetTransformerData
General Description
MultiSetTransformerData is a large dataset designed to train and validate neural Symbolic Regression models. It was designed to solve the Multi-Set Symbolic Skeleton Prediction (MSSP) problems, described in the paper "Univariate Skeleton Prediction in Multivariate Systems Using Transformers". However, it can be used for training generic SR models as well.
This dataset consists of artificially generated univariate symbolic skeletons, from which mathematical… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousGM/MultiSetTransformerData.JALMBench
About the Dataset
📦 JALMBench contains 245,355 audio samples and 11,316 text prompts to benchmark jailbreak attacks against audio-language models (ALMs). It consists of three main categories:
🔥 Harmful Query Category:Includes 246 harmful text queries ($T_{Harm}$), their corresponding audio ($A_{Harm}$), and a diverse audio set ($A_{Div}$) with 9 languages, 2 genders, 3 accents, and 3 TTS methods.
📒 Text-Transferred Jailbreak Category:Features adversarial texts generated by… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousUser000/JALMBench.GenIRCrowdEvalTSP_EXECUTION_RUNSfactorjepa-outputsvideo2mentalFlowGen
🌟 FlowGen
FlowGen is a controllable flowchart synthesizer that generates diagrams with tunable structural features and supports multiple rendering styles.
📑 Dataset description
This dataset contains different types of renderer flowchart images with different difficulty levels.
Types
Train (Easy)
Train (Medium)
Train (Hard)
Test (Graph Easy)
Test (Graph Medium)
Test (Graph Hard)
Test (Scanned Easy)
Test (Scanned Medium)
Test (Scanned Hard)
Mermaid
960
960
960… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous112233/FlowGen.LM-SimBench_rawMoM-CLAM-dataset
Melody or Machine: Benchmarking and Detecting Synthetic Music via Cross-Modal Contrastive Alignment
To support robust detection of AI-generated songs under diverse manipulations and model types, we present the Melody or Machine (MoM) Dataset—a large-scale benchmark reflecting the evolving landscape of song-level deepfakes. MoM spans three authenticity tiers: genuine recordings, synthetic audio with real lyrics/voice, and fully synthetic tracks, enabling evaluation across progressive… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2212/MoM-CLAM-dataset.HydroAgent-dataset
HydroAgent Dataset
Calibration assets for the HydroAgent
LLM agent — everything the agent (and the
EF5/CREST hydrologic simulator it controls)
needs to calibrate streamflow at a USGS gage:
CONUS terrain & default parameter rasters (basic_data/)
Per-gage MRMS hourly precipitation clips (data_mrms_clip/)
Hourly USGS streamflow observations (gauge/)
Daily potential ET rasters (pet/)
EF5 control-file template (docs/)
73 GPT-4o calibration trajectories across 29 gages (sets_for_SFT_RL/)… See the full description on the dataset page: https://huggingface.co/datasets/anonymousOwl/HydroAgent-dataset.pde-geo
Geometry-Aware PDE Benchmark Dataset
Dataset Description
This dataset contains geometry-aware partial differential equation (PDE)
simulation data for scientific machine learning, operator learning, and
generative PDE modeling experiments. It includes three subsets:
Darcy: static Darcy-flow samples on polygonal geometries. Each geometry
includes a triangular mesh, a 128 x 128 coefficient grid, signed-distance
information, and scalar solution values on mesh nodes.
Poisson:… See the full description on the dataset page: https://huggingface.co/datasets/An-onymous/pde-geo.DriftBench
DriftBench
A benchmark for measuring trajectory drift in multi-turn LLM-assisted
scientific ideation. When researchers iteratively refine ideas with an LLM,
do the models preserve fidelity to the original objective, or drift toward
locally coherent but globally misaligned elaborations?
Headline result (reproducible from this dataset)
All 7 evaluated models inflate complexity under iterative pressure.
5 of 7 models drift on at least 50% of briefs (constraint adherence < 3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-driftbench/DriftBench.review-dataset
MSIR-Bench Review Dataset
This repository contains an anonymized review snapshot of MSIR-Bench, a benchmark for identity-preserving style image retrieval.
Dataset Description
Each source identity is represented by an anonymous five-digit ID. Images are organized by split and identity folder. File names follow either <id>_<Style>.png, <id>_original.png, or legacy original.jpg for original reference images.
The dataset is intended for evaluating whether a retrieval… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-review-dataset-2026/review-dataset.the-stack-repoThis version of the dataset is strictly permitted for use exclusively in conjunction with the review process for the paper. Upon completion of the review process, a de-anonymized version of the dataset will be released under a license similar to that of The Stack, which can be found at https://huggingface.co/datasets/bigcode/the-stack.
cua-speedrun-trajectories
CUA Speedrun results
Trajectory archives and measured results for the benchmark sets used by CUA Speedrun.
Unanimous-295
Results on the 295-task OSWorld set used by CUA Speedrun. Except for the explicitly reported GPT-5.6 Luna aggregate row, scores are the mean normalized verifier score across the 295 tasks and may include partial credit. Times are measured task time per task.
Model
Average score
Average task time
Average cost per task
Kimi K3
85.07%… See the full description on the dataset page: https://huggingface.co/datasets/anonymousmypcbench/cua-speedrun-trajectories.vocalgrad
VocalGrad
VocalGrad is an audio benchmark for evaluating whether a model can detect the
direction of gradual perceptual change in speech. This public release contains
the test split only.
Each example contains one audio clip and one target attribute. The task is to
answer whether that attribute increases or decreases over time.
Task
Given an audio clip and an attribute name, predict one of two labels:
increase
decrease
The ground-truth label is derived from the metadata… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-user-592888/vocalgrad.CC3MStereoNav_DataSetProseOnlyRepair_linguistic_MQEDGAR_FILINGS_DATASET_2022_2026H1recap-t2i-evaluation-sample-2026
Recaptioned T2I Supervision Evaluation Sample
This repository is the small reviewer-inspection companion to the full anonymous caption-metadata release. The full release is hosted separately at https://huggingface.co/datasets/Anonymous1477/recap-t2i-evaluation-metadata-2026; this repository stays under the large-dataset sample threshold and gives reviewers a direct way to inspect redacted caption metadata, join structure, and selected image-conditioned audit packages.… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous1477/recap-t2i-evaluation-sample-2026.multisite-ppg-submission
Multisite PPG Dataset
A multisite photoplethysmography (PPG) dataset: long-duration recordings from four body locations, synchronized activity logs, and ECG-derived heart-rate ground truth. It supports PPG-based HR estimation, signal-quality assessment, motion-artifact handling, and cross-site generalization research.
For data preprocessing and baseline training code, see our GitHub repository: anonymous-ppg/wearable-ppg-dataset.
Submission note: This is the anonymized submission… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-ppg-dataset/multisite-ppg-submission.EDGAR_FILINGS_DATASET
SFD: SEC Filings Dataset (v1)
SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation.
This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in:
The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.llbench-dataset
LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models via Human Preferences
Anonymous release prepared for NeurIPS 2026 review. Please do not redistribute.
LL-Bench is a large-scale, human-preference benchmark for evaluating low-level
vision restoration in the era of large generative models (LGMs). It compares
10 LGMs with 16 specilist and 5 all-in-one models across 16 low-level vision tasks, paired with dense human annotations:pairwise… See the full description on the dataset page: https://huggingface.co/datasets/anonymousllbench/llbench-dataset.ProseOnlyRepair_linguistic_LQomicsbench-questions
omicsbench-questions
Questions used for producing the OmicsBench manuscript.
Causal_Plan
Causal Plan
Causal Plan is a unified multimodal dataset release for training and evaluating causal reasoning over visually grounded plans. The repository is organized as one entry point with three clearly separated resources:
Causal_Plan/
CausalPlan-1M-QA/
CausalPlan-1M-FourStage-Metadata/
Causal-Plan-Bench/
DATASET_MANIFEST.json
verify_alignment.py
README.md
The QA examples, item-level four-stage metadata, and benchmark package are stored in the same repository so that… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-causal-plan/Causal_Plan.
