CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anonymous-neurips26-ljasd /LM-SimBench LM-SimBench Dataset Description LM-SimBench is a large-scale training-performance profiling dataset for large language models. The dataset is collected from training runs based on the MindSpeed-LLM framework and the Ascend NPU development stack, covering multiple model families, context lengths, and distributed parallel configurations. Each model is sampled under feasible combinations of data parallelism (DP), tensor parallelism (TP), pipeline parallelism (PP), context… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips26-ljasd/LM-SimBench.tabulartabular-regressionn<1K0 likes619 downloads5mo agoHugging Face02anonymous-neurips26-ljasd /LM-SimBench_example LM-SimBench (Example Snapshot) Dataset Description This repository distributes a compact example snapshot of LM-SimBench, the structured CSV release of large-scale LLM training-performance profiling data. The snapshot is provided so reviewers and readers can inspect file layout, schemas, and representative records without downloading the multi–tens-of-gigabyte full release. The profiling methodology, software stack, and field definitions are the same as in the complete… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips26-ljasd/LM-SimBench_example.texttabular-regressionn<1K0 likes181 downloads5mo agoHugging Face03anon-neurips-ed-3552 /SASS-Bench SASS-Bench Hardware-grounded benchmark for execution prediction on NVIDIA GPU assembly (SASS). Given a compiled SASS kernel, launch configuration, raw inputs, and a list of checkpoint queries, a model must predict exact uint32 register values at specified instructions and the bit-pattern of the final output buffers. Ground truth is captured from a Blackwell sm_120 GPU via NVBit dynamic binary instrumentation. The initial release is 50 kernels with 2--3 input-regime variants each… See the full description on the dataset page: https://huggingface.co/datasets/anon-neurips-ed-3552/SASS-Bench.tabularothern<1K1 likes137 downloads5mo agoHugging Face04causalverify /causalverify-neurips2026 🎯 CausalVerify An Execution-Grounded Benchmark for LLM Causal Inference Workflows NeurIPS 2026 — Evaluations and Datasets Track · double-blind review · frozen at tag neurips2026-submission 💡 TL;DR A benchmark of 259 published economics papers (Experiment A — real-paper text-agreement diagnostic) and 100 fixed-seed synthetic data-generating processes (Experiment B — execution-grounded coefficient recovery), evaluating 7 frontier LLMs. The central… See the full description on the dataset page: https://huggingface.co/datasets/causalverify/causalverify-neurips2026.tabulartabular-regressionn<1K0 likes134 downloads5mo agoHugging Face05neurips-dataset-1211 /DIVE Dataset Card for DIVE This dataset contains safety ratings for image and text inputs. It contains 1000 adversarial prompts and 5 attention check prompts There are 35164 safety annotations from high-quality raters and 3246 safety annotations from low-quality raters The total number of ratings in the dataset is 38410 equal of the number of rows in this dataset. All the ratings in this dataset are provided by 707 demographically diverse raters - 637 are deemed high-quality and… See the full description on the dataset page: https://huggingface.co/datasets/neurips-dataset-1211/DIVE.tabulartext-classification10K<n<100K1 likes47 downloads8mo agoHugging Face06neurips2026lrtransfer /LR-Transfer-Trajectory Dataset Documentation Overview This dataset captures per-step training and validation metrics from training runs of a 12-layer GPT-style decoder-only transformer. Each run is stored as a single .csv file in which every row corresponds to one logged step, and several columns hold per-parameter measurements encoded as JSON. The dataset is designed to support post-hoc analysis of: Loss curves (train / val) Throughput and step latency Per-layer / per-parameter dynamics:… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026lrtransfer/LR-Transfer-Trajectory.tabular100K<n<1M1 likes32 downloads5mo agoHugging Face07for-anonymous-submission /NeurIPS-ED-2026-SubID-956 Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs 👁️Overview This repository contains the official dataset of VisFactor, a novel benchmark derived from the Factor-Referenced Cognitive Test (FRCT) that digitizes 20 vision-centric subtests from established cognitive psychology assessments. Our work systematically investigates the gap between human visual cognition and state-of-the-art Multimodal Large Language Models (MLLMs). 🎯 Key… See the full description on the dataset page: https://huggingface.co/datasets/for-anonymous-submission/NeurIPS-ED-2026-SubID-956.tabularvisual-question-answering1K<n<10K0 likes23 downloads2mo agoHugging Face08anon-neurips26-3390 /IndustryBench IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs Anonymized review copy — NeurIPS 2026 Evaluations & Datasets Track, Submission 3390 (under review). IndustryBench is a benchmark for evaluating the industrial procurement knowledge of large language models. It comprises 2,049 QA pairs grounded in Chinese national standards (GB/T) and structured industrial product records, with item-aligned renderings in Chinese, English, Russian, and Vietnamese. Erratum note:… See the full description on the dataset page: https://huggingface.co/datasets/anon-neurips26-3390/IndustryBench.textquestion-answering1K<n<10K0 likes22 downloads2mo agoHugging Face09facebook /beyond_the_lab_neurips_papertabularimage-classification100K<n<1M1 likes21 downloads5mo agoHugging Face10MajorTimberWolf /NeurIPS-Dataset NeurIPS Papers Dataset This dataset contains information about NeurIPS conference paper submissions including peer reviews, author rebuttals, and decision outcomes across multiple years. Files dataset.csv: Main dataset file containing all paper submission data Dataset Structure The CSV file contains the following columns: title: Paper title paper_decision: Decision outcome (Accept/Reject with specific categories) review_1, review_2, etc.: Peer reviews from… See the full description on the dataset page: https://huggingface.co/datasets/MajorTimberWolf/NeurIPS-Dataset.text1K<n<10K0 likes20 downloads1y agoHugging Face11neurips-ed2026-anon-checkpoints /checkpoint-zoo-metadata LowRankArena Checkpoint Zoo Metadata Catalog This repository is a companion metadata catalog for the checkpoint zoo hosted in the linked Hugging Face Model repository. Each artifact-catalog record points to a checkpoint or reconstruction artifact at a fixed source commit. This repository does not duplicate or modify the referenced artifacts, and its CC BY 4.0 license does not apply to checkpoint files or upstream model weights. Source commit and coverage The full… See the full description on the dataset page: https://huggingface.co/datasets/neurips-ed2026-anon-checkpoints/checkpoint-zoo-metadata.textn<1K0 likes20 downloads2mo agoHugging Face12anon123312 /retrieval-conditional-neurips2026 Dataset Release — Retrieval-Conditional NeurIPS 2026 This bundle accompanies the NeurIPS 2026 D&B Track submission "To Retrieve or Not to Retrieve? Most of the Benefit is Structural, Not Semantic." Contents File Config name Description data/per_task_outcomes.csv per_task_outcomes (default) Per-(backbone × env × condition × task) success/failure labels. 3,064 rows. data/stats_per_cell.csv stats_per_cell 54-cell aggregate success rates and pairwise contrasts.… See the full description on the dataset page: https://huggingface.co/datasets/anon123312/retrieval-conditional-neurips2026.tabularother1K<n<10K0 likes18 downloads5mo agoHugging Face13neurips26-sycophancy /GoalPrefBench GoalPref-Bench: Goal-Preference Alignment Benchmark GoalPref-Bench evaluates how AI systems handle conflicts between users' long-term goals and immediate preferences. The benchmark tests whether models prioritize helping users achieve their stated objectives or instead accommodate conflicting preferences that undermine those goals. Overview This benchmark addresses a core question in AI alignment — when a user's immediate preference conflicts with their long-term goal… See the full description on the dataset page: https://huggingface.co/datasets/neurips26-sycophancy/GoalPrefBench.textn<1K0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.