datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VietPET-RoI
VietPET-RoI
VietPET-RoI is a Vietnamese whole-body PET/CT dataset containing paired
cropped 3D volumes, regional reports, and modality-specific 3D ROI bounding
boxes. It is intended for medical multimodal research, report generation,
visual question answering, and ROI grounding.
Research use only. This dataset is not intended for diagnosis, treatment
decisions, or direct patient care.
Summary
Split
Patients
CT/PET region pairs
ROIs
Train
160
480
1,544… See the full description on the dataset page: https://huggingface.co/datasets/scarlettlin/VietPET-RoI.quote-repetition
quote-repetition (Joe Cavanagh, Andrew Gritsevskiy, and Derik Kauffman of Cavendish Labs)
General description
In this task, the authors ask language models to repeat back sentences given in the prompt, with few-shot examples to help it recognize the task. Each prompt contains a famous quote with a modified ending to mislead the model into completing the sequence with the famous ending rather than with the ending given in the prompt. The authors find that smaller models… See the full description on the dataset page: https://huggingface.co/datasets/inverse-scaling/quote-repetition.Void-Witch-Astra-Vanta
Void Witch Astra Vanta
Source-derived release with authored context (schema 4)
448 rows: 93 unchanged conversation exchanges and 355 document chunks.
All 1,623 nonblank authored source lines appear exactly once as body text.
No passages are omitted. The row count changed from 788 because passages,
headings and lists are now grouped by their source relationships.
The seven original .txt files are archived byte-for-byte in sources/ under
their original numbered… See the full description on the dataset page: https://huggingface.co/datasets/scarletdeath/Void-Witch-Astra-Vanta.inverse-scaling-ttc-main
Inverse Scaling in Test-Time Compute
Paper: Inverse Scaling in Test-Time Compute
Project Page: https://safety-research.github.io/inverse-scaling-ttc/
Abstract
We construct evaluation tasks where extending the reasoning length of Large Reasoning Models (LRMs) deteriorates performance, exhibiting an inverse scaling relationship between test-time compute and accuracy. Our evaluation tasks span four categories: simple counting tasks with distractors, regression tasks with… See the full description on the dataset page: https://huggingface.co/datasets/inverse-scaling-ttc/inverse-scaling-ttc-main.MV-ScanQANeQA
NeQA: Can Large Language Models Understand Negation in Multi-choice Questions? (Zhengping Zhou and Yuhui Zhang)
General description
This task takes an existing multiple-choice dataset and negates a part of each question to see if language models are sensitive to negation. The authors find that smaller language models display approximately random performance whereas the performance of larger models become significantly worse than random.
Language models failing to follow… See the full description on the dataset page: https://huggingface.co/datasets/inverse-scaling/NeQA.finewebs-scandeval-results
ScandEval Results on English NLU
We use ScandEval in revision 8766d2a to conduct experiments with our pretrained FineWeb LMs.
Additionally, results for BERT, RoBERTa and ELECTRA were also performed to have a nice comparison.
Model ID
Avg. Score
CoNLL-En
SST5
ScaLA-En
SQuAD
model-garden-lms/bert-base-finewebs-1m
69.03
88.98 ± 0.43 / 88.67 ± 0.36
58.11 ± 1.2 / 59.77 ± 1.49
57.29 ± 3.57 / 77.15 ± 2.17
55.82 ± 1.35 / 66.46 ± 1.51
model-garden-lms/bert-base-finewebs-951k… See the full description on the dataset page: https://huggingface.co/datasets/model-garden-lms/finewebs-scandeval-results.agent-env-scaling-data
Agent Environment Scaling Reviewed Data
This Dataset repository stores reviewed data artifacts separately from the implementation
repository. index.json identifies the current accepted snapshot. The tables/ directory exposes
small JSONL views for the Dataset Viewer; snapshots/ preserves portable content-addressed Stores,
readable previews, manifests, and exact validation boundaries.
Current snapshot
warehouse-tongyi-formal-20260806 is the newest complete result… See the full description on the dataset page: https://huggingface.co/datasets/graycatHCO3/agent-env-scaling-data.mcp-security-scan-2026
MCP Security Scan Dataset 2026
Security scan results for 4,867 MCP (Model Context Protocol) server repositories, scanned by MCPShield.
Dataset Description
This is the largest public labeled MCP security dataset. Each entry contains the security grade, score, and detailed findings for a GitHub repository implementing an MCP server.
Scanner
MCPShield v5.0 — Two-pass detection architecture:
Pass 1: 49 regex rules covering OWASP MCP Top 10 (94% detection on… See the full description on the dataset page: https://huggingface.co/datasets/MCPShield/mcp-security-scan-2026.scalewob-environments
ScaleWoB Environments
Private, versioned browser-runtime assets corresponding to the scalewob-verl dataset package.
Contents
scalewob-env-v0.1.0.tar.zst: deterministic archive containing the scalewob-env/ directory.
manifest.json: archive checksum, file counts, indexed environment count, and compatible runtime versions.
THIRD_PARTY_NOTICES.md: preliminary redistribution audit notes; not a complete license inventory.
The source tree is approximately 355 MB and… See the full description on the dataset page: https://huggingface.co/datasets/hysi-lab/scalewob-environments.arch-code-scale-lpi-260904T0500-code_block_only_superintelligence_n1000_e3_s1_a0.5SCALARnegative_resultsopenenv-scalingnuzzle-scan-saraprice-llama2-7b-backdoor-deploymentinverse-scaling-ttc-main
Inverse Scaling in Test-Time Compute
Note: This is an anonymized repository.
Abstract
We construct evaluation tasks where extending the reasoning length of Large Reasoning Models (LRMs) deteriorates performance, exhibiting an inverse scaling relationship between test-time compute and accuracy. Our evaluation tasks span four categories: simple counting tasks with distractors, regression tasks with spurious features, deduction tasks with constraint tracking, and advanced AI… See the full description on the dataset page: https://huggingface.co/datasets/anonscaling/inverse-scaling-ttc-main.dataset-injection-scan-study
Dataset Injection Scan — open study of popular HF datasets
from datasets import load_dataset
ds = load_dataset("fevziegeyurtsevenler/dataset-injection-scan-study")
Results of scanning 17,000 rows across 6 popular public instruction/prompt datasets for
smuggled prompt-injection with hf-dataset-scan
(invisible Unicode, injection phrasing EN+TR, exfil URLs).
Headline: no smuggled injection found
Dataset
Rows
Flagged
High
Med
Low
tatsu-lab/alpaca
3,000
0… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/dataset-injection-scan-study.SCALAR-VG
SCALAR_VG
While the world model continues to advance, existing datasets remain inadequate for supporting large-scale multi-modal training, particularly in comprehensive multi-dimensional scene-aware understanding. Therefore, we have built the SCALAR-VG through the SCALAR, integrating and extending many open-source image datasets to meet this demandImportantly, It contains about 240K images with comprehensive, hierarchical and multi-dimensional annotations.
Compared with existing… See the full description on the dataset page: https://huggingface.co/datasets/MiaoMiaoYang/SCALAR-VG.nuzzle-scan-openai-community-gpt2scanner-poisoned-iris-benchmark
Scanner Poisoned Iris Benchmark
This benchmark starts from the classic UCI Iris dataset and injects multiple synthetic poisoning patterns so dataset scanners can exercise duplicate, anomaly, missingness, skew, and divergence heuristics against a small tabular corpus.
Recommended Hugging Face repo slug: your-org/scanner-poisoned-iris-benchmark
What It Is For
benchmarking dataset quality and poisoning detection workflows
regression-testing scanner heuristics on a… See the full description on the dataset page: https://huggingface.co/datasets/jgracie52/scanner-poisoned-iris-benchmark.pqc-ssl-scans
PQC Vulnerability Scan Dataset
SSL/TLS certificate scans of 45 major finance, healthcare, and government domains, scored for post-quantum cryptography (PQC) migration urgency.
Dataset Description
Each row represents a live SSL certificate scan performed on 2026-03-24 using hiero-cli-pqc.
Features
Feature
Type
Description
domain
string
Scanned domain name
key_algorithm
string
Public key algorithm (RSA, ECDSA, Ed25519)
key_size
int
Key size in bits… See the full description on the dataset page: https://huggingface.co/datasets/Q-GRID/pqc-ssl-scans.eval_educational_promptsythetic_casual_relation_medium_scaleserrm-icml-26731-scaled-reproInverse-scaling-testJimmy19991222__llama-3-8b-instruct-gapo-v2-rougeL-beta10-gamma0.3-lr1.0e-6-scale-log-details
Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-rougeL-beta10-gamma0.3-lr1.0e-6-scale-log
Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-rougeL-beta10-gamma0.3-lr1.0e-6-scale-log
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-rougeL-beta10-gamma0.3-lr1.0e-6-scale-log-details.scanner_annJimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-details
Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log
Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-details.Scale-SWE-Agent_AweAgent_vllm_16kproteins_scanprosite
