datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kitti-yolo11n-robustness-benchmark
KITTI YOLO11n Robustness & Adversarial Benchmark Suite
This dataset contains 649,425 benchmark samples evaluating the perception robustness of YOLO11n (Ultralytics YOLOv11 nano in original FP32 precision) on the official KITTI Object Detection train set (3,711 images) under 35 attack & corruption techniques across 5 severity levels.
?? Benchmark Leaderboard (mAP@0.5 Drop on YOLO11n)
Clean Baseline AP50: 0.3555
Evaluation Model: YOLO11n (Original weights:… See the full description on the dataset page: https://huggingface.co/datasets/VietPhong/kitti-yolo11n-robustness-benchmark.staining-robustness-evaluation
A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models
This repository provides the stain references, pretrained models, and experimental results required to:
Define custom staining references using our PLISM reference library
Reproduce our published controlled staining robustness experiments
👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main
👉 Associated publication: Paper
Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.lighting-invariant-bedroom-perception-robustness-benchmark
Lighting-Invariant Bedroom Perception & Robustness Benchmark
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/lighting-invariant-bedroom-perception-robustness-benchmark.bangla-noise-robustness-datacode-switching-tokenizer-robustness
Code-Switching Dataset for Tokenizer Robustness Analysis
Dataset Description
This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models.
Purpose
Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.path-vqa-robustnesstokenization_robustness_v102
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/tokenization_robustness_v102.slake-robustnessprobe-robustness-rotten_tomatoes
rotten_tomatoes
Dataset repo: wrynx/probe-robustness-rotten_tomatoes
Auto-generated by prepare_datasets.py. Do not hand-edit -- regenerate by re-running the script (with --force) instead.
Stats
Total records: 10662
Records per split:
test: 1066
train: 8530
valid: 1066
Number of classes: 2
Records per class:
0: 5331
1: 5331
Records per class per split:
test:
0: 533
1: 533
train:
0: 4265
1: 4265
valid:
0: 533
1: 533
Original dataset README… See the full description on the dataset page: https://huggingface.co/datasets/wrynx/probe-robustness-rotten_tomatoes.lighting-invariant-bedroom-perception-robustness-benchmark-next-pack-1917c2cb-f1975230
Home Object Detection, Grasping and Sorting Eval — YOLOv8
Evaluation dataset for a robotic arm that detects, grasps, and sorts objects by type in home environments. 30 renders at 640x640 across kitchen, entry, living room and dressing spaces, staged with everyday household objects. Includes RGB plus metric depth, world-space normals (OpenGL, linear), albedo and material index passes, per-frame annotations, and midday lighting. Targets a YOLOv8 model.
This dataset mirrors public… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/lighting-invariant-bedroom-perception-robustness-benchmark-next-pack-1917c2cb-f1975230.openvla-oft-plus-robustness-recovery
OpenVLA-OFT+ Robustness and Recovery Dataset
成功した操作だけを増やせば、ロボット方策は頑健になるに違いない。
しかし、OpenVLA-OFT+の挙動を条件別に確かめると、失敗は一様ではなかった。
Spatialではsafe successが63.3%まで下がり、collision rateは35.0%に達した。
Goalのsafe success 90.0%、Objectの95.0%と比べても、空間関係を扱う操作の不安定さが際立つ。
このデータセットは、その差を学習データへ戻すために作成した。
遠距離の空間操作に加え、把持に失敗した状態、持ち上げられなかった状態、運搬中に物体を落とした状態から、成功まで立て直す軌跡を収録している。
弱点の観測
各suiteは20 scene、3条件、合計60エピソードで測定した。
95%信頼区間はscene単位のcluster bootstrap(10,000回)で算出している。
Suite
Episodes
Safe… See the full description on the dataset page: https://huggingface.co/datasets/argo11/openvla-oft-plus-robustness-recovery.probe-robustness-sms_spam
sms_spam
Dataset repo: wrynx/probe-robustness-sms_spam
Auto-generated by prepare_datasets.py. Do not hand-edit -- regenerate by re-running the script (with --force) instead.
Stats
Total records: 5159
Records per split:
test: 804
train: 3604
valid: 751
Number of classes: 2
Records per class:
ham: 4517
spam: 642
Records per class per split:
test:
ham: 708
spam: 96
train:
ham: 3143
spam: 461
valid:
ham: 666
spam: 85
Original dataset README… See the full description on the dataset page: https://huggingface.co/datasets/wrynx/probe-robustness-sms_spam.chiboard-1.1-robustness-sft
Chiboard 1.1 robustness supplemental SFT — Plan 04.6 candidate
This immutable candidate successor preserves the Plan 04.5 rows and adds the
document-unique Plan 04.6 capacity build. Use the plan04_6_eligible config for
the one-row-per-source-document candidate arm. The legacy rows remain available
for provenance and historical comparison but are not eligible in the maximum
Plan 05 schedule.
The prompt contract remains qwen35-chiboard-field-tokens-v2; completion alone
receives… See the full description on the dataset page: https://huggingface.co/datasets/johnbean393/chiboard-1.1-robustness-sft.Nemotron-RL-Jailbreak-Robustness-v1
Dataset Description:
The Nemotron-RL-Jailbreak-Robustness-v1 data is designed to (1) strengthen model robustness against a variety of adversarial jailbreak techniques and (2) at the same time improve adherence to behavioral policies.
This dataset is a collection of hybrid (open-source and synthetically generated) collection of adversarial prompts designed to elicit undesirable behavior from large language models. That's it, just prompts, responses are generated during training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Jailbreak-Robustness-v1.Robustness
ELOQUENT Robustness and Consistency Task
This dataset contains the sample and test datasets for the Robustness and Consistency task, which is part of the ELOQUENT lab. This dataset is for participants to generate texts for prompt variants, to investigate prompt style conditioned variation.
Robustness task
ELOQUENT lab
CLEF conference 9-12 September 2025
The task in brief (this is a simple task to execute!)
This dataset provides a number of questions in several… See the full description on the dataset page: https://huggingface.co/datasets/Eloquent/Robustness.omni-med-vqa-mini-robustnesspmc-vqa-robustness2026.RA.Five-Seat-Qwen3-8B-Robustness
Five-Seat Qwen3-8B Private Robustness Subset
⚠ ERRATUM (2026-08-10) — the +0.449 one-oracle result is measured on a spoiled ballot
OmniscientBestResponsePolicy, the computable seat in the one-oracle arm, cast its forced-final vote on
whichever live offer it valued most instead of on the one under the up/down vote; the protocol rejected that as
a legality error and the turn was recorded as a silent pass. Fixed in commit ca20157 (2026-08-10), after
these episodes… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Five-Seat-Qwen3-8B-Robustness.repro-learning-rate-annealing-improves-tuning-robustness-in-stochastic-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
vqa-med-robustnessnl2sh-chatter-robustness
Chatter / robustness pairs for NL->shell models
246 hand-written (natural language, shell command) pairs teaching the
"boring reflex": greetings, small talk, identity questions and nonsense
input map to harmless commands (echo hello, pwd) instead of garbage
or network-touching behavior.
Generated by organic_augment.py (deterministic, seed 42). Used in the
training pool of barbarabhb/nl2sh-qwen25-coder-1.5b-GGUF.
probe-robustness-truthfulqa
truthfulqa
Dataset repo: wrynx/probe-robustness-truthfulqa
Auto-generated by prepare_datasets.py. Do not hand-edit -- regenerate by re-running the script (with --force) instead.
Stats
Total records: 817
Records per split:
test: 114
train: 591
valid: 112
Label column: not present / not populated for this dataset.
Original dataset README (from truthfulqa/truthful_qa)
Reproduced here from the source dataset's own card (license, citation, task… See the full description on the dataset page: https://huggingface.co/datasets/wrynx/probe-robustness-truthfulqa.robotrace-vla-robustness-traces
RoboTrace Evidence Bundle
This dataset repository contains the public evidence bundle for RoboTrace, a low-cost deployment-stress evaluation scaffold for robot-learning and VLA-style inference pipelines.
The current release evaluates lerobot/pusht and includes reports, metrics, plots, summaries, and release manifests from a complete staged run.
What this bundle is for
Use this repository to inspect evidence from RoboTrace:
action-trace stability metrics
visual… See the full description on the dataset page: https://huggingface.co/datasets/i-am-shaurya05/robotrace-vla-robustness-traces.repro-towards-optimal-robustness-in-learning-augmented-paging-traces
Agent traces
Agent sessions published from a Trackio Logbook.
farsi_tokenizer_robustness
TokSuite Benchmark (Farsi Collection)
Dataset Description
This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Farsi (Persian) language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness.
Curated by: R3 Research Team
Language(s): Farsi/Persian (fa)
License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/r-three/farsi_tokenizer_robustness.tokenizer-robustness-mmlu
Tokenizer Robustness MMLU Dataset
This dataset contains MMLU-formatted questions and answers designed to test tokenizer robustness across different text formats and languages.
Dataset Description
The dataset consists of the same questions presented in 6 different formats, with both test (20 questions) and development (5 questions) sets:
original - Standard formatted questions
minor_spelling_errors - Questions with minor misspellings
spoken_language - Questions in casual… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/tokenizer-robustness-mmlu.text-to-sql-phrasing-robustness
Does sloppy phrasing break text-to-SQL?
The enterprise text-to-SQL benchmark
lists its own biggest caveat: every question is template-generated, so real user phrasing is untested.
This is the test. 35 test questions (one per template), each sent to the deployed
pipeline four ways: as written, with a typo, in business shorthand, and stripped to a terse fragment.
24 questions and 85 answers survive the filter described under Setup; every answer was
executed against the database.… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-phrasing-robustness.probe-robustness-ag_news
ag_news
Dataset repo: wrynx/probe-robustness-ag_news
Auto-generated by prepare_datasets.py. Do not hand-edit -- regenerate by re-running the script (with --force) instead.
Stats
Total records: 127600
Records per split:
test: 25644
train: 84108
valid: 17848
Number of classes: 4
Records per class:
0: 31900
1: 31900
2: 31900
3: 31900
Records per class per split:
test:
0: 6451
1: 6412
2: 6312
3: 6469
train:
0: 20968
1: 21030
2: 21078
3: 21032
valid:
0: 4481… See the full description on the dataset page: https://huggingface.co/datasets/wrynx/probe-robustness-ag_news.certainty-robustness-llm-evaluation
Certainty Robustness Benchmark
This repository accompanies the paper:
Certainty robustness: Evaluating LLM stability under self-challenging promptsMohammadreza Saadat, Steve NemzerarXiv:2603.03330, 2026https://arxiv.org/abs/2603.03330
Overview
The Certainty Robustness Benchmark evaluates how large language models (LLMs) behave when their initial answers are challenged by follow-up prompts such as:
“Are you sure?”
“You are wrong!”
confidence elicitation prompts
Rather… See the full description on the dataset page: https://huggingface.co/datasets/Reza-Telus/certainty-robustness-llm-evaluation.tts-robustness-benchmark
TTS Robustness Benchmark
Overview
The TTS Robustness Benchmark is a set of evaluation samples used to calculate domain-wise CER (Character Error Rate) scores across 7 critical stress-test categories. This benchmark ensures that each sentence-language pair appears exactly once, providing a clean and reliable metric for TTS robustness.
Blog Post: Read the official announcement
Column Descriptions
Column
Description
text
The original input… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/tts-robustness-benchmark.
