CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01VietPhong /kitti-yolo11n-robustness-benchmark KITTI YOLO11n Robustness & Adversarial Benchmark Suite This dataset contains 649,425 benchmark samples evaluating the perception robustness of YOLO11n (Ultralytics YOLOv11 nano in original FP32 precision) on the official KITTI Object Detection train set (3,711 images) under 35 attack & corruption techniques across 5 severity levels. ?? Benchmark Leaderboard (mAP@0.5 Drop on YOLO11n) Clean Baseline AP50: 0.3555 Evaluation Model: YOLO11n (Original weights:… See the full description on the dataset page: https://huggingface.co/datasets/VietPhong/kitti-yolo11n-robustness-benchmark.tabularobject-detection100K<n<1M0 likes3.3k downloads25d agoHugging Face02CTPLab-DBE-UniBas /staining-robustness-evaluation A Protocol for Evaluating Robustness to H&E Staining Variation in Computational Pathology Models This repository provides the stain references, pretrained models, and experimental results required to: Define custom staining references using our PLISM reference library Reproduce our published controlled staining robustness experiments 👉 Code repository: https://github.com/lely475/staining-robustness-evaluation/tree/main 👉 Associated publication: Paper Overview: How… See the full description on the dataset page: https://huggingface.co/datasets/CTPLab-DBE-UniBas/staining-robustness-evaluation.tabular100K<n<1M1 likes2.7k downloads4mo agoHugging Face03physicl /lighting-invariant-bedroom-perception-robustness-benchmark Lighting-Invariant Bedroom Perception & Robustness Benchmark Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/lighting-invariant-bedroom-perception-robustness-benchmark.imagen<1K0 likes2.1k downloads3mo agoHugging Face04JabaleNurAdnan /bangla-noise-robustness-datatext100K<n<1M0 likes566 downloads9d agoHugging Face05Malikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes290 downloads1y agoHugging Face06simwit /path-vqa-robustnessimage10K<n<100K0 likes249 downloads1y agoHugging Face07r-three /tokenization_robustness_v102 Dataset Card for Tokenization Robustness A comprehensive evaluation dataset for testing robustness of different tokenization strategies. Dataset Details Dataset Description This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling. Curated by: R3 Funded by [optional]: [More Information Needed] Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/tokenization_robustness_v102.tabularmultiple-choicen<1K2 likes227 downloads1y agoHugging Face08simwit /slake-robustnessimage1K<n<10K0 likes181 downloads1y agoHugging Face09wrynx /probe-robustness-rotten_tomatoes rotten_tomatoes Dataset repo: wrynx/probe-robustness-rotten_tomatoes Auto-generated by prepare_datasets.py. Do not hand-edit -- regenerate by re-running the script (with --force) instead. Stats Total records: 10662 Records per split: test: 1066 train: 8530 valid: 1066 Number of classes: 2 Records per class: 0: 5331 1: 5331 Records per class per split: test: 0: 533 1: 533 train: 0: 4265 1: 4265 valid: 0: 533 1: 533 Original dataset README… See the full description on the dataset page: https://huggingface.co/datasets/wrynx/probe-robustness-rotten_tomatoes.tabular10K<n<100K0 likes168 downloads22d agoHugging Face10physicl-community /lighting-invariant-bedroom-perception-robustness-benchmark-next-pack-1917c2cb-f1975230 Home Object Detection, Grasping and Sorting Eval — YOLOv8 Evaluation dataset for a robotic arm that detects, grasps, and sorts objects by type in home environments. 30 renders at 640x640 across kitchen, entry, living room and dressing spaces, staged with everyday household objects. Includes RGB plus metric depth, world-space normals (OpenGL, linear), albedo and material index passes, per-frame annotations, and midday lighting. Targets a YOLOv8 model. This dataset mirrors public… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/lighting-invariant-bedroom-perception-robustness-benchmark-next-pack-1917c2cb-f1975230.imagen<1K0 likes163 downloads20d agoHugging Face11argo11 /openvla-oft-plus-robustness-recovery OpenVLA-OFT+ Robustness and Recovery Dataset 成功した操作だけを増やせば、ロボット方策は頑健になるに違いない。 しかし、OpenVLA-OFT+の挙動を条件別に確かめると、失敗は一様ではなかった。 Spatialではsafe successが63.3%まで下がり、collision rateは35.0%に達した。 Goalのsafe success 90.0%、Objectの95.0%と比べても、空間関係を扱う操作の不安定さが際立つ。 このデータセットは、その差を学習データへ戻すために作成した。 遠距離の空間操作に加え、把持に失敗した状態、持ち上げられなかった状態、運搬中に物体を落とした状態から、成功まで立て直す軌跡を収録している。 弱点の観測 各suiteは20 scene、3条件、合計60エピソードで測定した。 95%信頼区間はscene単位のcluster bootstrap(10,000回)で算出している。 Suite Episodes Safe… See the full description on the dataset page: https://huggingface.co/datasets/argo11/openvla-oft-plus-robustness-recovery.tabularrobotics1K<n<10K1 likes155 downloads2mo agoHugging Face12wrynx /probe-robustness-sms_spam sms_spam Dataset repo: wrynx/probe-robustness-sms_spam Auto-generated by prepare_datasets.py. Do not hand-edit -- regenerate by re-running the script (with --force) instead. Stats Total records: 5159 Records per split: test: 804 train: 3604 valid: 751 Number of classes: 2 Records per class: ham: 4517 spam: 642 Records per class per split: test: ham: 708 spam: 96 train: ham: 3143 spam: 461 valid: ham: 666 spam: 85 Original dataset README… See the full description on the dataset page: https://huggingface.co/datasets/wrynx/probe-robustness-sms_spam.text1K<n<10K0 likes133 downloads22d agoHugging Face13johnbean393 /chiboard-1.1-robustness-sft Chiboard 1.1 robustness supplemental SFT — Plan 04.6 candidate This immutable candidate successor preserves the Plan 04.5 rows and adds the document-unique Plan 04.6 capacity build. Use the plan04_6_eligible config for the one-row-per-source-document candidate arm. The legacy rows remain available for provenance and historical comparison but are not eligible in the maximum Plan 05 schedule. The prompt contract remains qwen35-chiboard-field-tokens-v2; completion alone receives… See the full description on the dataset page: https://huggingface.co/datasets/johnbean393/chiboard-1.1-robustness-sft.tabular1M<n<10M0 likes129 downloads19d agoHugging Face14nvidia /Nemotron-RL-Jailbreak-Robustness-v1 Dataset Description: The Nemotron-RL-Jailbreak-Robustness-v1 data is designed to (1) strengthen model robustness against a variety of adversarial jailbreak techniques and (2) at the same time improve adherence to behavioral policies. This dataset is a collection of hybrid (open-source and synthetically generated) collection of adversarial prompts designed to elicit undesirable behavior from large language models. That's it, just prompts, responses are generated during training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Jailbreak-Robustness-v1.textreinforcement-learning1K<n<10K1 likes119 downloads4mo agoHugging Face15Eloquent /Robustness ELOQUENT Robustness and Consistency Task This dataset contains the sample and test datasets for the Robustness and Consistency task, which is part of the ELOQUENT lab. This dataset is for participants to generate texts for prompt variants, to investigate prompt style conditioned variation. Robustness task ELOQUENT lab CLEF conference 9-12 September 2025 The task in brief (this is a simple task to execute!) This dataset provides a number of questions in several… See the full description on the dataset page: https://huggingface.co/datasets/Eloquent/Robustness.textn<1K0 likes113 downloads1y agoHugging Face16simwit /omni-med-vqa-mini-robustnessimage10K<n<100K0 likes106 downloads1y agoHugging Face17simwit /pmc-vqa-robustnessimage10K<n<100K0 likes102 downloads1y agoHugging Face18siddharthmb /2026.RA.Five-Seat-Qwen3-8B-Robustness Five-Seat Qwen3-8B Private Robustness Subset ⚠ ERRATUM (2026-08-10) — the +0.449 one-oracle result is measured on a spoiled ballot OmniscientBestResponsePolicy, the computable seat in the one-oracle arm, cast its forced-final vote on whichever live offer it valued most instead of on the one under the up/down vote; the protocol rejected that as a legality error and the turn was recorded as a silent pass. Fixed in commit ca20157 (2026-08-10), after these episodes… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Five-Seat-Qwen3-8B-Robustness.tabularn<1K0 likes92 downloads1mo agoHugging Face19GAMI000 /repro-learning-rate-annealing-improves-tuning-robustness-in-stochastic-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabular1K<n<10K0 likes85 downloads2mo agoHugging Face20simwit /vqa-med-robustnessimage1K<n<10K0 likes77 downloads1y agoHugging Face21barbarabhb /nl2sh-chatter-robustness Chatter / robustness pairs for NL->shell models 246 hand-written (natural language, shell command) pairs teaching the "boring reflex": greetings, small talk, identity questions and nonsense input map to harmless commands (echo hello, pwd) instead of garbage or network-touching behavior. Generated by organic_augment.py (deterministic, seed 42). Used in the training pool of barbarabhb/nl2sh-qwen25-coder-1.5b-GGUF. texttext-generationn<1K0 likes76 downloads29d agoHugging Face22wrynx /probe-robustness-truthfulqa truthfulqa Dataset repo: wrynx/probe-robustness-truthfulqa Auto-generated by prepare_datasets.py. Do not hand-edit -- regenerate by re-running the script (with --force) instead. Stats Total records: 817 Records per split: test: 114 train: 591 valid: 112 Label column: not present / not populated for this dataset. Original dataset README (from truthfulqa/truthful_qa) Reproduced here from the source dataset's own card (license, citation, task… See the full description on the dataset page: https://huggingface.co/datasets/wrynx/probe-robustness-truthfulqa.textn<1K0 likes70 downloads24d agoHugging Face23i-am-shaurya05 /robotrace-vla-robustness-traces RoboTrace Evidence Bundle This dataset repository contains the public evidence bundle for RoboTrace, a low-cost deployment-stress evaluation scaffold for robot-learning and VLA-style inference pipelines. The current release evaluates lerobot/pusht and includes reports, metrics, plots, summaries, and release manifests from a complete staged run. What this bundle is for Use this repository to inspect evidence from RoboTrace: action-trace stability metrics visual… See the full description on the dataset page: https://huggingface.co/datasets/i-am-shaurya05/robotrace-vla-robustness-traces.imageroboticsn<1K0 likes66 downloads3mo agoHugging Face24Auenchanters /repro-towards-optimal-robustness-in-learning-augmented-paging-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes62 downloads2mo agoHugging Face25r-three /farsi_tokenizer_robustness TokSuite Benchmark (Farsi Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Farsi (Persian) language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness. Curated by: R3 Research Team Language(s): Farsi/Persian (fa) License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/r-three/farsi_tokenizer_robustness.tabularmultiple-choicen<1K1 likes59 downloads11mo agoHugging Face26Malikeh1375 /tokenizer-robustness-mmlu Tokenizer Robustness MMLU Dataset This dataset contains MMLU-formatted questions and answers designed to test tokenizer robustness across different text formats and languages. Dataset Description The dataset consists of the same questions presented in 6 different formats, with both test (20 questions) and development (5 questions) sets: original - Standard formatted questions minor_spelling_errors - Questions with minor misspellings spoken_language - Questions in casual… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/tokenizer-robustness-mmlu.textn<1K0 likes53 downloads1y agoHugging Face27hari-krishna-ai /text-to-sql-phrasing-robustness Does sloppy phrasing break text-to-SQL? The enterprise text-to-SQL benchmark lists its own biggest caveat: every question is template-generated, so real user phrasing is untested. This is the test. 35 test questions (one per template), each sent to the deployed pipeline four ways: as written, with a typo, in business shorthand, and stripped to a terse fragment. 24 questions and 85 answers survive the filter described under Setup; every answer was executed against the database.… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-phrasing-robustness.tabulartext-generationn<1K0 likes52 downloads5d agoHugging Face28wrynx /probe-robustness-ag_news ag_news Dataset repo: wrynx/probe-robustness-ag_news Auto-generated by prepare_datasets.py. Do not hand-edit -- regenerate by re-running the script (with --force) instead. Stats Total records: 127600 Records per split: test: 25644 train: 84108 valid: 17848 Number of classes: 4 Records per class: 0: 31900 1: 31900 2: 31900 3: 31900 Records per class per split: test: 0: 6451 1: 6412 2: 6312 3: 6469 train: 0: 20968 1: 21030 2: 21078 3: 21032 valid: 0: 4481… See the full description on the dataset page: https://huggingface.co/datasets/wrynx/probe-robustness-ag_news.tabular100K<n<1M0 likes51 downloads24d agoHugging Face29Reza-Telus /certainty-robustness-llm-evaluation Certainty Robustness Benchmark This repository accompanies the paper: Certainty robustness: Evaluating LLM stability under self-challenging promptsMohammadreza Saadat, Steve NemzerarXiv:2603.03330, 2026https://arxiv.org/abs/2603.03330 Overview The Certainty Robustness Benchmark evaluates how large language models (LLMs) behave when their initial answers are challenged by follow-up prompts such as: “Are you sure?” “You are wrong!” confidence elicitation prompts Rather… See the full description on the dataset page: https://huggingface.co/datasets/Reza-Telus/certainty-robustness-llm-evaluation.documentn<1K0 likes49 downloads6mo agoHugging Face30sarvamai /tts-robustness-benchmark TTS Robustness Benchmark Overview The TTS Robustness Benchmark is a set of evaluation samples used to calculate domain-wise CER (Character Error Rate) scores across 7 critical stress-test categories. This benchmark ensures that each sentence-language pair appears exactly once, providing a clean and reliable metric for TTS robustness. Blog Post: Read the official announcement Column Descriptions Column Description text The original input… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/tts-robustness-benchmark.texttext-to-speechn<1K6 likes46 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.