CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LLMSafety /MaliciousInstructtextn<1K0 likes257 downloads7mo agoHugging Face02jang1563 /llm-sfm-safety-eval LLM x SFM Safety Evaluation When a general-purpose language model interprets the output of a specialist science foundation model (a protein, genomic, RNA, or chemistry model), does its safety behavior recognize the scientific content, or only the surface form of the request? This repository is the empirical core of a study of that question: the evaluation harness, the redacted aggregate results, and the measurement specifications behind four findings about how deployed Claude… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/llm-sfm-safety-eval.texttext-classification10K<n<100K0 likes202 downloads13d agoHugging Face03Uranus /llm-safety-instruct-v7text100K<n<1M0 likes133 downloads2y agoHugging Face04llm-semantic-router /mlcommons-ai-safety-synth MLCommons AI Safety Synthesized Dataset Synthesized training data for AI safety classifiers based on the MLCommons AI Safety Hazard Taxonomy. Dataset Description This dataset contains 12,000 synthesized unsafe prompts across 6 hazard categories, designed to augment training data for content safety classifiers. Each category contains 2,000 balanced samples. Hazard Categories (MLCommons AI Safety Taxonomy) Category Description Samples… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/mlcommons-ai-safety-synth.texttext-classification10K<n<100K1 likes104 downloads8mo agoHugging Face05Charly-X /bayesian-llm-safety-inference Bayesian Latent Safety-Trait Dataset Summary This dataset supports Bayesian latent-trait analysis of language-model safety behavior. It contains 90 benchmark-derived roots, three matched prompt variants per root, responses from four target models over five runs, two independent LLM ratings per response, and one human rating for a stratified 540-response calibration subset. The three dimensions are harmful compliance, sycophancy, and agentic protocol violation.… See the full description on the dataset page: https://huggingface.co/datasets/Charly-X/bayesian-llm-safety-inference.tabulartext-classification10K<n<100K0 likes76 downloads2mo agoHugging Face06TrustSafeAI /llm_physical_safety_benchmark LLM Physical Safety Benchmark in Drone Control This benchmark consists of four datasets designed to evaluate the performance of Large Language Models (LLMs) in controlling drones and their vulnerability to physical attacks. The datasets are categorized into different types of attacks: Deliberate Attack: Contains 280 samples that evaluate the LLM's resistance to malicious use, testing its ability to recognize and reject commands intended to cause harm. Subcategories include Direct… See the full description on the dataset page: https://huggingface.co/datasets/TrustSafeAI/llm_physical_safety_benchmark.textquestion-answeringn<1K0 likes44 downloads2y agoHugging Face07kumitang /llm_physical_safety_benchmark LLM Physical Safety Benchmark in Drone Control This benchmark consists of four datasets designed to evaluate the performance of Large Language Models (LLMs) in controlling drones and their vulnerability to physical attacks. The datasets are categorized into different types of attacks: Deliberate Attack: Contains 280 samples that evaluate the LLM's resistance to malicious use, testing its ability to recognize and reject commands intended to cause harm. Subcategories include Direct… See the full description on the dataset page: https://huggingface.co/datasets/kumitang/llm_physical_safety_benchmark.textquestion-answeringn<1K0 likes40 downloads2y agoHugging Face08LLMSafety /SorryBenchtextn<1K0 likes38 downloads2mo agoHugging Face09LLMSafety /HarmBenchtextn<1K0 likes34 downloads7mo agoHugging Face10sadecebirisii /turkish-llm-authority-bypass-safety-sft Turkish LLM Safety Dataset — Authority & System Command Bypass Refusal Kod adı: TR-Auth-Bypass-Refusal-v1 Dil: Türkçe (tr) Format: Hugging Face / Unsloth chat template uyumlu 🇹🇷 Türkçe Açıklama Amaç Bu veri seti, büyük dil modellerinin (LLM) güvenlik bariyerlerini (guardrails) aşmaya yönelik yetki süistimali ve sistem komutu bypass saldırılarını tespit edip güvenli biçimde reddetmesi için hazırlanmış bir Supervised Fine-Tuning (SFT) veri setidir.… See the full description on the dataset page: https://huggingface.co/datasets/sadecebirisii/turkish-llm-authority-bypass-safety-sft.texttext-generationn<1K0 likes26 downloads2mo agoHugging Face11LLMSafety /XSTesttextn<1K0 likes25 downloads4mo agoHugging Face12TianfuXinqu /pwc747_a10865__paper__P02__2023__high__llm_safety Northwind Support Tickets Archive A derived dataset combining service interaction logs with survey responses for support ticket analysis. Upstream Sources This dataset is derived from the following upstream source datasets: Northwind Service Interaction Logs (TianfuXinqu/pwc747_a10865__paper__P05__2022__high__llm_safety) Northwind Customer Survey Responses (TianfuXinqu/pwc747_a10865__paper__P06__2018__low__federated_learning) Commercial Use… See the full description on the dataset page: https://huggingface.co/datasets/TianfuXinqu/pwc747_a10865__paper__P02__2023__high__llm_safety.textn<1K0 likes25 downloads1mo agoHugging Face13Uranus /llm-safety-dpo-v3text10K<n<100K0 likes23 downloads2y agoHugging Face14LLMSafety /or-benchtext1K<n<10K0 likes23 downloads2mo agoHugging Face15Uranus /llm-safety-instruct-v6text100K<n<1M1 likes19 downloads2y agoHugging Face16LLMSafety /NuminaMath-CoT-100ktext100K<n<1M0 likes19 downloads5mo agoHugging Face17LLMSafety /OpenCodeInstruct-50ktext10K<n<100K0 likes18 downloads5mo agoHugging Face18amalia-llm /amalia-Nemotron-SFT-Safety-v1 AMALIA Nemotron-SFT-Safety-v1 Version of the nvidia/Nemotron-SFT-Safety-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage. This dataset went through a processing pipeline to remove entries that reference other LLMs or research labs; Original Dataset: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Safety-v1 This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model. Citation… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SFT-Safety-v1.texttext-generation10K<n<100K0 likes18 downloads3mo agoHugging Face19LLMSafety /hh-dpotext10K<n<100K0 likes16 downloads7mo agoHugging Face20LLMSafety /ArenaHardtextn<1K0 likes15 downloads2mo agoHugging Face21LLMSafety /MT-Benchtextn<1K0 likes12 downloads2mo agoHugging Face22LLMSafety /sorry-bench-roletextn<1K0 likes11 downloads4mo agoHugging Face23APTO-001 /llm-safety-japanese-multiturn-datasetgated株式会社APTOはLLMの安全性の性能を改善させるためのデータセットの開発を行いました。 近年、LLMの性能は飛躍的に向上しており、また、LLMの安全性は長らく注目を集めています。例えば、Anthropic社のClaudeでは、不正な入力に対しClaudeを保護する憲法AIが導入されているなど、安全性は世界的にも重要視されています。※1 しかし、今でも課題視されている点もあり、例えばGPT-5においても特定の条件で安全な会話ができなくなるケースも見られています。※2 日本国内でも例外なくLLMの安全性に対して課題感を持ちながらも安全性の向上への取り組みが見られます。 そのような国内のニーズに答えるべく、日本語で構成された安全性向上のためのデータセットを公開いたしました。 データの内容 データセットの構成として、ターン数がかさむほど安全性が劣化するという点に着目して、マルチターンのデータセットを作成しました。以下のサンプルのようなフォーマットのデータ構成となっております。 { "question_turn1":… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/llm-safety-japanese-multiturn-dataset.textn<1K0 likes10 downloads11mo agoHugging Face24LLMSafety /AdvBenchtextn<1K0 likes9 downloads7mo agoHugging Face25sabaridsnfuji /repro-a-coin-flip-for-safety-llm-judges-fail-to-reliably-measure-adversarial-robustness Reproduction: A Coin Flip for Safety - LLM Judges Fail to Reliably Measure Adversarial Robustness Paper Information Title: A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness OpenReview ID: RgnoWsmYBM Conference: ICML 2026 Task: Evaluate 4 LLM-based safety judges on 6,642 human-verified adversarial prompts Reproduction Summary This paper's all 6 claims require LLM inference on proprietary adversarial prompts and… See the full description on the dataset page: https://huggingface.co/datasets/sabaridsnfuji/repro-a-coin-flip-for-safety-llm-judges-fail-to-reliably-measure-adversarial-robustness.textn<1K0 likes7 downloads2mo agoHugging Face26umairpy /LLM-AI-Safety-Response-Classification ASRCD — AI Safety Response Classification Dataset Understanding AI Decision Making, Harm Detection, and Response Strategy in Real-World LLM Interactions Author: Umair SaeedVersion: 1.0Total Rows: 1,000Format: CSVLanguage: EnglishTask Type: Multi-label Text ClassificationLicense: Research and Educational Use Only About Dataset ASRCD — AI Safety Response Classification Dataset Overview This dataset is designed to train and evaluate… See the full description on the dataset page: https://huggingface.co/datasets/umairpy/LLM-AI-Safety-Response-Classification.texttext-classification1K<n<10K0 likes6 downloads3mo agoHugging Face27FreDERicia /LLM_SAFETY_CHECK_OODgatedtext10M<n<100M0 likes1 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.