CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LIBERO-Safety /libero_safetyvideo10K<n<100K2 likes32k downloads5mo agoHugging Face02PKU-Alignment /PKU-SafeRLHF Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. [🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset] Citation If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.tabulartext-generation100K<n<1M196 likes14k downloads2y agoHugging Face03albertklorer /safedocs-1M-muse-spark-1.3-judged SafeDocs: Muse Spark 1.3 judge annotations Incrementally published, one complete shard per commit. All original source columns, images, complete Paddle JSON, rows and row order are preserved. No language or quality filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status, and judge_error. Operational failures retain the original page with a null verdict and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs. Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.tabular100K<n<1M0 likes12k downloads2d agoHugging Face04czxlovesu03 /tfds_out_safe0 likes8.3k downloads9mo agoHugging Face05nvidia /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K110 likes8.2k downloads1y agoHugging Face06deepghs /safebooru-webp-4Mpixelgated Safebooru 4M Re-encoded Dataset This is the re-encoded dataset of deepghs/safebooru_full. And all the resized images are maintained here. There are 5756655 images in total. The maximum ID of these images is 5974383. Last updated at 2025-08-06 08:31:53 JST. How to Painlessly Use This Use cheesechaser to quickly get images from this repository. Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/safebooru-webp-4Mpixel.image-classification1M<n<10M2 likes7.6k downloads1y agoHugging Face07ai-safety-institute /AgentHarm AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Maksym Andriushchenko1,†,*, Alexandra Souly2,* Mateusz Dziemian1, Derek Duenas1, Maxwell Lin1, Justin Wang1, Dan Hendrycks1,§, Andy Zou1,¶,§, Zico Kolter1,¶, Matt Fredrikson1,¶,* Eric Winsor2, Jerome Wynne2, Yarin Gal2,♯, Xander Davies2,♯,* 1Gray Swan AI, 2UK AI Safety Institute, *Core Contributor †EPFL, §Center for AI Safety, ¶Carnegie Mellon University, ♯University of Oxford Paper: https://arxiv.org/abs/2410.09024… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/AgentHarm.textn<1K62 likes6.7k downloads2y agoHugging Face08safe-autonomous-systems /fluidgym-data0 likes5.5k downloads6mo agoHugging Face09nvidia /Nemotron-AIQ-Agentic-Safety-Dataset-1.0 Nemotron-AIQ Agentic Safety Dataset Dataset Summary Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.texttext-generation10K<n<100K18 likes4.7k downloads10mo agoHugging Face10physicl /indoor-safety-hazard-detection-and-work-zone-monitoring Indoor Safety Hazard Detection & Work-Zone Monitoring Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/indoor-safety-hazard-detection-and-work-zone-monitoring.imagen<1K0 likes4.3k downloads3mo agoHugging Face11albertklorer /safedocs-1M1 likes3.6k downloads1d agoHugging Face12nvidia /Aegis-AI-Content-Safety-Dataset-1.0 🛡️ Nemotron Content Safety Dataset V1 Nemotron Content Safety Dataset V1, formerly known as Aegis AI Content Safety Dataset, is an open-source content safety dataset (CC-BY-4.0), which adheres to Nvidia's content safety taxonomy, covering 13 critical risk categories (see Dataset Description). Dataset Details Dataset Description Nemotron Content Safety Dataset V1 is comprised of approximately 11,000 manually annotated interactions between humans and LLMs, split… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-1.0.texttext-classification10K<n<100K61 likes3.5k downloads1y agoHugging Face13datamol-io /safe-gpt SAFE Molecules Dataset (v2) A large-scale molecular dataset containing approximately 1.17 billion unique molecules, each represented with both canonical SMILES and SAFE (Sequential Attachment-based Fragment Embedding) strings. This dataset is intended to support large-scale pretraining and evaluation of chemical language models, including generative, conditional, and structure-aware modeling tasks. Note This is version 2 of the SAFE dataset. The original v1 release contained… See the full description on the dataset page: https://huggingface.co/datasets/datamol-io/safe-gpt.texttext-generation1B<n<10B4 likes3.1k downloads8mo agoHugging Face14physicl /kitchen-workspace-understanding-safe-manipulation Kitchen Workspace Understanding & Safe Manipulation Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/kitchen-workspace-understanding-safe-manipulation.imagen<1K0 likes3k downloads3mo agoHugging Face15ai-safety-institute /lie-detection-rollouts Lie Detection Rollouts Assistant completions across many open-weight models on the lie-detection evaluation suite used by the deception research pipeline. One subset per model, one split per task. Columns messages — list of OpenAI-style messages. Each message has: role: system | user | assistant content: final message text reasoning_content: chain-of-thought for reasoning models, None otherwise is_lie — ground-truth label from the is_deceptive scorer: lie |… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/lie-detection-rollouts.text1M<n<10M0 likes2.8k downloads3mo agoHugging Face16Voxel51 /Safe_and_Unsafe_Behaviours Dataset Card for safe_unsafe_behaviours This is a FiftyOne dataset with 691 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/Safe_and_Unsafe_Behaviours") # Launch the App session = fo.launch_app(dataset) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Safe_and_Unsafe_Behaviours.videoimage-classificationn<1K3 likes2.7k downloads8mo agoHugging Face17ErnestBeckham /gridstar-safety-data1 likes2.5k downloads9d agoHugging Face18aisingapore /Safety-Toxicity-Detectiongated SEA Toxicity Detection SEA Toxicity Detection evaluates a model's ability to identify toxic content such as hate speech and abusive language in text. It is sampled from MLHSD for Indonesian, TTD for Thai, and ViHSD for Vietnamese. Supported Tasks and Leaderboards SEA Toxicity Detection is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore. Languages Indonesian (id) Thai… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Safety-Toxicity-Detection.texttext-generation1K<n<10K0 likes2.4k downloads9mo agoHugging Face19laion /relaion2B-en-research-safegatedimage1B<n<10B226 likes2.3k downloads2y agoHugging Face20xTRam1 /safe-guard-prompt-injectionWe formulated the prompt injection detector problem as a classification problem and trained our own language model to detect whether a given user prompt is an attack or safe. First, to train our own prompt injection detector, we required high-quality labelled data; however, existing prompt injection datasets were either too small (on the magnitude of O(100)) or didn’t cover a broad spectrum of prompt injection attacks. To this end, inspired by the GLAN paper, we created a custom synthetic… See the full description on the dataset page: https://huggingface.co/datasets/xTRam1/safe-guard-prompt-injection.text10K<n<100K34 likes2.2k downloads2y agoHugging Face21xing-shadow /Safety-helmet-datasetimageobject-detection10K<n<100K0 likes2.1k downloads4mo agoHugging Face22asatheesh /latent-mas-safety-dataset-seq-qwen3-4b LatentMAS Safety Dataset — Phase 0 Latent states, model completions, and safety labels from a Qwen3-4B latent-MAS pipeline (Planner → Critic → Refiner → Judger, inter-agent messages passed as hidden-state vectors) evaluated on prompts from public safety benchmarks. Intended for training a latent safety value model and for probing / interpretability work on multi-agent latent reasoning. What's in it 195,589 rollouts from 14 prompt sources, greedy decode… See the full description on the dataset page: https://huggingface.co/datasets/asatheesh/latent-mas-safety-dataset-seq-qwen3-4b.100K<n<1M0 likes1.9k downloads1mo agoHugging Face23PKU-Alignment /MM-SafetyBenchWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements. Usage and License Notices: The dataset is intended and licensed for research use only. They are also restricted to uses that follow the license agreement GPT-4 and Stable Diffusion. The dataset is CC BY NC 4.0 (allowing only non-commercial use). Data Source: For more information about the dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/MM-SafetyBench.image1K<n<10K8 likes1.9k downloads2y agoHugging Face24locuslab /safewebtext10M<n<100M3 likes1.7k downloads4mo agoHugging Face25AILabDsUnipi /SafeQIL-dataset Human-Generated Demonstrations for Safe Reinforcement Learning Paper: Learning to maintain safety through expert demonstrations in settings with unknown constraints: A Q-learning perspective Code: AILabDsUnipi/SafeQIL Dataset Description This dataset consists of human-generated demonstrations collected across four challenging constrained environments from the Safety-Gymnasium benchmark (SafetyPointGoal1-v0, SafetyCarPush2-v0, SafetyPointCircle2-v0, and… See the full description on the dataset page: https://huggingface.co/datasets/AILabDsUnipi/SafeQIL-dataset.reinforcement-learning100K<n<1M0 likes1.7k downloads6mo agoHugging Face26PKU-Alignment /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K62 likes1.6k downloads3y agoHugging Face27phat06 /Safety-helmet-datasetimageobject-detection1K<n<10K2 likes1.4k downloads10mo agoHugging Face28laion /relaion2B-multi-research-safegatedimage1B<n<10B48 likes1.3k downloads2y agoHugging Face29PKU-Alignment /PKU-SafeRLHF-30K Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. Dataset Summary The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-30K.tabulartext-generation10K<n<100K14 likes1.2k downloads3y agoHugging Face30Snowstorm1492 /SafeIMG &nbsp;SafeIMG AI-generated Images Challenge Visual Trust in High-risk Scenarios Yi-Zhi Wang1,2, Yichen Xiao1,2, Linan Yue1,2, Weibo Gao3, Yichao Du4, Pengfei Fang1,2, Shimin Di1,2, Min-Ling Zhang1,2 1 Southeast University &nbsp;&nbsp; 2 Key Laboratory of Computer Network and Information Integration, Ministry of Education3 The Hong Kong Polytechnic University &nbsp;&nbsp; 4 School of Artificial Intelligence, Wuhan University Overview · Dataset · Results · Quick Start ·… See the full description on the dataset page: https://huggingface.co/datasets/Snowstorm1492/SafeIMG.image1K<n<10K0 likes1.2k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.