CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01thu-coai /ShieldVLMimage1K<n<10K1 likes486 downloads1y agoHugging Face02shields /catalan_commonvoice Dataset Card for "catalan_commonvoice" More Information needed audio100K<n<1M1 likes412 downloads4y agoHugging Face03jsl5710 /Shield DIA-GUARD — dia_splits Canonical train/val/test splits for the DIA-GUARD safety-guard training pipeline. Generated on 2026-03-30 | Seed: 42 | Ratios: 70 / 15 / 15 The split files are hosted on HuggingFace: https://huggingface.co/datasets/jsl5710/Shield Download via the HuggingFace Hub: from huggingface_hub import snapshot_download snapshot_download(repo_id="jsl5710/Shield", repo_type="dataset", local_dir="dataset/dia_splits") Or with the CLI: huggingface-cli download… See the full description on the dataset page: https://huggingface.co/datasets/jsl5710/Shield.texttext-classification1M<n<10M0 likes372 downloads6mo agoHugging Face04eminorhan /shieldtext1K<n<10K0 likes322 downloads1y agoHugging Face05shields /catalan_commonvoice_first15hr_processed Dataset Card for "catalan_commonvoice_first15hr_processed" More Information needed 10K<n<100K0 likes170 downloads4y agoHugging Face06auren-research /pii-shield PII Shield: Multilingual PII Detection Dataset PII Shield is a large-scale, multilingual dataset for training and evaluating Personally Identifiable Information (PII) detection models. Built by Auren Research, it combines real-world documents from diverse domains with high-quality span-level PII annotations produced by fastino/gliner2-privacy-filter-PII-multi— achieving the highest F1 on the SPY benchmark among open-source PII detectors. The dataset is designed to… See the full description on the dataset page: https://huggingface.co/datasets/auren-research/pii-shield.texttoken-classification1M<n<10M2 likes125 downloads4mo agoHugging Face07shields /catalan_commonvoice_first15hr_processed_with_noise Dataset Card for "catalan_commonvoice_first15hr_processed_with_noise" More Information needed 10K<n<100K0 likes117 downloads4y agoHugging Face08dmilush /shieldlm-prompt-injection ShieldLM Prompt Injection Dataset A unified prompt injection detection dataset with 54,162 samples spanning three attack categories: direct injection, indirect injection, and jailbreak. Curated from 11 source datasets with a 3-level hierarchical label schema. Dataset Description Purpose Training and evaluating prompt injection classifiers for production deployment. Designed to address gaps in existing datasets: Indirect injection coverage (via InjecAgent… See the full description on the dataset page: https://huggingface.co/datasets/dmilush/shieldlm-prompt-injection.texttext-classification10K<n<100K1 likes112 downloads7mo agoHugging Face09uwsthoughts /dolly_shield Dataset Card for Project Dolly Shield This is a collection of Beatport and Spotify data that I found on Kaggle. No additional Spotify or Beatport data was collected from their platforms directly. Spotify bans the use of their data for training AI models. I've decided I can use a dataset from Kaggle for my AI project but will not collect additional data from Spotify via their API services for AI model training. More information on Spotify's AI policies can be found here. The… See the full description on the dataset page: https://huggingface.co/datasets/uwsthoughts/dolly_shield.tabular10M<n<100M0 likes102 downloads2y agoHugging Face10QLWD /code_shieldtext100K<n<1M0 likes95 downloads2y agoHugging Face11wordsum /for-the-small-shield-chapters Foreword The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster. I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.tabulartext-retrieval1K<n<10K0 likes94 downloads1mo agoHugging Face12Khalilah-Shields /MSU-Benchmark MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto. Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie** Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/Khalilah-Shields/MSU-Benchmark.audioaudio-classification1K<n<10K0 likes81 downloads2mo agoHugging Face13TonyYun /pii-shield-benchmark PII Shield Benchmark Five public PII-detection datasets rewritten into one record shape, so models can be trained and evaluated across all of them without writing five parsers. 2,301,892 records carrying 13,338,839 labeled spans in 25 European languages. Nothing here is new text. Every document comes unchanged from one of the five source datasets below; this repository harmonizes the containers — one schema, one label taxonomy, one file format (Parquet, zstd) — and publishes a… See the full description on the dataset page: https://huggingface.co/datasets/TonyYun/pii-shield-benchmark.token-classification1M<n<10M0 likes81 downloads1mo agoHugging Face14shields /whisper-small-hindi Dataset Card for "whisper-small-hindi" More Information needed audio1K<n<10K6 likes65 downloads4y agoHugging Face15woozziam /depth_png_shieldvideon<1K0 likes61 downloads3mo agoHugging Face16ata8e /SHIELDINGThis dataset was generated as a part of our paper "Towards Secure Prompt Processing: A Unified Framework to Detect, Sanitize, and Prevent Adversarial Prompts via Dual-Optimized Threshold-Aware Learning". SHIELDING dataset was generated by combining and filtering the following datasets: Prompt Injection Hackaprompt GPT35, https://huggingface.co/datasets/imoxto/prompt_injection_hackaprompt_gpt35. JailBreak V-28k, https://huggingface.co/datasets/JailbreakV-28K/JailBreakV-28k. Open Platypus… See the full description on the dataset page: https://huggingface.co/datasets/ata8e/SHIELDING.texttext-classification10K<n<100K0 likes59 downloads3mo agoHugging Face17Abdennebi /shieldlm-prompt-injection ShieldLM Prompt Injection Dataset A unified prompt injection detection dataset with 54,162 samples spanning three attack categories: direct injection, indirect injection, and jailbreak. Curated from 11 source datasets with a 3-level hierarchical label schema. Dataset Description Purpose Training and evaluating prompt injection classifiers for production deployment. Designed to address gaps in existing datasets: Indirect injection coverage (via InjecAgent… See the full description on the dataset page: https://huggingface.co/datasets/Abdennebi/shieldlm-prompt-injection.texttext-classification10K<n<100K0 likes57 downloads7mo agoHugging Face18ShieldX /Rash-Driving-Detection-on-Bikes-for-ML Dataset Card for Rash Driving Detection on Bikes Using Mobile and Sensor Data This dataset is designed to aid the detection of rash driving behavior on bikes using data collected from mobile and sensor-based systems. It includes sensor readings such as accelerometer values, orientation (azimuth, pitch, roll), and speed, with labels indicating whether the riding behavior is classified as rash or not. Dataset Details Dataset Description This dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/ShieldX/Rash-Driving-Detection-on-Bikes-for-ML.tabular10K<n<100K0 likes43 downloads9mo agoHugging Face19axolotl-ai-co /text-vision-shieldstral-2k-testA 2k sample dataset for testing the Shieldstral multimodal moderation format. Each sample is a fixed system prompt, a [text, image, text] user message, and a single yes/no answer. Load in Axolotl via: datasets: - path: Nanobit/text-vision-shieldstral-2k-test type: chat_template Make sure to download the image via: wget https://huggingface.co/datasets/Nanobit/text-vision-shieldstral-2k-test/resolve/main/African_elephant.jpg Image source:… See the full description on the dataset page: https://huggingface.co/datasets/axolotl-ai-co/text-vision-shieldstral-2k-test.text1K<n<10K0 likes37 downloads2mo agoHugging Face20shields /catalan_commonvoice_15_beforeMappingaudio1M<n<10M0 likes33 downloads2y agoHugging Face21ShieldX /Context-Aware-Repository-Prompt-Injection Overview This dataset is designed for training and evaluating AI security scanners that detect repository-aware prompt injection attacks in software development and code-assistant environments. Repository-aware prompt injections are malicious instructions embedded in code repositories, documentation, comments, configuration files, issue trackers, or other project artifacts that attempt to manipulate an AI system's behavior, override its instructions, exfiltrate sensitive… See the full description on the dataset page: https://huggingface.co/datasets/ShieldX/Context-Aware-Repository-Prompt-Injection.text1K<n<10K1 likes28 downloads4mo agoHugging Face22smolify /smolified-sentinel-privacy-shield 🤏 smolified-sentinel-privacy-shield Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model smolify/smolified-sentinel-privacy-shield. 📦 Asset Details Origin: Smolify Foundry (Job ID: 41a1525a) Records: 1440 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by smolify. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes26 downloads6mo agoHugging Face23open-llm-leaderboard-old /details_ShieldX__manovyadh-1.1B-v1-chat Dataset Card for Evaluation run of ShieldX/manovyadh-1.1B-v1-chat Dataset automatically created during the evaluation run of model ShieldX/manovyadh-1.1B-v1-chat on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ShieldX__manovyadh-1.1B-v1-chat.0 likes25 downloads3y agoHugging Face24hirushafernando /slm-shield-ds-role-and-instruction-violationtext10K<n<100K0 likes24 downloads3mo agoHugging Face25ata8e /SHIELDING_LLaMA-3_ResponseThis dataset is a part of our SHIELDING dataset that progressed to LLaMA-3 and checked each response humanly, as a part of our paper "Towards Secure Prompt Processing: A Unified Framework to Detect, Sanitize, and Prevent Adversarial Prompts via Dual-Optimized Threshold-Aware Learning". texttext-classification10K<n<100K0 likes21 downloads3mo agoHugging Face26hirushafernando /slm-shield-ds-obfuscation-and-evasion-patternstext10K<n<100K0 likes21 downloads3mo agoHugging Face27axolotl-ai-co /text-shieldstral-2k-testA 2k sample dataset for testing the Shieldstral text moderation format. Each sample is a fixed system prompt, an <Instruct>/<Query>/<Document> user message, and a single yes/no answer. Load in Axolotl via: datasets: - path: Nanobit/text-shieldstral-2k-test type: chat_template Derived from PKU-Alignment/BeaverTails (30k_train, shuffled with seed 42), mapping its is_safe flag to the answer. Inherits its CC BY-NC 4.0 license. BeaverTails labels are noisy, so this is a format/smoke test… See the full description on the dataset page: https://huggingface.co/datasets/axolotl-ai-co/text-shieldstral-2k-test.text1K<n<10K1 likes20 downloads2mo agoHugging Face28Jumbol /ShieldBreaker_Benchmark_Datasetgated ShieldBreaker Benchmark Dataset Overview The ShieldBreaker Benchmark Dataset is a comprehensive collection of anti-CRISPR protein sequences and structures, designed for machine learning research in CRISPR-Cas system inhibition. This dataset contains both positive (anti-CRISPR) and negative (non-anti-CRISPR) samples with dual-modal data representations. Dataset Structure ShieldBreaker_Upload/ ├── positive/ │ ├── fasta/ │ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/Jumbol/ShieldBreaker_Benchmark_Dataset.texttext-classification10K<n<100K0 likes19 downloads9mo agoHugging Face29shields /catalan_commonvoice_first15hr Dataset Card for "catalan_commonvoice_first15hr" More Information needed audio10K<n<100K0 likes18 downloads4y agoHugging Face30ShieldX /manovyadh-3.5ktexttext-classification1K<n<10K1 likes14 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.