datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ShieldVLMcatalan_commonvoice
Dataset Card for "catalan_commonvoice"
More Information needed
Shield
DIA-GUARD — dia_splits
Canonical train/val/test splits for the DIA-GUARD safety-guard training pipeline.
Generated on 2026-03-30 | Seed: 42 | Ratios: 70 / 15 / 15
The split files are hosted on HuggingFace:
https://huggingface.co/datasets/jsl5710/Shield
Download via the HuggingFace Hub:
from huggingface_hub import snapshot_download
snapshot_download(repo_id="jsl5710/Shield", repo_type="dataset", local_dir="dataset/dia_splits")
Or with the CLI:
huggingface-cli download… See the full description on the dataset page: https://huggingface.co/datasets/jsl5710/Shield.shieldcatalan_commonvoice_first15hr_processed
Dataset Card for "catalan_commonvoice_first15hr_processed"
More Information needed
pii-shield
PII Shield: Multilingual PII Detection Dataset
PII Shield is a large-scale, multilingual dataset for training and evaluating Personally Identifiable Information (PII) detection models. Built by Auren Research, it combines real-world documents from diverse domains with high-quality span-level PII annotations produced by fastino/gliner2-privacy-filter-PII-multi— achieving the highest F1 on the SPY benchmark among open-source PII detectors.
The dataset is designed to… See the full description on the dataset page: https://huggingface.co/datasets/auren-research/pii-shield.catalan_commonvoice_first15hr_processed_with_noise
Dataset Card for "catalan_commonvoice_first15hr_processed_with_noise"
More Information needed
shieldlm-prompt-injection
ShieldLM Prompt Injection Dataset
A unified prompt injection detection dataset with 54,162 samples spanning three attack categories: direct injection, indirect injection, and jailbreak. Curated from 11 source datasets with a 3-level hierarchical label schema.
Dataset Description
Purpose
Training and evaluating prompt injection classifiers for production deployment. Designed to address gaps in existing datasets:
Indirect injection coverage (via InjecAgent… See the full description on the dataset page: https://huggingface.co/datasets/dmilush/shieldlm-prompt-injection.dolly_shield
Dataset Card for Project Dolly Shield
This is a collection of Beatport and Spotify data that I found on Kaggle.
No additional Spotify or Beatport data was collected from their platforms directly. Spotify bans the use of their data for training AI models. I've decided I can use a dataset from Kaggle for my AI project but will not collect additional data from Spotify via their API services for AI model training.
More information on Spotify's AI policies can be found here.
The… See the full description on the dataset page: https://huggingface.co/datasets/uwsthoughts/dolly_shield.code_shieldfor-the-small-shield-chapters
Foreword
The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster.
I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct
I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.MSU-Benchmark
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto.
Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie**
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China
School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/Khalilah-Shields/MSU-Benchmark.pii-shield-benchmark
PII Shield Benchmark
Five public PII-detection datasets rewritten into one record shape, so models can be
trained and evaluated across all of them without writing five parsers. 2,301,892 records
carrying 13,338,839 labeled spans in 25 European languages.
Nothing here is new text. Every document comes unchanged from one of the five source
datasets below; this repository harmonizes the containers — one schema, one label
taxonomy, one file format (Parquet, zstd) — and publishes a… See the full description on the dataset page: https://huggingface.co/datasets/TonyYun/pii-shield-benchmark.whisper-small-hindi
Dataset Card for "whisper-small-hindi"
More Information needed
depth_png_shieldSHIELDINGThis dataset was generated as a part of our paper "Towards Secure Prompt Processing: A Unified Framework to Detect, Sanitize, and Prevent Adversarial Prompts via Dual-Optimized Threshold-Aware Learning".
SHIELDING dataset was generated by combining and filtering the following datasets:
Prompt Injection Hackaprompt GPT35, https://huggingface.co/datasets/imoxto/prompt_injection_hackaprompt_gpt35.
JailBreak V-28k, https://huggingface.co/datasets/JailbreakV-28K/JailBreakV-28k.
Open Platypus… See the full description on the dataset page: https://huggingface.co/datasets/ata8e/SHIELDING.shieldlm-prompt-injection
ShieldLM Prompt Injection Dataset
A unified prompt injection detection dataset with 54,162 samples spanning three attack categories: direct injection, indirect injection, and jailbreak. Curated from 11 source datasets with a 3-level hierarchical label schema.
Dataset Description
Purpose
Training and evaluating prompt injection classifiers for production deployment. Designed to address gaps in existing datasets:
Indirect injection coverage (via InjecAgent… See the full description on the dataset page: https://huggingface.co/datasets/Abdennebi/shieldlm-prompt-injection.Rash-Driving-Detection-on-Bikes-for-ML
Dataset Card for Rash Driving Detection on Bikes Using Mobile and Sensor Data
This dataset is designed to aid the detection of rash driving behavior on bikes using data collected from mobile and sensor-based systems. It includes sensor readings such as accelerometer values, orientation (azimuth, pitch, roll), and speed, with labels indicating whether the riding behavior is classified as rash or not.
Dataset Details
Dataset Description
This dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/ShieldX/Rash-Driving-Detection-on-Bikes-for-ML.text-vision-shieldstral-2k-testA 2k sample dataset for testing the Shieldstral multimodal moderation format. Each sample is a fixed system prompt, a [text, image, text] user message, and a single yes/no answer.
Load in Axolotl via:
datasets:
- path: Nanobit/text-vision-shieldstral-2k-test
type: chat_template
Make sure to download the image via:
wget https://huggingface.co/datasets/Nanobit/text-vision-shieldstral-2k-test/resolve/main/African_elephant.jpg
Image source:… See the full description on the dataset page: https://huggingface.co/datasets/axolotl-ai-co/text-vision-shieldstral-2k-test.catalan_commonvoice_15_beforeMappingContext-Aware-Repository-Prompt-Injection
Overview
This dataset is designed for training and evaluating AI security scanners that detect repository-aware prompt injection attacks in software development and code-assistant environments.
Repository-aware prompt injections are malicious instructions embedded in code repositories, documentation, comments, configuration files, issue trackers, or other project artifacts that attempt to manipulate an AI system's behavior, override its instructions, exfiltrate sensitive… See the full description on the dataset page: https://huggingface.co/datasets/ShieldX/Context-Aware-Repository-Prompt-Injection.smolified-sentinel-privacy-shield
🤏 smolified-sentinel-privacy-shield
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-sentinel-privacy-shield.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 41a1525a)
Records: 1440
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
details_ShieldX__manovyadh-1.1B-v1-chat
Dataset Card for Evaluation run of ShieldX/manovyadh-1.1B-v1-chat
Dataset automatically created during the evaluation run of model ShieldX/manovyadh-1.1B-v1-chat on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ShieldX__manovyadh-1.1B-v1-chat.slm-shield-ds-role-and-instruction-violationSHIELDING_LLaMA-3_ResponseThis dataset is a part of our SHIELDING dataset that progressed to LLaMA-3 and checked each response humanly, as a part of our paper "Towards Secure Prompt Processing: A Unified Framework to Detect, Sanitize, and Prevent Adversarial Prompts via Dual-Optimized Threshold-Aware Learning".
slm-shield-ds-obfuscation-and-evasion-patternstext-shieldstral-2k-testA 2k sample dataset for testing the Shieldstral text moderation format. Each sample is a fixed system prompt, an <Instruct>/<Query>/<Document> user message, and a single yes/no answer.
Load in Axolotl via:
datasets:
- path: Nanobit/text-shieldstral-2k-test
type: chat_template
Derived from PKU-Alignment/BeaverTails (30k_train, shuffled with seed 42), mapping its is_safe flag to the answer. Inherits its CC BY-NC 4.0 license. BeaverTails labels are noisy, so this is a format/smoke test… See the full description on the dataset page: https://huggingface.co/datasets/axolotl-ai-co/text-shieldstral-2k-test.ShieldBreaker_Benchmark_Dataset
ShieldBreaker Benchmark Dataset
Overview
The ShieldBreaker Benchmark Dataset is a comprehensive collection of anti-CRISPR protein sequences and structures, designed for machine learning research in CRISPR-Cas system inhibition. This dataset contains both positive (anti-CRISPR) and negative (non-anti-CRISPR) samples with dual-modal data representations.
Dataset Structure
ShieldBreaker_Upload/
├── positive/
│ ├── fasta/
│ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/Jumbol/ShieldBreaker_Benchmark_Dataset.catalan_commonvoice_first15hr
Dataset Card for "catalan_commonvoice_first15hr"
More Information needed
manovyadh-3.5k
