shield
Datasets
All datasets matching “shield”ShieldVLMcatalan_commonvoice
Dataset Card for "catalan_commonvoice"
More Information needed
Shield
DIA-GUARD — dia_splits
Canonical train/val/test splits for the DIA-GUARD safety-guard training pipeline.
Generated on 2026-03-30 | Seed: 42 | Ratios: 70 / 15 / 15
The split files are hosted on HuggingFace:
https://huggingface.co/datasets/jsl5710/Shield
Download via the HuggingFace Hub:
from huggingface_hub import snapshot_download
snapshot_download(repo_id="jsl5710/Shield", repo_type="dataset", local_dir="dataset/dia_splits")
Or with the CLI:
huggingface-cli download… See the full description on the dataset page: https://huggingface.co/datasets/jsl5710/Shield.shieldcatalan_commonvoice_first15hr_processed
Dataset Card for "catalan_commonvoice_first15hr_processed"
More Information needed
pii-shield
PII Shield: Multilingual PII Detection Dataset
PII Shield is a large-scale, multilingual dataset for training and evaluating Personally Identifiable Information (PII) detection models. Built by Auren Research, it combines real-world documents from diverse domains with high-quality span-level PII annotations produced by fastino/gliner2-privacy-filter-PII-multi— achieving the highest F1 on the SPY benchmark among open-source PII detectors.
The dataset is designed to… See the full description on the dataset page: https://huggingface.co/datasets/auren-research/pii-shield.
