datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
log-redaction-trajectories
Log Redaction Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/log-redaction-trajectories.NinjaMasker-PII-RedactionDocPII-redaction-benchmark
DocPII: Contextual Redaction Benchmark Dataset
Dataset Description
DocPII contains 1101 high-quality document samples enriched with embedded personally identifiable information (PII). Designed to evaluate context-aware redaction systems, it provides realistic, full-document contexts—a notable advancement over sentence-level datasets.
All documents have been manually reviewed for accuracy, coherence, and redaction alignment, ensuring data quality for benchmarking and… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/DocPII-redaction-benchmark.redaction-demo
