redaction
RedactionBench
Dataset Card for RedactionBench
RedactionBench is an evaluation-only benchmark for character-level redaction across eleven document categories. Each of the 200 documents is manually-annotated with character spans that are either mandatory (must redact) or contextual.
RedactionBench mixes 101 real-world documents manually sourced from the public web (transcribed, augmented) with 99 synthetic documents authored to fill categories where synthetic data is more appropriate.
The above is… See the full description on the dataset page: https://huggingface.co/datasets/RedactionBench/RedactionBench.RedactionBench
Dataset Card for RedactionBench
RedactionBench is an evaluation-only benchmark for character-level redaction across eleven document categories. Each of the 200 documents is manually-annotated with character spans that are either mandatory (must redact) or contextual.
RedactionBench mixes 101 real-world documents manually sourced from the public web (transcribed, augmented) with 99 synthetic documents authored to fill categories where synthetic data is more appropriate.
The… See the full description on the dataset page: https://huggingface.co/datasets/A10Networks/RedactionBench.log-redaction-trajectories
Log Redaction Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/log-redaction-trajectories.RedactionBench
Dataset Card for RedactionBench
RedactionBench is an evaluation-only benchmark for character-level redaction across eleven document categories. Each of the 200 documents is manually-annotated with character spans that are either mandatory (must redact) or contextual.
RedactionBench mixes 101 real-world documents manually sourced from the public web (transcribed, augmented) with 99 synthetic documents authored to fill categories where synthetic data is more appropriate.
The… See the full description on the dataset page: https://huggingface.co/datasets/aibotjock/RedactionBench.NinjaMasker-PII-RedactionNinjaMasker-PII-Redaction-Dataset
