datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stack-exchange-preferences-20230914-clean-anonymization
Dataset Card for "stack-exchange-preferences-20230914-clean-anonymization"
More Information needed
text-anonymization-benchmark
Dataset Card for the Text Anonymization Benchmark (TAB)
Dataset Summary
This repository contains the v1.0 release of the Text Anonymization Benchmark, a corpus for text anonymization.
The corpus comprises 1,268 English-language court cases from the European Court for Human Rights (ECHR). The documents were manually annotated with information about personal identifiers (including their semantic category and need for masking), confidential attributes and co-reference… See the full description on the dataset page: https://huggingface.co/datasets/ildpil/text-anonymization-benchmark.OR_anonymizationtext-anonymization-benchmark-val-test
Dataset card for Text Anonymization Benchmark (TAB) Validation & Test
Dataset Summary
This is the validation and test split of the Text Anonymisation Benchmark.
As the title says it's a dataset focused on text anonymisation, specifcially European Court Documents, which contain labels by mutltiple annotators.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/mattmdjaga/text-anonymization-benchmark-val-test.text-anonymization-benchmark-train
Dataset card for Text Anonymization Benchmark (TAB) train
Dataset Summary
This is the training split of the Text Anonymisation Benchmark.
As the title says it's a dataset focused on text anonymisation, specifcially European Court Documents, which contain labels by mutltiple annotators.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/mattmdjaga/text-anonymization-benchmark-train.face-anonymization-public-evaluation-v58
Origin Data Lab — Face Anonymization Public Evaluation V58
Privacy processing for real-world video datasets
Automated face anonymization combined with targeted Human QA for autonomous driving, robotics, computer vision, urban mobility, and AI data teams.
This repository presents a public engineering evaluation of Origin Data Lab's face anonymization pipeline in dense, low-light urban traffic conditions.
Evaluate Your Own Video
Need privacy processing for traffic… See the full description on the dataset page: https://huggingface.co/datasets/origindatalab/face-anonymization-public-evaluation-v58.GeoR-Bench
GeoR-Bench
GeoR-Bench is an anonymous geoscience image-text benchmark for evaluating multimodal Intelligence on Earth science tasks. It contains 440 instance, where each instance contributes one input+output+prompt. Each instance includes paired input and target images, prompt and judge rubrics spanning satellite imagery, maps, schematics, and other geoscience visual representations.",
Dataset structure
Each benchmark example is stored as a directory containing five… See the full description on the dataset page: https://huggingface.co/datasets/Anonymization-312/GeoR-Bench.anonymization-resumes-datasetanonymization-before-after
Anonymization Before/After
A small paired tabular dataset showing the same records before and after
a 10-step anonymization pipeline. Useful as a teaching fixture for privacy
courses, a benchmark for anonymization toolkits, and a sanity-check input
for red-team / membership-inference experiments.
Important: the PII in sample_raw.csv is entirely synthetic.
Names follow the pattern Person_001, emails are person_001@example.com,
phone numbers are 555-00XX, and "national IDs" are… See the full description on the dataset page: https://huggingface.co/datasets/t22000t/anonymization-before-after.MAPA_Anonymization_package
[!NOTE]
Dataset origin: https://elrc-share.eu/repository/browse/mapa-anonymization-package-french/2769ba3a8a8411ec9c1a00155d0267062553e7eec46c4dec878f6d1cc079f24e/
chessgerman-medical-anonymization-dataset-v0.1text-anonymization-benchmark
Dataset Card for the Text Anonymization Benchmark (TAB)
Dataset Summary
This repository contains the v1.0 release of the Text Anonymization Benchmark, a corpus for text anonymization.
The corpus comprises 1,268 English-language court cases from the European Court for Human Rights (ECHR). The documents were manually annotated with information about personal identifiers (including their semantic category and need for masking), confidential attributes and co-reference… See the full description on the dataset page: https://huggingface.co/datasets/doshimit3015/text-anonymization-benchmark.threevis-aware-anonymization-eval
Evaluation Data for "Visibility-Aware Diffusion-Based Face Anonymization for Real-World Deployment" (ICPR 2026)
This dataset contains the input samples, sample manifests, and generated outputs used to
produce the quantitative results reported in Tables 2, 3, and 4 of the paper. It is
intended to let others reproduce the exact numbers reported in the paper, not to
serve as a general-purpose dataset.
Code: vis-aware-diffusion-anonymization
Contents
COCO/
input/… See the full description on the dataset page: https://huggingface.co/datasets/med-jaouad/vis-aware-anonymization-eval.bridge-anonymization-semanticSound-Event-Anonymization
