mattmdjaga/text-anonymization-benchmark-train
Dataset card for Text Anonymization Benchmark (TAB) train Dataset Summary This is the training split of the Text Anonymisation Benchmark. As the title says it's a dataset focused on text anonymisation, specifcially European Court Documents, which contain labels by mutltiple annotators. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data… See the full description on the dataset page: https://huggingface.co/datasets/mattmdjaga/text-anonymization-benchmark-train.
Dataset card for Text Anonymization Benchmark (TAB) train
Table of Contents
- Table of Contents
- Dataset Description
- Dataset Summary
- Supported Tasks and Leaderboards
- Languages
- Dataset Structure
- Data Instances
- Data Fields
- Data Splits
- Dataset Creation
- Curation Rationale
- Source Data
- Annotations
- Personal and Sensitive Information
- Considerations for Using the Data
- Social Impact of Dataset
- Discussion of Biases
- Other Known Limitations
- Additional Information
- Dataset Curators
- Licensing Information
- Citation Information
- Contributions
Dataset Description
- Homepage:
- Repository:
- Paper:
- Leaderboard:
- Point of Contact:
Dataset Summary
This is the training split of the Text Anonymisation Benchmark. As the title says it's a dataset focused on text anonymisation, specifcially European Court Documents, which contain labels by mutltiple annotators.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data
Initial Data Collection and Normalization
[More Information Needed]
Who are the source language producers?
[More Information Needed]
Annotations
Annotation process
[More Information Needed]
Who are the annotators?
[More Information Needed]
Personal and Sensitive Information
[More Information Needed]
Considerations for Using the Data
Social Impact of Dataset
[More Information Needed]
Discussion of Biases
[More Information Needed]
Other Known Limitations
[More Information Needed]
Additional Information
Dataset Curators
[More Information Needed]
Licensing Information
TAB is released under an MIT License. The MIT License is a short and simple permissive license allowing both commercial and non-commercial use of the software.
Citation Information
[More Information Needed]
Contributions
@article{DBLP:journals/corr/abs-2202-00443,
author = {Ildik{\'{o}} Pil{\'{a}}n and
Pierre Lison and
Lilja {\O}vrelid and
Anthi Papadopoulou and
David S{\'{a}}nchez and
Montserrat Batet},
title = {The Text Anonymization Benchmark {(TAB):} {A} Dedicated Corpus and
Evaluation Framework for Text Anonymization},
journal = {CoRR},
volume = {abs/2202.00443},
year = {2022},
url = {https://arxiv.org/abs/2202.00443},
eprinttype = {arXiv},
eprint = {2202.00443},
timestamp = {Wed, 09 Feb 2022 15:43:35 +0100},
biburl = {https://dblp.org/rec/journals/corr/abs-2202-00443.bib},
bibsource = {dblp computer science bibliography, https://dblp.org}
}