datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mapa
Dataset Card for Multilingual European Datasets for Sensitive Entity Detection in the Legal Domain
Dataset Summary
The dataset consists of 12 documents (9 for Spanish due to parsing errors) taken from EUR-Lex, a multilingual corpus of court
decisions and legal dispositions in the 24 official languages of the European Union. The documents have been annotated
for named entities following the guidelines of the MAPA project which foresees two
annotation level, a general and a… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/mapa.mapa-eur-lex-pii
Dataset Card for MAPA EUR-LEX PII
A small, multilingual PII-detection benchmark built from the dglover1/mapa-eur-lex test split (itself derived from joelniklaus/mapa). Two English EUR-LEX legal documents were re-annotated by hand with a fine-grained PII scheme, and those labels were then projected onto their professional human translations in 20 other EU languages.
Dataset Details
Dataset Description
The dataset contains 42 documents: two source… See the full description on the dataset page: https://huggingface.co/datasets/piimb/mapa-eur-lex-pii.mapa-eur-lex
Dataset Card for Multilingual European Datasets for Sensitive Entity Detection in the Legal Domain
Dataset Summary
This dataset is a completed version of the MAPA EUR-LEX dataset, originally converted to Huggingface format by joelniklaus. See the dataset card for more information about MAPA.
3 of the (Spanish) EUR-LEX WebAnno TSV files in the source MAPA repository are malformed, so they were omitted from the original conversion, causing under-representation of the… See the full description on the dataset page: https://huggingface.co/datasets/dglover1/mapa-eur-lex.self-instruct-seed-ca
Catalan self-instruct seed
Manual translation of the seed instructions from self-instruct.
Note that some examples could not be literally translated (e.g. jokes, puns, code) and had to be adapted to the target language.
