datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eur-lex
EUR-Lex EN–DA (Parallel Legal Text)
A parallel corpus of EU legal documents in English and Danish. Contains only samples where both languages are present.
Dataset Structure
Features
Field
Type
Description
celex
string
CELEX document identifier
resource_type
string
Type of legal document (caselaw, decision, directive, intagr, recommendation, regulation)
url
string
Source URL
title_en
string
English title
title_da
string
Danish title
text_en… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/eur-lex.eurlex_sdg_coverage
EUR-Lex SDG-Annotated Dataset
Dataset Summary
The EUR-Lex SDG-Annotated Dataset enhances the original EUR-Lex dataset by mapping European Union legislation documents to the United Nations Sustainable Development Goals (SDGs). Each document in this dataset has been annotated using a rule-based keyword matching approach, leveraging SDG-related terms and synonyms extracted from an Excel-based term matrix.
This dataset provides researchers, policymakers, and data scientists… See the full description on the dataset page: https://huggingface.co/datasets/razaulhaq/eurlex_sdg_coverage.eurlexsum_ita_cleaned_16384_184
Dataset Card for "eurlexsum_ita_cleaned_16384_184"
More Information needed
eurlexsum_ita_cleaned_32768_299
Dataset Card for "eurlexsum_ita_cleaned_32768_299"
More Information needed
eurlexsum_ita_cleaned_8192_232
Dataset Card for "eurlexsum_ita_cleaned_8192_232"
More Information needed
LGZ_eurlexsum
Dataset Card for "LGZ_eurlexsum"
More Information needed
LexGenZero_eurlexsum
Dataset Card for "LexGenZero_eurlexsum"
More Information needed
eurlexsum_ita_cleaned_8192_86
Dataset Card for "eurlexsum_ita_cleaned_8192_86"
More Information needed
