CoolFace
Datasetpublic

dennlinger/eur-lex-sum

The EUR-Lex-Sum dataset is a multilingual resource intended for text summarization in the legal domain. It is based on human-written summaries of legal acts issued by the European Union. It distinguishes itself by introducing a smaller set of high-quality human-written samples, each of which have much longer references (and summaries!) than comparable datasets. Additionally, the underlying legal acts provide a challenging domain-specific application to legal texts, which are so far underrepresented in non-English languages. For each legal act, the sample can be available in up to 24 languages (the officially recognized languages in the European Union); the validation and test samples consist entirely of samples available in all languages, and are aligned across all languages at the paragraph level.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
51likes1.7kdownloads

dennlinger/eur-lex-sum · main · files are served by the source, never re-hosted here