lexlms/lex_files_preprocessed
Dataset Card for "LexFiles" Dataset Summary Disclaimer: This is a pre-proccessed version of the LexFiles corpus (https://huggingface.co/datasets/lexlms/lexfiles), where documents are pre-split in chunks of 512 tokens. The LeXFiles is a new diverse English multinational legal corpus that we created including 11 distinct sub-corpora that cover legislation and case law from 6 primarily English-speaking legal systems (EU, CoE, Canada, US, UK, India). The corpus… See the full description on the dataset page: https://huggingface.co/datasets/lexlms/lex_files_preprocessed.
This repository belongs to lexlms on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
