CoolFace
Datasetpublic

joelniklaus/mc4_legal

Dataset Card for MC4_Legal: A Corpus Covering the Legal Part of MC4 for European Languages Dataset Summary This dataset contains large text resources (~133GB in total) from mc4 filtered for legal data that can be used for pretraining language models. Use the dataset like this: from datasets import load_dataset dataset = load_dataset("joelito/mc4_legal", "de", split='train', streaming=True) Supported Tasks and Leaderboards The dataset supports the… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/mc4_legal.

sourceHugging Facecc-by-4.0updated 4y agoView on Hugging Face
7likes415downloads
settings

This repository belongs to joelniklaus on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namemc4_legal
visibilitypublic
licencecc-by-4.0
gatedno
ownerjoelniklaus
Account settings
joelniklaus/mc4_legal · CoolFace