CoolFace
Datasetpublic

faur-ai/fulg

❄️FuLG The FuLG dataset is a comprehensive Romanian language corpus comprising 150 billion tokens, carefully extracted from Common Crawl. This extensive dataset is the result of rigorous filtering and deduplication processes applied to 95 Common Crawl snapshots. The compressed dataset has 289 GB. For more details, check the arXiv preprint. How do I download this? Using 🤗 Datasets from datasets import load_dataset # Full dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/fulg.

sourceHugging Faceodc-byupdated 2y agoView on Hugging Face
10likes1kdownloads
settings

This repository belongs to faur-ai on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namefulg
visibilitypublic
licenceodc-by
gatedno
ownerfaur-ai
Account settings
faur-ai/fulg · CoolFace