faur-ai/fulg
❄️FuLG The FuLG dataset is a comprehensive Romanian language corpus comprising 150 billion tokens, carefully extracted from Common Crawl. This extensive dataset is the result of rigorous filtering and deduplication processes applied to 95 Common Crawl snapshots. The compressed dataset has 289 GB. For more details, check the arXiv preprint. How do I download this? Using 🤗 Datasets from datasets import load_dataset # Full dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/fulg.
This repository belongs to faur-ai on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
