CoolFace
Datasetpublic

catallama/Catalan-Raw-Text

Dataset Summary The Catalan Raw Text Dataset is a subset of the projecte-aina/catalan_general_crawling. It is licensed under a Creative Commons Attribution 4.0 International license, just like the origin dataset. The dataset consists of 404k samples (roughly 20% of the original), totalling 331M tokens after tokenizing it with the Llama-3 Tokenizer. Languages The dataset is in Catalan (ca-ES). Data Fields text (str): Text. Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/catallama/Catalan-Raw-Text.

sourceHugging Facecc-by-sa-4.0updated 2y agoView on Hugging Face
0likes88downloads
Dataset Card

Dataset Summary

The Catalan Raw Text Dataset is a subset of the projecte-aina/catalan_general_crawling.

It is licensed under a Creative Commons Attribution 4.0 International license, just like the origin dataset.

The dataset consists of 404k samples (roughly 20% of the original), totalling 331M tokens after tokenizing it with the Llama-3 Tokenizer.

Languages

The dataset is in Catalan (ca-ES).

Data Fields

  • text (str): Text.

Data Splits

The dataset contains two splits: train and test.

Contributions

Thanks to projecte-aina for providing the original dataset.

Please visit the Source Dataset for more information about how it was collected and curated.