catallama/Catalan-Raw-Text
Dataset Summary The Catalan Raw Text Dataset is a subset of the projecte-aina/catalan_general_crawling. It is licensed under a Creative Commons Attribution 4.0 International license, just like the origin dataset. The dataset consists of 404k samples (roughly 20% of the original), totalling 331M tokens after tokenizing it with the Llama-3 Tokenizer. Languages The dataset is in Catalan (ca-ES). Data Fields text (str): Text. Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/catallama/Catalan-Raw-Text.
Dataset Summary
The Catalan Raw Text Dataset is a subset of the projecte-aina/catalan_general_crawling.
It is licensed under a Creative Commons Attribution 4.0 International license, just like the origin dataset.
The dataset consists of 404k samples (roughly 20% of the original), totalling 331M tokens after tokenizing it with the Llama-3 Tokenizer.
Languages
The dataset is in Catalan (ca-ES).
Data Fields
text(str): Text.
Data Splits
The dataset contains two splits: train and test.
Contributions
Thanks to projecte-aina for providing the original dataset.
Please visit the Source Dataset for more information about how it was collected and curated.
