CoolFace
Datasetpublic

catallama/Catalan-Raw-Text

Dataset Summary The Catalan Raw Text Dataset is a subset of the projecte-aina/catalan_general_crawling. It is licensed under a Creative Commons Attribution 4.0 International license, just like the origin dataset. The dataset consists of 404k samples (roughly 20% of the original), totalling 331M tokens after tokenizing it with the Llama-3 Tokenizer. Languages The dataset is in Catalan (ca-ES). Data Fields text (str): Text. Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/catallama/Catalan-Raw-Text.

sourceHugging Facecc-by-sa-4.0updated 2y agoView on Hugging Face
0likes102downloads
7 commits on main
bbbb2782y ago

Update README.md

laurentiubp
c7e55772y ago

Update README.md

laurentiubp
ccc1da52y ago

Update README.md

laurentiubp
a29a3022y ago

Update README.md

laurentiubp
8c1762b2y ago

Update README.md

laurentiubp
2622fe02y ago

Upload dataset

laurentiubp
26c08882y ago

initial commit

laurentiubp