CoolFace
Datasetpublic

projecte-aina/CATalog

Dataset Summary CATalog is a diverse, open-source Catalan corpus for language modelling. It consists of text documents from 26 different sources, including web crawling, news, forums, digital libraries and public institutions, totaling in 17.45 billion words. Supported Tasks and Leaderboards Fill-Mask Text Generation other:Language-Modelling: The dataset is suitable for training a model in Language Modelling, predicting the next word in a given context. Success… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CATalog.

sourceHugging Faceupdated 1y agoView on Hugging Face
8likes4.1kdownloads

projecte-aina/CATalog · main · files are served by the source, never re-hosted here