CoolFace
Datasetpublic

projecte-aina/CATalog

Dataset Summary CATalog is a diverse, open-source Catalan corpus for language modelling. It consists of text documents from 26 different sources, including web crawling, news, forums, digital libraries and public institutions, totaling in 17.45 billion words. Supported Tasks and Leaderboards Fill-Mask Text Generation other:Language-Modelling: The dataset is suitable for training a model in Language Modelling, predicting the next word in a given context. Success… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CATalog.

sourceHugging Faceupdated 1y agoView on Hugging Face
8likes4.1kdownloads
settings

This repository belongs to projecte-aina on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameCATalog
visibilitypublic
licencenot set
gatedno
ownerprojecte-aina
Account settings
projecte-aina/CATalog · CoolFace