softcatala/ca_text_corpus
Dataset Card for ca-text-corpus Dataset Summary Public domain corpus of Catalan text. Supported Tasks and Leaderboards This dataset can be used as a small Catalan text corpus for language modeling, text generation experiments, sentence selection, and prompt sentence sourcing for speech datasets. It is not associated with a public leaderboard. Languages Catalan (ca). Dataset Structure Data Instances Each… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/ca_text_corpus.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face