datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tecla
Dataset Card for TeCla
Dataset Summary
TeCla (Text Classification) is a Catalan News corpus for thematic multi-class Text Classification tasks. The present version (2.0) contains 113.376 articles classified under a hierarchical class structure consisting of a coarse-grained and a fine-grained class. Each of the 4 coarse-grained classes accept a subset of fine-grained ones, 53 in total.
The previous version (1.0.1) can still be found at https://zenodo.org/record/4761505… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/tecla.teclaflan_combined_task1588_tecla_classificationroots_ca_teclaROOTS Subset: roots_ca_tecla
TeCla: Text Classification Catalan dataset
Dataset uid: tecla
Description
TeCla is a Catalan News corpus for thematic Text Classification tasks. It contains 153.265 articles classified under 30 different categories.
The source data is crawled from the ACN (Catalan News Agency) site: http://www.acn.cat, and used under CC-BY-NC-ND 4.0 licence. The dataset is released under the same licence, and is intended exclusively for training Machine… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_ca_tecla.
