datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
c4_validationcarol-subcorpora
Subcorpora Carolina
Contains Carol·B and Carol·(D+B), respectively balanced and deduplicated subcorpora of the Carolina Corpus (Bea) version.
Carol·B was balanced in terms of tokens per domain from Carolina Corpus. This means that it contains approximately the same number of tokens (~60,2M) from each of Carolina’s largest domains: Legislative, Instructional, Entertainment, Journalistic, Juridical and Virtual Forum. Carol·B has, in total, 361,071,147 tokens and 5,5 GB.
Carol·(D+B)… See the full description on the dataset page: https://huggingface.co/datasets/carolina-c4ai/carol-subcorpora.
