universals
universal_spanish_chilean_corpus
Universal Chilean Spanish Corpus
Este dataset se compone de 37_213_992 textos correspondientes a español de Chile y a español multidialectal.
Los textos en español multidialectal provienen del spanish books.
Los textos en español de Chile vienen de los dominios .cl del mc4 dataset y de tweets, noticias y reclamos de l chilean-spanish-corpus
Name
Count
Source
books
87967
spanish books
mc4
8706681
from mc4 (.cl domains) in chilean-spanish-corpus
twitter
27306583… See the full description on the dataset page: https://huggingface.co/datasets/jorgeortizfuentes/universal_spanish_chilean_corpus.UniversalScienceKownledge-finetome-top-20kUniversal_Segmentations_1.0
[!NOTE]
Dataset origin: https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-4629
Description
Universal Segmentations (UniSegments) is a collection of lexical resources capturing morphological segmentations harmonised into a cross-linguistically consistent annotation scheme for many languages. The annotation scheme consists of simple tab-separated columns that stores a word and its morphological segmentations, including pieces of information about the word and the segmented… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Universal_Segmentations_1.0.universal-subset-testuniversal-subset-fixeduniversal-studio-session
