CoolFace
20 results

universals

jorgeortizfuentes /universal_spanish_chilean_corpus Universal Chilean Spanish Corpus Este dataset se compone de 37_213_992 textos correspondientes a español de Chile y a español multidialectal. Los textos en español multidialectal provienen del spanish books. Los textos en español de Chile vienen de los dominios .cl del mc4 dataset y de tweets, noticias y reclamos de l chilean-spanish-corpus Name Count Source books 87967 spanish books mc4 8706681 from mc4 (.cl domains) in chilean-spanish-corpus twitter 27306583… See the full description on the dataset page: https://huggingface.co/datasets/jorgeortizfuentes/universal_spanish_chilean_corpus.texttext-generation10M<n<100M8 likes1.4k downloads3y agoHugging Face0xZee /UniversalScienceKownledge-finetome-top-20ktext10K<n<100K0 likes58 downloads2y agoHugging FaceFrancophonIA /Universal_Segmentations_1.0 [!NOTE] Dataset origin: https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-4629 Description Universal Segmentations (UniSegments) is a collection of lexical resources capturing morphological segmentations harmonised into a cross-linguistically consistent annotation scheme for many languages. The annotation scheme consists of simple tab-separated columns that stores a word and its morphological segmentations, including pieces of information about the word and the segmented… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Universal_Segmentations_1.0.0 likes9 downloads1y agoHugging FaceSraghvi /universal-subset-testtabular1K<n<10K0 likes7 downloads1y agoHugging FaceSraghvi /universal-subset-fixedtabular1K<n<10K0 likes7 downloads1y agoHugging Facetutubool /universal-studio-sessiongatedtextn<1K0 likes3 downloads10mo agoHugging Face