CoolFace
3 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tatiana-merz /cyrillic_turkic_langs Cyrillic dataset of 8 Turkic languages spoken in Russia and former USSR Dataset Description The dataset is a part of the [Leipzig Corpora (Wiki) Collection]: https://corpora.uni-leipzig.de/ For the text-classification comparison, Russian has been included to the dataset. Paper: Dirk Goldhahn, Thomas Eckart and Uwe Quasthoff (2012): Building Large Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to 200 Languages. In: Proceedings of the Eighth… See the full description on the dataset page: https://huggingface.co/datasets/tatiana-merz/cyrillic_turkic_langs.texttext-classification10K<n<100K1 likes105 downloads4y agoHugging Face02Electrotubbie /classification_Turkic_languages Description A dataset with texts and the categories to which these texts belong. Usage This dataset can be used to check language models for the correct classification of texts by category. Dataset structure: lang: the language to which the text source belongs; title: the title of the text; original_text: original text taken from a web page; processed_text: processed text using preprocessing functions; category: the category to which the text belongs;… See the full description on the dataset page: https://huggingface.co/datasets/Electrotubbie/classification_Turkic_languages.texttext-classification100K<n<1M1 likes35 downloads3y agoHugging Face03Electrotubbie /triplets_Turkic_languages Triplets for Turkic languages language models Description This dataset is designed to test models for working with Next Sentence Prediction (NSP) and Sentence Order Prediction (SOP). It includes two sub-sets with triplets of texts.. Usage This dataset can be used to train and evaluate models capable of performing NSP and SAP tasks. Dataset structure: Each entry in the dataset represents three values: text: a triplet of text; flag: a flag indicating… See the full description on the dataset page: https://huggingface.co/datasets/Electrotubbie/triplets_Turkic_languages.texttext-classification10K<n<100K1 likes23 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.