CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01muhammedturan /turk-ictihat-kararlari-fulltext Türk İçtihat Kararları - Tam Metin (UYAP e-Bedesten) Adalet Bakanlığı UYAP Mevzuat ve İçtihat Programı (bedesten.adalet.gov.tr) üzerinden derlenen Yargıtay kararları veri seti. Kavram bazlı arama sorguları ile toplanmıştır. Boyut Kayıt: 9,899,589 benzersiz karar Terim dağılımı: {"bosanma": 100052, "karar": 9894542, "nafaka": 66241, "tazminat": 100052} Yıl aralığı: 1993-2026 Tam metni gelen karar: 39,301 (artımlı olarak dolduruluyor) Şema… See the full description on the dataset page: https://huggingface.co/datasets/muhammedturan/turk-ictihat-kararlari-fulltext.tabulartext-generation1M<n<10M1 likes584 downloads11d agoHugging Face02futureDoctor /turkic_tts_dataset Turkic TTS Dataset A multilingual TTS corpus covering Turkic languages. Languages Subset Source Speakers azerbaijani BHOSAI/Azerbaijani_News_TTS 1 (female) bashkir AigizK/bashkort_tts_dataset 8 (7F + 1M, ElevenLabs cloned) Schema Column Type Description audio Audio Speech sample text string Transcription source_link string Original dataset URL speaker_idstring Speaker identifier (and style if applicable) gender string… See the full description on the dataset page: https://huggingface.co/datasets/futureDoctor/turkic_tts_dataset.audiotext-to-speech10K<n<100K0 likes280 downloads4mo agoHugging Face03alenisaw /turkicocr-cyrillic TurkicOCR Synthetic Cyrillic Dataset A large-scale synthetic dataset for document AI research in underrepresented Turkic languages — Kazakh and Kyrgyz. Built to cover the full document understanding pipeline: text detection, recognition (OCR), layout analysis, and visual document understanding (VDU). Pages span 29 authentic document archetypes across administrative, educational, and commercial domains, rendered with 7 procedural degradation profiles that simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/alenisaw/turkicocr-cyrillic.textimage-to-text100K<n<1M0 likes249 downloads1mo agoHugging Face04yiyic /Turkic_traintext1M<n<10M2 likes233 downloads2y agoHugging Face05yiyic /Turkic_train_lang_script_idtext1M<n<10M0 likes161 downloads2y agoHugging Face06tatiana-merz /cyrillic_turkic_langs Cyrillic dataset of 8 Turkic languages spoken in Russia and former USSR Dataset Description The dataset is a part of the [Leipzig Corpora (Wiki) Collection]: https://corpora.uni-leipzig.de/ For the text-classification comparison, Russian has been included to the dataset. Paper: Dirk Goldhahn, Thomas Eckart and Uwe Quasthoff (2012): Building Large Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to 200 Languages. In: Proceedings of the Eighth… See the full description on the dataset page: https://huggingface.co/datasets/tatiana-merz/cyrillic_turkic_langs.texttext-classification10K<n<100K1 likes105 downloads4y agoHugging Face07gaydmi /old-turkic-corpustabular1K<n<10K1 likes39 downloads2y agoHugging Face08Electrotubbie /classification_Turkic_languages Description A dataset with texts and the categories to which these texts belong. Usage This dataset can be used to check language models for the correct classification of texts by category. Dataset structure: lang: the language to which the text source belongs; title: the title of the text; original_text: original text taken from a web page; processed_text: processed text using preprocessing functions; category: the category to which the text belongs;… See the full description on the dataset page: https://huggingface.co/datasets/Electrotubbie/classification_Turkic_languages.texttext-classification100K<n<1M1 likes35 downloads3y agoHugging Face09Electrotubbie /triplets_Turkic_languages Triplets for Turkic languages language models Description This dataset is designed to test models for working with Next Sentence Prediction (NSP) and Sentence Order Prediction (SOP). It includes two sub-sets with triplets of texts.. Usage This dataset can be used to train and evaluate models capable of performing NSP and SAP tasks. Dataset structure: Each entry in the dataset represents three values: text: a triplet of text; flag: a flag indicating… See the full description on the dataset page: https://huggingface.co/datasets/Electrotubbie/triplets_Turkic_languages.texttext-classification10K<n<100K1 likes23 downloads3y agoHugging Face10mteb /TurkicClassification TurkicClassification An MTEB dataset Massive Text Embedding Benchmark A dataset of news classification in three Turkic languages. Task category t2c Domains News, Written Reference https://huggingface.co/datasets/Electrotubbie/classification_Turkic_languages/ Source datasets: Electrotubbie/classification_Turkic_languages How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/TurkicClassification.texttext-classification1K<n<10K1 likes19 downloads1y agoHugging Face11yiyic /Turkic_devtextn<1K0 likes13 downloads2y agoHugging Face12mteb /cyrillic_turkic_langtext10K<n<100K0 likes12 downloads1y agoHugging Face13mteb /turkictext100K<n<1M0 likes9 downloads1y agoHugging Face14yiyic /Turkic_dev_lang_script_idtextn<1K0 likes8 downloads2y agoHugging Face15Nigina-Rinatova /turkic-nlp-corpustext10M<n<100M0 likes6 downloads4mo agoHugging Face16Nigina-Rinatova /turkic-eval-benchmarktext1K<n<10K0 likes5 downloads4mo agoHugging Face17yasalma /useful_turkic_corporagatedtext10M<n<100M0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.