datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turk-ictihat-kararlari-fulltext
Türk İçtihat Kararları - Tam Metin (UYAP e-Bedesten)
Adalet Bakanlığı UYAP Mevzuat ve İçtihat Programı (bedesten.adalet.gov.tr)
üzerinden derlenen Yargıtay kararları veri seti. Kavram bazlı arama
sorguları ile toplanmıştır.
Boyut
Kayıt: 9,899,589 benzersiz karar
Terim dağılımı: {"bosanma": 100052, "karar": 9894542, "nafaka": 66241, "tazminat": 100052}
Yıl aralığı: 1993-2026
Tam metni gelen karar: 39,301 (artımlı olarak dolduruluyor)
Şema… See the full description on the dataset page: https://huggingface.co/datasets/muhammedturan/turk-ictihat-kararlari-fulltext.turkic_tts_dataset
Turkic TTS Dataset
A multilingual TTS corpus covering Turkic languages.
Languages
Subset
Source
Speakers
azerbaijani
BHOSAI/Azerbaijani_News_TTS
1 (female)
bashkir
AigizK/bashkort_tts_dataset
8 (7F + 1M, ElevenLabs cloned)
Schema
Column
Type
Description
audio
Audio
Speech sample
text
string
Transcription
source_link
string
Original dataset URL
speaker_idstring
Speaker identifier (and style if applicable)
gender
string… See the full description on the dataset page: https://huggingface.co/datasets/futureDoctor/turkic_tts_dataset.turkicocr-cyrillic
TurkicOCR Synthetic Cyrillic Dataset
A large-scale synthetic dataset for document AI research in underrepresented Turkic languages — Kazakh and Kyrgyz. Built to cover the full document understanding pipeline: text detection, recognition (OCR), layout analysis, and visual document understanding (VDU).
Pages span 29 authentic document archetypes across administrative, educational, and commercial domains, rendered with 7 procedural degradation profiles that simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/alenisaw/turkicocr-cyrillic.Turkic_trainTurkic_train_lang_script_idcyrillic_turkic_langs
Cyrillic dataset of 8 Turkic languages spoken in Russia and former USSR
Dataset Description
The dataset is a part of the [Leipzig Corpora (Wiki) Collection]: https://corpora.uni-leipzig.de/
For the text-classification comparison, Russian has been included to the dataset.
Paper:
Dirk Goldhahn, Thomas Eckart and Uwe Quasthoff (2012): Building Large Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to 200 Languages. In: Proceedings of the Eighth… See the full description on the dataset page: https://huggingface.co/datasets/tatiana-merz/cyrillic_turkic_langs.old-turkic-corpusclassification_Turkic_languages
Description
A dataset with texts and the categories to which these texts belong.
Usage
This dataset can be used to check language models for the correct classification of texts by category.
Dataset structure:
lang: the language to which the text source belongs;
title: the title of the text;
original_text: original text taken from a web page;
processed_text: processed text using preprocessing functions;
category: the category to which the text belongs;… See the full description on the dataset page: https://huggingface.co/datasets/Electrotubbie/classification_Turkic_languages.triplets_Turkic_languages
Triplets for Turkic languages language models
Description
This dataset is designed to test models for working with Next Sentence Prediction (NSP) and Sentence Order Prediction (SOP). It includes two sub-sets with triplets of texts..
Usage
This dataset can be used to train and evaluate models capable of performing NSP and SAP tasks.
Dataset structure:
Each entry in the dataset represents three values:
text: a triplet of text;
flag: a flag indicating… See the full description on the dataset page: https://huggingface.co/datasets/Electrotubbie/triplets_Turkic_languages.TurkicClassification
TurkicClassification
An MTEB dataset
Massive Text Embedding Benchmark
A dataset of news classification in three Turkic languages.
Task category
t2c
Domains
News, Written
Reference
https://huggingface.co/datasets/Electrotubbie/classification_Turkic_languages/
Source datasets:
Electrotubbie/classification_Turkic_languages
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/TurkicClassification.Turkic_devcyrillic_turkic_langturkicTurkic_dev_lang_script_idturkic-nlp-corpusturkic-eval-benchmarkuseful_turkic_corpora
