datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multidomain-kazakh-dataset
⚡ Each donation funds the next large quant.
I host free GGUF or MoE quants as independent research.
Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro.
Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.
🎉 Boosty🦄 |
☕ Buy Me a Coffee🦄 |
⭐ DonationAlerts🦄
💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/multidomain-kazakh-dataset.multidomain-kazakh-dataset
Dataset Description
Point of Contact: Sanzhar Murzakhmetov, Besultan Sagyndyk
Dataset Summary
MDBKD | Multi-Domain Bilingual Kazakh Dataset is a Kazakh-language dataset containing just over 24 883 808 unique texts from multiple domains.
Supported Tasks
'MLM/CLM': can be used to train a model for casual and masked languange modeling
Languages
The kk code for Kazakh as generally spoken in the Kazakhstan
Data Instances
For each instance… See the full description on the dataset page: https://huggingface.co/datasets/kz-transformers/multidomain-kazakh-dataset.KazakhLawCorpus-clean
KazakhLawCorpus-clean
Dataset Summary
KazakhLawCorpus-clean is a cleaned, Kazakh-only corpus of legislative documents from the Republic of Kazakhstan. It is a processed derivative of the original Arailym-tleubayeva/KazakhLawCorpus dataset.
The original dataset repository was downloaded from Hugging Face and used as the source for this release. Its laws_metadata.csv file contained 223,245 legislative records with multilingual fields and source-oriented metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhLawCorpus-clean.kazakhstan-sociology-llm-benchmark
Kazakhstan Sociology Consultant — LLM Benchmark Dataset
Benchmark dataset for evaluating Large Language Models on sociological survey data analysis tasks (Kazakhstan).
Diploma thesis: "Implementation of a visual-statistical analytics module in a digital sociology consultant system"
Dataset Description
This benchmark evaluates LLMs on their ability to:
Parse natural language queries (Russian) about sociological data
Generate correct SQLite SQL queries with JOINs
Choose… See the full description on the dataset page: https://huggingface.co/datasets/lesakbota/kazakhstan-sociology-llm-benchmark.entertainment-reviews-kazakhKazakh-Literature-Collectionsmall_kazakh_corpus
Dataset Card for Small Kazakh Language Corpus
The Small Kazakh Language Corpus is a specialized collection of textual data designed for training and research of natural language processing (NLP) models in the Kazakh language. The corpus is structured to ensure high text quality and comprehensive representation of diverse linguistic constructs.
Dataset Details
Dataset Description
The dataset consists of Kazakh language texts with annotations that support tasks… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/small_kazakh_corpus.KazakhTextDuplicates
Dataset Card for KazakhTextDuplicates
Dataset Details
Dataset Description
The KazakhTextDuplicates dataset is a collection of Kazakh-language texts containing duplicates with different levels of modification. The dataset includes exact duplicates, contextual duplicates, and partial duplicates, making it valuable for research in text similarity, duplicate detection, information retrieval, and plagiarism detection.
Developed by: Arailym Tleubayeva
Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhTextDuplicates.kazakh-iftKazakh-IFT 🇰🇿
Authors: Nurkhan Laiyk, Daniil Orel, Rituraj Joshi, Maiya Goloburda, Yuxia Wang, Preslav Nakov, Fajri Koto
Dataset Summary
Instruction tuning in low-resource languages remains challenging due to limited coverage of region-specific institutional and cultural knowledge. To address this gap, we introduce a large-scale instruction-following dataset (~10,600 samples) focused on Kazakhstan, spanning domains such as governance, legal processes, cultural practices, and… See the full description on the dataset page: https://huggingface.co/datasets/nurkhan5l/kazakh-ift.sist-kazakh-corpus
SIST Kazakh Corpus
Description
SIST Kazakh Corpus is a curated dataset of Kazakh scientific articles
collected for research in text similarity detection, plagiarism analysis,
and low-resource NLP tasks.
The dataset was created to support:
Text similarity detection in agglutinative languages
Kazakh NLP benchmarking
Scientific text analysis
Retrieval-Augmented Generation (RAG) research
Dataset Structure
The dataset is provided in CSV format.
Columns may… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/sist-kazakh-corpus.kazakh-ai-detect
🇰🇿 KazAI-Detect: Kazakh AI-Generated Text Detection Benchmark
KazAI-Detect is the first comprehensive multi-domain benchmark dataset designed specifically for training and evaluating AI-generated text detectors in the Kazakh language.
📌 Dataset Overview
Languages: Kazakh (kk)
Task: Binary Text Classification (0: Human, 1: AI)
Domains:
Consumer Reviews (sourced from authentic KazSAnDRA user reviews)
Formal News (Informburo, Egemen Qazaqstan)
Academic &… See the full description on the dataset page: https://huggingface.co/datasets/nKa1i/kazakh-ai-detect.kazakh_calibration_datasetkazakh_reviews_2gisKazakh-Speech-Dataset
🎧 Kazakh Speech Dataset
The Kazakh Speech Dataset is a high-quality speech audio dataset developed to provide structured and scalable audio data for AI and machine learning applications. It includes 130 hours of audio data distributed across 672 files, delivered in MP3 and WAV formats, with a total size of 123 MB. This well-balanced audio dataset ensures diverse and representative voice data, featuring 54% female and 46% male speakers, with an age range spanning from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Kazakh-Speech-Dataset.Roleplay-Kazakh
RolePlay-Kazakh
Roleplay-Kazakh Dataset is a dataset for roleplaying in the Kazakh language for the Large Language Model.
The base dataset is the GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, see this github repo.
For… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Kazakh.kazakh-morpho-experiments
kazakh-morpho-experiments
Қазақ морфологиясына арналған тәжірибе материалдары · Материалы экспериментов по казахской морфологии · Kazakh morphology experiment material
Қазақша · Русский · English
Қазақша
kazakh-morpho-experiments — қазақ тілінің морфологиялық талдауын оқытуға және бағалауға қолданылған деректер, скрипттер мен нәтижелер мұрағаты. Репозиторий көлемі — 13.7 МБ; ол тәжірибені қайта қарауға және таңбаларды түзету барысын зерттеуге арналған.… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/kazakh-morpho-experiments.kazakh_dollyThis dataset is a Kazakh translation of databricks/databricks-dolly-15k dataset developed by Databricks, Inc.
Languages: Kazakh Version: 1.0
This dataset was translated from the original using Google Translate with few minor manual adjustments by a native speaker (me).
Original dataset: Copyright (2023) Databricks, Inc. (https://www.databricks.com)
This translated dataset is subject to the CC BY-SA 3.0 license, as per the original dataset's licensing terms.
For more information on the CC BY-SA… See the full description on the dataset page: https://huggingface.co/datasets/sabinaasker/kazakh_dolly.kazakh-morpho-1200-sentences
kazakh-morpho-1200-sentences
Морфологиялық белгіленген қазақша сөйлемдер · Казахские предложения с морфологической разметкой · Morphologically annotated Kazakh sentences
Қазақша · Русский · English
Қазақша
kazakh-morpho-1200-sentences — морфологиялық талдауға толық белгіленген 1 200 қазақша сөйлемнен тұратын, көлемі 0.8 МБ датасет. Жинақ сөйлем деңгейіндегі морфологиялық зерттеулерге арналған бастапқы дерек ретінде қолданылады.
Құрамы
Толық дерек… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/kazakh-morpho-1200-sentences.kazakh_names
