CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AMAImedia /multidomain-kazakh-dataset ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/multidomain-kazakh-dataset.text10M<n<100M0 likes701 downloads8d agoHugging Face02kz-transformers /multidomain-kazakh-dataset Dataset Description Point of Contact: Sanzhar Murzakhmetov, Besultan Sagyndyk Dataset Summary MDBKD | Multi-Domain Bilingual Kazakh Dataset is a Kazakh-language dataset containing just over 24 883 808 unique texts from multiple domains. Supported Tasks 'MLM/CLM': can be used to train a model for casual and masked languange modeling Languages The kk code for Kazakh as generally spoken in the Kazakhstan Data Instances For each instance… See the full description on the dataset page: https://huggingface.co/datasets/kz-transformers/multidomain-kazakh-dataset.texttext-generation10M<n<100M30 likes304 downloads1y agoHugging Face03Arailym-tleubayeva /KazakhLawCorpus-clean KazakhLawCorpus-clean Dataset Summary KazakhLawCorpus-clean is a cleaned, Kazakh-only corpus of legislative documents from the Republic of Kazakhstan. It is a processed derivative of the original Arailym-tleubayeva/KazakhLawCorpus dataset. The original dataset repository was downloaded from Hugging Face and used as the source for this release. Its laws_metadata.csv file contained 223,245 legislative records with multilingual fields and source-oriented metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhLawCorpus-clean.tabulartext-retrieval100K<n<1M1 likes101 downloads15d agoHugging Face04lesakbota /kazakhstan-sociology-llm-benchmark Kazakhstan Sociology Consultant — LLM Benchmark Dataset Benchmark dataset for evaluating Large Language Models on sociological survey data analysis tasks (Kazakhstan). Diploma thesis: "Implementation of a visual-statistical analytics module in a digital sociology consultant system" Dataset Description This benchmark evaluates LLMs on their ability to: Parse natural language queries (Russian) about sociological data Generate correct SQLite SQL queries with JOINs Choose… See the full description on the dataset page: https://huggingface.co/datasets/lesakbota/kazakhstan-sociology-llm-benchmark.textn<1K1 likes84 downloads6mo agoHugging Face05R3iwan /entertainment-reviews-kazakhtexttext-classification1K<n<10K0 likes82 downloads9mo agoHugging Face06SagiAbd /Kazakh-Literature-Collectiontextn<1K1 likes51 downloads2y agoHugging Face07Arailym-tleubayeva /small_kazakh_corpus Dataset Card for Small Kazakh Language Corpus The Small Kazakh Language Corpus is a specialized collection of textual data designed for training and research of natural language processing (NLP) models in the Kazakh language. The corpus is structured to ensure high text quality and comprehensive representation of diverse linguistic constructs. Dataset Details Dataset Description The dataset consists of Kazakh language texts with annotations that support tasks… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/small_kazakh_corpus.textmask-generation10K<n<100K1 likes28 downloads2y agoHugging Face08Arailym-tleubayeva /KazakhTextDuplicates Dataset Card for KazakhTextDuplicates Dataset Details Dataset Description The KazakhTextDuplicates dataset is a collection of Kazakh-language texts containing duplicates with different levels of modification. The dataset includes exact duplicates, contextual duplicates, and partial duplicates, making it valuable for research in text similarity, duplicate detection, information retrieval, and plagiarism detection. Developed by: Arailym Tleubayeva Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhTextDuplicates.text10K<n<100K1 likes26 downloads2y agoHugging Face09nurkhan5l /kazakh-iftgatedKazakh-IFT 🇰🇿 Authors: Nurkhan Laiyk, Daniil Orel, Rituraj Joshi, Maiya Goloburda, Yuxia Wang, Preslav Nakov, Fajri Koto Dataset Summary Instruction tuning in low-resource languages remains challenging due to limited coverage of region-specific institutional and cultural knowledge. To address this gap, we introduce a large-scale instruction-following dataset (~10,600 samples) focused on Kazakhstan, spanning domains such as governance, legal processes, cultural practices, and… See the full description on the dataset page: https://huggingface.co/datasets/nurkhan5l/kazakh-ift.text10K<n<100K0 likes20 downloads1y agoHugging Face10Arailym-tleubayeva /sist-kazakh-corpus SIST Kazakh Corpus Description SIST Kazakh Corpus is a curated dataset of Kazakh scientific articles collected for research in text similarity detection, plagiarism analysis, and low-resource NLP tasks. The dataset was created to support: Text similarity detection in agglutinative languages Kazakh NLP benchmarking Scientific text analysis Retrieval-Augmented Generation (RAG) research Dataset Structure The dataset is provided in CSV format. Columns may… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/sist-kazakh-corpus.tabularsentence-similarityn<1K1 likes19 downloads8mo agoHugging Face11nKa1i /kazakh-ai-detect 🇰🇿 KazAI-Detect: Kazakh AI-Generated Text Detection Benchmark KazAI-Detect is the first comprehensive multi-domain benchmark dataset designed specifically for training and evaluating AI-generated text detectors in the Kazakh language. 📌 Dataset Overview Languages: Kazakh (kk) Task: Binary Text Classification (0: Human, 1: AI) Domains: Consumer Reviews (sourced from authentic KazSAnDRA user reviews) Formal News (Informburo, Egemen Qazaqstan) Academic &… See the full description on the dataset page: https://huggingface.co/datasets/nKa1i/kazakh-ai-detect.texttext-classification10K<n<100K0 likes19 downloads1mo agoHugging Face12akylbekmaxutov /kazakh_calibration_datasettext100K<n<1M0 likes17 downloads2y agoHugging Face13kkenbbi /kazakh_reviews_2gistextn<1K1 likes14 downloads1y agoHugging Face14Speech-data /Kazakh-Speech-Dataset 🎧 Kazakh Speech Dataset The Kazakh Speech Dataset is a high-quality speech audio dataset developed to provide structured and scalable audio data for AI and machine learning applications. It includes 130 hours of audio data distributed across 672 files, delivered in MP3 and WAV formats, with a total size of 123 MB. This well-balanced audio dataset ensures diverse and representative voice data, featuring 54% female and 46% male speakers, with an age range spanning from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Kazakh-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes12 downloads6mo agoHugging Face15jojo-ai-mst /Roleplay-Kazakh RolePlay-Kazakh Roleplay-Kazakh Dataset is a dataset for roleplaying in the Kazakh language for the Large Language Model. The base dataset is the GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API. For more information and other language datasets for roleplay, see this github repo. For… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Kazakh.texttext-generation1K<n<10K0 likes10 downloads2y agoHugging Face16TilQazyna /kazakh-morpho-experimentsgated kazakh-morpho-experiments Қазақ морфологиясына арналған тәжірибе материалдары · Материалы экспериментов по казахской морфологии · Kazakh morphology experiment material Қазақша · Русский · English Қазақша kazakh-morpho-experiments — қазақ тілінің морфологиялық талдауын оқытуға және бағалауға қолданылған деректер, скрипттер мен нәтижелер мұрағаты. Репозиторий көлемі — 13.7 МБ; ол тәжірибені қайта қарауға және таңбаларды түзету барысын зерттеуге арналған.… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/kazakh-morpho-experiments.texttoken-classification1K<n<10K0 likes10 downloads2mo agoHugging Face17sabinaasker /kazakh_dollyThis dataset is a Kazakh translation of databricks/databricks-dolly-15k dataset developed by Databricks, Inc. Languages: Kazakh Version: 1.0 This dataset was translated from the original using Google Translate with few minor manual adjustments by a native speaker (me). Original dataset: Copyright (2023) Databricks, Inc. (https://www.databricks.com) This translated dataset is subject to the CC BY-SA 3.0 license, as per the original dataset's licensing terms. For more information on the CC BY-SA… See the full description on the dataset page: https://huggingface.co/datasets/sabinaasker/kazakh_dolly.text10K<n<100K1 likes9 downloads2y agoHugging Face18TilQazyna /kazakh-morpho-1200-sentencesgated kazakh-morpho-1200-sentences Морфологиялық белгіленген қазақша сөйлемдер · Казахские предложения с морфологической разметкой · Morphologically annotated Kazakh sentences Қазақша · Русский · English Қазақша kazakh-morpho-1200-sentences — морфологиялық талдауға толық белгіленген 1 200 қазақша сөйлемнен тұратын, көлемі 0.8 МБ датасет. Жинақ сөйлем деңгейіндегі морфологиялық зерттеулерге арналған бастапқы дерек ретінде қолданылады. Құрамы Толық дерек… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/kazakh-morpho-1200-sentences.tabulartoken-classification10K<n<100K0 likes7 downloads2mo agoHugging Face19R3iwan /kazakh_namestexttext-generation1K<n<10K0 likes4 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.