datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KazakhLawCorpus-clean
KazakhLawCorpus-clean
Dataset Summary
KazakhLawCorpus-clean is a cleaned, Kazakh-only corpus of legislative documents from the Republic of Kazakhstan. It is a processed derivative of the original Arailym-tleubayeva/KazakhLawCorpus dataset.
The original dataset repository was downloaded from Hugging Face and used as the source for this release. Its laws_metadata.csv file contained 223,245 legislative records with multilingual fields and source-oriented metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhLawCorpus-clean.sist-kazakh-corpus
SIST Kazakh Corpus
Description
SIST Kazakh Corpus is a curated dataset of Kazakh scientific articles
collected for research in text similarity detection, plagiarism analysis,
and low-resource NLP tasks.
The dataset was created to support:
Text similarity detection in agglutinative languages
Kazakh NLP benchmarking
Scientific text analysis
Retrieval-Augmented Generation (RAG) research
Dataset Structure
The dataset is provided in CSV format.
Columns may… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/sist-kazakh-corpus.kazakh-morpho-1200-sentences
kazakh-morpho-1200-sentences
Морфологиялық белгіленген қазақша сөйлемдер · Казахские предложения с морфологической разметкой · Morphologically annotated Kazakh sentences
Қазақша · Русский · English
Қазақша
kazakh-morpho-1200-sentences — морфологиялық талдауға толық белгіленген 1 200 қазақша сөйлемнен тұратын, көлемі 0.8 МБ датасет. Жинақ сөйлем деңгейіндегі морфологиялық зерттеулерге арналған бастапқы дерек ретінде қолданылады.
Құрамы
Толық дерек… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/kazakh-morpho-1200-sentences.
