Arailym-tleubayeva/sist-kazakh-corpus
SIST Kazakh Corpus Description SIST Kazakh Corpus is a curated dataset of Kazakh scientific articles collected for research in text similarity detection, plagiarism analysis, and low-resource NLP tasks. The dataset was created to support: Text similarity detection in agglutinative languages Kazakh NLP benchmarking Scientific text analysis Retrieval-Augmented Generation (RAG) research Dataset Structure The dataset is provided in CSV format. Columns… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/sist-kazakh-corpus.
This repository belongs to Arailym-tleubayeva on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
