datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twitter-hate-speech-en-240ksamplesThis dataset is a combination of the three datasets listed below:
tdavidson/hate_speech_offensive
LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset
ucberkeley-dlab/measuring-hate-speech
It has only two columns, "tweet" and "labels", and 242738 rows of uncleaned data.
KazakhLawCorpus-clean
KazakhLawCorpus-clean
Dataset Summary
KazakhLawCorpus-clean is a cleaned, Kazakh-only corpus of legislative documents from the Republic of Kazakhstan. It is a processed derivative of the original Arailym-tleubayeva/KazakhLawCorpus dataset.
The original dataset repository was downloaded from Hugging Face and used as the source for this release. Its laws_metadata.csv file contained 223,245 legislative records with multilingual fields and source-oriented metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhLawCorpus-clean.legalup-laws
Kazakhstan Legal Acts Dataset (LegalUp)
Dataset Summary
The LegalUp dataset contains structured metadata for legislative documents of the Republic of Kazakhstan.
The current release includes 392,084 legislative document records extracted from a PostgreSQL database.
The dataset is designed for:
Legal Retrieval-Augmented Generation (Legal RAG)
Information Retrieval
Legal Search
Question Answering
Semantic Search
Legal NLP
Benchmark Construction
Academic Research… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/legalup-laws.NK-Oil-Well-Sensor-Monitoring
Oil Well Sensor Monitoring Dataset - NK Field
Dataset Description
This dataset contains hourly sensor readings from 10 oil wells at the NK field.
The monitoring period covers approximately seven months, from January 1, 2026, to July 20, 2026.
The data were provided by Galaz and Company LLP
(ТОО «Галаз и Компания») within the research project:
“Development and Implementation of Control Algorithms for Low-Production-Rate Wells in Mechanized Oil Production Systems… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/NK-Oil-Well-Sensor-Monitoring.KazOilWellOps_Dataset
Kazakhstan Oil Well Operational Dataset
Description
This dataset contains structured operational and production parameters of sucker rod pump (SRP) oil wells in Kazakhstan.
It is intended for industrial AI research, oil production analysis, production forecasting, and predictive modeling of well performance under real field operating conditions.
Location: North-West Konys oil field, Kyzylorda Region, Kazakhstan (≈150 km NW of Kyzylorda city).
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazOilWellOps_Dataset.small_kazakh_corpus
Dataset Card for Small Kazakh Language Corpus
The Small Kazakh Language Corpus is a specialized collection of textual data designed for training and research of natural language processing (NLP) models in the Kazakh language. The corpus is structured to ensure high text quality and comprehensive representation of diverse linguistic constructs.
Dataset Details
Dataset Description
The dataset consists of Kazakh language texts with annotations that support tasks… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/small_kazakh_corpus.KazakhTextDuplicates
Dataset Card for KazakhTextDuplicates
Dataset Details
Dataset Description
The KazakhTextDuplicates dataset is a collection of Kazakh-language texts containing duplicates with different levels of modification. The dataset includes exact duplicates, contextual duplicates, and partial duplicates, making it valuable for research in text similarity, duplicate detection, information retrieval, and plagiarism detection.
Developed by: Arailym Tleubayeva
Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhTextDuplicates.sist-kazakh-corpus
SIST Kazakh Corpus
Description
SIST Kazakh Corpus is a curated dataset of Kazakh scientific articles
collected for research in text similarity detection, plagiarism analysis,
and low-resource NLP tasks.
The dataset was created to support:
Text similarity detection in agglutinative languages
Kazakh NLP benchmarking
Scientific text analysis
Retrieval-Augmented Generation (RAG) research
Dataset Structure
The dataset is provided in CSV format.
Columns may… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/sist-kazakh-corpus.herb_sheetsAITUAdmissionsGuideDataset
AITU Admissions Guide Dataset
Dataset Details
Dataset Description
This dataset contains questions, answers, and categories related to the admission process at Astana IT University (AITU). It is designed to assist in automating applicant consultations and can be used for chatbot training, recommendation systems, and NLP-based question-answering models.
Curated by: Astana IT University
Funded by [optional]: Arailym Tleubayeva, Alina Mitroshina, Alpar Arman… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/AITUAdmissionsGuideDataset.sist-english-corpusopus100-multilingualg23-project
