datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
space-track-tle-history
Space-Track TLE History
Complete archive of Two-Line Element (TLE) orbital data for every tracked object in Earth orbit, from 1959 to 2026. Sourced from Space-Track.org bulk exports.
Quick Start
from datasets import load_dataset
# Load a specific year
ds = load_dataset("juliensimon/space-track-tle-history", data_files="data/tle_2024.parquet")
# Load everything (238M rows — use streaming for large-scale analysis)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/space-track-tle-history.tlessstarlink-tle-latest
Latest Starlink & GPS TLEs
Credit: NASA
Part of the Orbital Mechanics Datasets collection on Hugging Face.
Dataset description
Latest Two-Line Element (TLE) orbital data for the Starlink and GPS constellations, sourced daily from CelesTrak.
Two-Line Element sets (TLEs) are the standard format for representing satellite orbital elements, developed by NORAD in the 1960s and still used universally today. Each TLE encodes six Keplerian orbital elements plus… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/starlink-tle-latest.constellation-tle-latest
Constellation TLEs -- 18 Satellite Constellations
Credit: NASA
Part of the Orbital Mechanics Datasets collection on Hugging Face.
Dataset description
Daily Two-Line Element (TLE) snapshots for 18 satellite constellations sourced from CelesTrak. Covers GNSS navigation (GPS, Galileo, BeiDou, GLONASS, SBAS), LEO broadband (OneWeb, Kuiper, Qianfan, Hulianwang), LEO communications (Iridium, Globalstar, ORBCOMM), Earth observation (Planet Labs, Spire Global)… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/constellation-tle-latest.space-track-tle-history
Space-Track TLE History
Complete archive of Two-Line Element (TLE) orbital data for every tracked object in Earth orbit, from 1959 to 2026. Sourced from Space-Track.org bulk exports.
Quick Start
from datasets import load_dataset
# Load a specific year
ds = load_dataset("juliensimon/space-track-tle-history", data_files="data/tle_2024.parquet")
# Load everything (238M rows — use streaming for large-scale analysis)ds =… See the full description on the dataset page: https://huggingface.co/datasets/oxzoid/space-track-tle-history.Smiles2Dockhttps://arxiv.org/pdf/2406.05738
twitter-hate-speech-en-240ksamplesThis dataset is a combination of the three datasets listed below:
tdavidson/hate_speech_offensive
LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset
ucberkeley-dlab/measuring-hate-speech
It has only two columns, "tweet" and "labels", and 242738 rows of uncleaned data.
KazakhLawCorpus-clean
KazakhLawCorpus-clean
Dataset Summary
KazakhLawCorpus-clean is a cleaned, Kazakh-only corpus of legislative documents from the Republic of Kazakhstan. It is a processed derivative of the original Arailym-tleubayeva/KazakhLawCorpus dataset.
The original dataset repository was downloaded from Hugging Face and used as the source for this release. Its laws_metadata.csv file contained 223,245 legislative records with multilingual fields and source-oriented metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhLawCorpus-clean.KazakhLawCorpus
Current Release
Current version contains three datasets.
data/
├── laws_metadata.csv
├── law_history.csv
└── law_references.csv
Dataset Description
1. laws_metadata.csv
Contains metadata describing legal acts.
Current size:
223,245 legal acts
Main fields include:
Column
Description
source_id
Internal database identifier
law_id
Stable legal act identifier
title
Original title
title_kk
Kazakh title
title_ru
Russian title… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhLawCorpus.KazakhTextDuplicatesv2.0
KazakhTextDuplicates v2.0
KazakhTextDuplicates v2.0 is a large-scale dataset for duplicate detection, near-duplicate retrieval, semantic textual similarity (STS), and plagiarism detection in the Kazakh language.
Version 2.0 significantly extends the dataset with:
a large augmented training corpus (200K+ pairs),
a continuous semantic similarity score (similarity_score),
multiple difficulty levels of noisy duplicates,
a clean train/validation/test split without identifier overlap.
The… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhTextDuplicatesv2.0.legalup-laws
Kazakhstan Legal Acts Dataset (LegalUp)
Dataset Summary
The LegalUp dataset contains structured metadata for legislative documents of the Republic of Kazakhstan.
The current release includes 392,084 legislative document records extracted from a PostgreSQL database.
The dataset is designed for:
Legal Retrieval-Augmented Generation (Legal RAG)
Information Retrieval
Legal Search
Question Answering
Semantic Search
Legal NLP
Benchmark Construction
Academic Research… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/legalup-laws.bop_distrib_tless
BOP-Distrib dataset
Dataset Description
BOP-Distrib: Revisiting 6D Pose Estimation Benchmarks for Better Evaluation under Visual Ambiguities.
This dataset contains the per-image pose distribution annotation for the T-LESS dataset (available here).
The project page can be found at: https://cea-list.github.io/BOP-Distrib/
Citation Information
If you use the BOP-Distrib annotations in your research, please cite the BOP-Distrib paper:… See the full description on the dataset page: https://huggingface.co/datasets/CEAai/bop_distrib_tless.NK-Oil-Well-Sensor-Monitoring
Oil Well Sensor Monitoring Dataset - NK Field
Dataset Description
This dataset contains hourly sensor readings from 10 oil wells at the NK field.
The monitoring period covers approximately seven months, from January 1, 2026, to July 20, 2026.
The data were provided by Galaz and Company LLP
(ТОО «Галаз и Компания») within the research project:
“Development and Implementation of Control Algorithms for Low-Production-Rate Wells in Mechanized Oil Production Systems… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/NK-Oil-Well-Sensor-Monitoring.this-person-does-not-existThis dataset consists of 8892 AI-generated profile pictures downloaded from (https://www.kaggle.com/datasets/pablobedolla/this-person-does-not-exist-data)
KazOilWellOps_Dataset
Kazakhstan Oil Well Operational Dataset
Description
This dataset contains structured operational and production parameters of sucker rod pump (SRP) oil wells in Kazakhstan.
It is intended for industrial AI research, oil production analysis, production forecasting, and predictive modeling of well performance under real field operating conditions.
Location: North-West Konys oil field, Kyzylorda Region, Kazakhstan (≈150 km NW of Kyzylorda city).
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazOilWellOps_Dataset.LegalRAG
Kazakh Legal Text Chunks
Dataset Summary
Kazakh Legal Text Chunks is a processed corpus of official legal texts of the Republic of Kazakhstan, prepared for retrieval-augmented generation (RAG), legal information retrieval, and grounded legal question answering in the Kazakh language.
The dataset contains structure-preserving text chunks derived from publicly available legal and normative documents. It is intended for research and development in:
legal retrieval,
legal QA… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/LegalRAG.small_kazakh_corpus
Dataset Card for Small Kazakh Language Corpus
The Small Kazakh Language Corpus is a specialized collection of textual data designed for training and research of natural language processing (NLP) models in the Kazakh language. The corpus is structured to ensure high text quality and comprehensive representation of diverse linguistic constructs.
Dataset Details
Dataset Description
The dataset consists of Kazakh language texts with annotations that support tasks… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/small_kazakh_corpus.KazakhTextDuplicates
Dataset Card for KazakhTextDuplicates
Dataset Details
Dataset Description
The KazakhTextDuplicates dataset is a collection of Kazakh-language texts containing duplicates with different levels of modification. The dataset includes exact duplicates, contextual duplicates, and partial duplicates, making it valuable for research in text similarity, duplicate detection, information retrieval, and plagiarism detection.
Developed by: Arailym Tleubayeva
Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhTextDuplicates.sist-kazakh-corpus
SIST Kazakh Corpus
Description
SIST Kazakh Corpus is a curated dataset of Kazakh scientific articles
collected for research in text similarity detection, plagiarism analysis,
and low-resource NLP tasks.
The dataset was created to support:
Text similarity detection in agglutinative languages
Kazakh NLP benchmarking
Scientific text analysis
Retrieval-Augmented Generation (RAG) research
Dataset Structure
The dataset is provided in CSV format.
Columns may… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/sist-kazakh-corpus.tless_env_datasetherb_sheetsAITUAdmissionsGuideDataset
AITU Admissions Guide Dataset
Dataset Details
Dataset Description
This dataset contains questions, answers, and categories related to the admission process at Astana IT University (AITU). It is designed to assist in automating applicant consultations and can be used for chatbot training, recommendation systems, and NLP-based question-answering models.
Curated by: Astana IT University
Funded by [optional]: Arailym Tleubayeva, Alina Mitroshina, Alpar Arman… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/AITUAdmissionsGuideDataset.AncestryOmicsUKBtlem-leaderboardtless-5-objects-examplejozsef_attila_osszesoi_docs_datasetsist-english-corpusoi_docs_synthetic_alpacahungary_history_alpaca
