CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01juliensimon /space-track-tle-history Space-Track TLE History Complete archive of Two-Line Element (TLE) orbital data for every tracked object in Earth orbit, from 1959 to 2026. Sourced from Space-Track.org bulk exports. Quick Start from datasets import load_dataset # Load a specific year ds = load_dataset("juliensimon/space-track-tle-history", data_files="data/tle_2024.parquet") # Load everything (238M rows — use streaming for large-scale analysis) ds =… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/space-track-tle-history.tabulartime-series-forecasting100M<n<1B2 likes2.9k downloads9h agoHugging Face02bop-benchmark /tless2 likes1.7k downloads2y agoHugging Face03juliensimon /starlink-tle-latest Latest Starlink & GPS TLEs Credit: NASA Part of the Orbital Mechanics Datasets collection on Hugging Face. Dataset description Latest Two-Line Element (TLE) orbital data for the Starlink and GPS constellations, sourced daily from CelesTrak. Two-Line Element sets (TLEs) are the standard format for representing satellite orbital elements, developed by NORAD in the 1960s and still used universally today. Each TLE encodes six Keplerian orbital elements plus… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/starlink-tle-latest.texttabular-regression10K<n<100K0 likes1.3k downloads18h agoHugging Face04juliensimon /constellation-tle-latest Constellation TLEs -- 18 Satellite Constellations Credit: NASA Part of the Orbital Mechanics Datasets collection on Hugging Face. Dataset description Daily Two-Line Element (TLE) snapshots for 18 satellite constellations sourced from CelesTrak. Covers GNSS navigation (GPS, Galileo, BeiDou, GLONASS, SBAS), LEO broadband (OneWeb, Kuiper, Qianfan, Hulianwang), LEO communications (Iridium, Globalstar, ORBCOMM), Earth observation (Planet Labs, Spire Global)… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/constellation-tle-latest.texttabular-regression1K<n<10K1 likes1.3k downloads17h agoHugging Face05oxzoid /space-track-tle-history Space-Track TLE History Complete archive of Two-Line Element (TLE) orbital data for every tracked object in Earth orbit, from 1959 to 2026. Sourced from Space-Track.org bulk exports. Quick Start from datasets import load_dataset # Load a specific year ds = load_dataset("juliensimon/space-track-tle-history", data_files="data/tle_2024.parquet") # Load everything (238M rows — use streaming for large-scale analysis)ds =… See the full description on the dataset page: https://huggingface.co/datasets/oxzoid/space-track-tle-history.tabulartime-series-forecasting100M<n<1B0 likes549 downloads6mo agoHugging Face06tlemenestrel /Smiles2Dockhttps://arxiv.org/pdf/2406.05738 text10M<n<100M1 likes433 downloads1y agoHugging Face07TLeonidas /twitter-hate-speech-en-240ksamplesThis dataset is a combination of the three datasets listed below: tdavidson/hate_speech_offensive LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset ucberkeley-dlab/measuring-hate-speech It has only two columns, "tweet" and "labels", and 242738 rows of uncleaned data. text100K<n<1M1 likes131 downloads2y agoHugging Face08Arailym-tleubayeva /KazakhLawCorpus-clean KazakhLawCorpus-clean Dataset Summary KazakhLawCorpus-clean is a cleaned, Kazakh-only corpus of legislative documents from the Republic of Kazakhstan. It is a processed derivative of the original Arailym-tleubayeva/KazakhLawCorpus dataset. The original dataset repository was downloaded from Hugging Face and used as the source for this release. Its laws_metadata.csv file contained 223,245 legislative records with multilingual fields and source-oriented metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhLawCorpus-clean.tabulartext-retrieval100K<n<1M1 likes117 downloads12d agoHugging Face09Arailym-tleubayeva /KazakhLawCorpus Current Release Current version contains three datasets. data/ ├── laws_metadata.csv ├── law_history.csv └── law_references.csv Dataset Description 1. laws_metadata.csv Contains metadata describing legal acts. Current size: 223,245 legal acts Main fields include: Column Description source_id Internal database identifier law_id Stable legal act identifier title Original title title_kk Kazakh title title_ru Russian title… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhLawCorpus.text-retrieval100K<n<1M1 likes74 downloads12d agoHugging Face10Arailym-tleubayeva /KazakhTextDuplicatesv2.0 KazakhTextDuplicates v2.0 KazakhTextDuplicates v2.0 is a large-scale dataset for duplicate detection, near-duplicate retrieval, semantic textual similarity (STS), and plagiarism detection in the Kazakh language. Version 2.0 significantly extends the dataset with: a large augmented training corpus (200K+ pairs), a continuous semantic similarity score (similarity_score), multiple difficulty levels of noisy duplicates, a clean train/validation/test split without identifier overlap. The… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhTextDuplicatesv2.0.sentence-similarity100K<n<1M1 likes52 downloads9mo agoHugging Face11Arailym-tleubayeva /legalup-laws Kazakhstan Legal Acts Dataset (LegalUp) Dataset Summary The LegalUp dataset contains structured metadata for legislative documents of the Republic of Kazakhstan. The current release includes 392,084 legislative document records extracted from a PostgreSQL database. The dataset is designed for: Legal Retrieval-Augmented Generation (Legal RAG) Information Retrieval Legal Search Question Answering Semantic Search Legal NLP Benchmark Construction Academic Research… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/legalup-laws.textquestion-answering100K<n<1M0 likes50 downloads12d agoHugging Face12CEAai /bop_distrib_tless BOP-Distrib dataset Dataset Description BOP-Distrib: Revisiting 6D Pose Estimation Benchmarks for Better Evaluation under Visual Ambiguities. This dataset contains the per-image pose distribution annotation for the T-LESS dataset (available here). The project page can be found at: https://cea-list.github.io/BOP-Distrib/ Citation Information If you use the BOP-Distrib annotations in your research, please cite the BOP-Distrib paper:… See the full description on the dataset page: https://huggingface.co/datasets/CEAai/bop_distrib_tless.0 likes44 downloads3mo agoHugging Face13Arailym-tleubayeva /NK-Oil-Well-Sensor-Monitoring Oil Well Sensor Monitoring Dataset - NK Field Dataset Description This dataset contains hourly sensor readings from 10 oil wells at the NK field. The monitoring period covers approximately seven months, from January 1, 2026, to July 20, 2026. The data were provided by Galaz and Company LLP (ТОО «Галаз и Компания») within the research project: “Development and Implementation of Control Algorithms for Low-Production-Rate Wells in Mechanized Oil Production Systems… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/NK-Oil-Well-Sensor-Monitoring.texttime-series-forecasting100K<n<1M0 likes40 downloads2mo agoHugging Face14TLeonidas /this-person-does-not-existThis dataset consists of 8892 AI-generated profile pictures downloaded from (https://www.kaggle.com/datasets/pablobedolla/this-person-does-not-exist-data) image1K<n<10K0 likes36 downloads2y agoHugging Face15Arailym-tleubayeva /KazOilWellOps_Dataset Kazakhstan Oil Well Operational Dataset Description This dataset contains structured operational and production parameters of sucker rod pump (SRP) oil wells in Kazakhstan. It is intended for industrial AI research, oil production analysis, production forecasting, and predictive modeling of well performance under real field operating conditions. Location: North-West Konys oil field, Kyzylorda Region, Kazakhstan (≈150 km NW of Kyzylorda city). Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazOilWellOps_Dataset.tabulartabular-regressionn<1K0 likes33 downloads7mo agoHugging Face16Arailym-tleubayeva /LegalRAG Kazakh Legal Text Chunks Dataset Summary Kazakh Legal Text Chunks is a processed corpus of official legal texts of the Republic of Kazakhstan, prepared for retrieval-augmented generation (RAG), legal information retrieval, and grounded legal question answering in the Kazakh language. The dataset contains structure-preserving text chunks derived from publicly available legal and normative documents. It is intended for research and development in: legal retrieval, legal QA… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/LegalRAG.tabularquestion-answering10K<n<100K0 likes33 downloads7mo agoHugging Face17Arailym-tleubayeva /small_kazakh_corpus Dataset Card for Small Kazakh Language Corpus The Small Kazakh Language Corpus is a specialized collection of textual data designed for training and research of natural language processing (NLP) models in the Kazakh language. The corpus is structured to ensure high text quality and comprehensive representation of diverse linguistic constructs. Dataset Details Dataset Description The dataset consists of Kazakh language texts with annotations that support tasks… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/small_kazakh_corpus.textmask-generation10K<n<100K1 likes27 downloads2y agoHugging Face18Arailym-tleubayeva /KazakhTextDuplicates Dataset Card for KazakhTextDuplicates Dataset Details Dataset Description The KazakhTextDuplicates dataset is a collection of Kazakh-language texts containing duplicates with different levels of modification. The dataset includes exact duplicates, contextual duplicates, and partial duplicates, making it valuable for research in text similarity, duplicate detection, information retrieval, and plagiarism detection. Developed by: Arailym Tleubayeva Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhTextDuplicates.text10K<n<100K1 likes20 downloads2y agoHugging Face19Arailym-tleubayeva /sist-kazakh-corpus SIST Kazakh Corpus Description SIST Kazakh Corpus is a curated dataset of Kazakh scientific articles collected for research in text similarity detection, plagiarism analysis, and low-resource NLP tasks. The dataset was created to support: Text similarity detection in agglutinative languages Kazakh NLP benchmarking Scientific text analysis Retrieval-Augmented Generation (RAG) research Dataset Structure The dataset is provided in CSV format. Columns may… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/sist-kazakh-corpus.tabularsentence-similarityn<1K1 likes19 downloads7mo agoHugging Face20nsimonato25 /tless_env_datasetimage10K<n<100K0 likes16 downloads2mo agoHugging Face21tleegwater /herb_sheetstextn<1K0 likes15 downloads1y agoHugging Face22Arailym-tleubayeva /AITUAdmissionsGuideDataset AITU Admissions Guide Dataset Dataset Details Dataset Description This dataset contains questions, answers, and categories related to the admission process at Astana IT University (AITU). It is designed to assist in automating applicant consultations and can be used for chatbot training, recommendation systems, and NLP-based question-answering models. Curated by: Astana IT University Funded by [optional]: Arailym Tleubayeva, Alina Mitroshina, Alpar Arman… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/AITUAdmissionsGuideDataset.texttable-question-answeringn<1K0 likes14 downloads2y agoHugging Face23tlemenestrel /AncestryOmicsUKB0 likes13 downloads1y agoHugging Face24SUSTech /tlem-leaderboardtabularn<1K0 likes11 downloads3y agoHugging Face25federicoarenas-ai /tless-5-objects-exampleimagen<1K0 likes10 downloads1y agoHugging Face26tleo /jozsef_attila_osszestextn<1K0 likes5 downloads6mo agoHugging Face27tleo /oi_docs_datasettextn<1K0 likes4 downloads1y agoHugging Face28Arailym-tleubayeva /sist-english-corpustabulartext-classificationn<1K0 likes4 downloads1y agoHugging Face29tleo /oi_docs_synthetic_alpacatext1K<n<10K0 likes3 downloads1y agoHugging Face30tleo /hungary_history_alpacatext10K<n<100K0 likes2 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.