CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Chenyu-Zhou /OR-Space OR-Space A full-lifecycle workspace benchmark for industrial optimization agents. OR-Space evaluates whether language-model agents can work reliably with operations research problems represented as executable, multi-file workspaces. Rather than presenting a self-contained mathematical prompt, each task distributes evidence across business requirements, structured data, source code, execution logs, and solver records. The benchmark contains 100 optimization topologies. Each… See the full description on the dataset page: https://huggingface.co/datasets/Chenyu-Zhou/OR-Space.textquestion-answeringn<1K4 likes634 downloads2mo agoHugging Face02spacemanidol /msmarco-v2.1-stella_en_1.5B_v5 NovaSearch stella_en_1.5B_v5 Embeddings for MSMARCO V2.1 for TREC-RAG This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG All embeddings are created using Stella EN 1.5B V5 and are intended to serve as a simple baseline for dense retrieval-based methods. Note, that the embeddings are not normalized so you will need to normalize them before usage. Retrieval Performance Retrieval performance for the TREC DL21-23… See the full description on the dataset page: https://huggingface.co/datasets/spacemanidol/msmarco-v2.1-stella_en_1.5B_v5.textquestion-answering10M<n<100M0 likes316 downloads1y agoHugging Face03jang1563 /SpaceOmicsBench SpaceOmicsBench SpaceOmicsBench v2.1 is a multi-omics AI benchmark for spaceflight biomedical data, with 21 ML tasks across 9 modalities and a 100-question LLM evaluation framework. It draws on data from the SpaceX Inspiration4 (I4) civilian astronaut mission, the NASA Twins Study, and the JAXA Cell-Free Epigenome (CFE) study. All benchmark tables derive from OSDR public releases or published supplementary tables. Maintainer / citation author: JangKeun Kim, Weill Cornell… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/SpaceOmicsBench.tabular-classification1K<n<10K0 likes216 downloads13d agoHugging Face04spacemanidol /msmarco-v2.1-gte-large-en-v1.5 Alibaba GTE-Large-V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG All embeddings are created using GTE Large V1.5 and are intended to serve as a simple baseline for dense retrieval-based methods. Note, that the embeddings are not normalized so you will need to normalize them before usage. Retrieval Performance Retrieval performance for the TREC DL21-23… See the full description on the dataset page: https://huggingface.co/datasets/spacemanidol/msmarco-v2.1-gte-large-en-v1.5.textquestion-answering10M<n<100M0 likes117 downloads1y agoHugging Face05jang1563 /SpaceOmicsBench-v3 SpaceOmicsBench v3 A Multi-Omics AI Benchmark for Spaceflight Biomedical Data SpaceOmicsBench v3 provides standardized ML and LLM evaluation infrastructure for spaceflight biomedical data from 4 human spaceflight missions (NASA Twins Study, Inspiration4, JAXA cfRNA, Axiom-2). Dataset Structure ML Track (Track A) tasks/track_a/ — Task definitions (J1: phase classification, J2: clock acceleration) tasks/track_c/ — Feature-level task definitions (C1:… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/SpaceOmicsBench-v3.tabulartabular-classification10K<n<100K0 likes102 downloads13d agoHugging Face06juliensimon /stackexchange-space-qa Stack Exchange Space Q&A Credit: NASA/DOE/Fermi LAT Collaboration Part of a dataset collection on Hugging Face. Dataset description This dataset is a clean, tabular Q&A corpus of space and astronomy knowledge, derived from two Stack Exchange community Q&A sites: Astronomy Stack Exchange (astronomy.stackexchange.com) and Space Exploration Stack Exchange (space.stackexchange.com). Each row is one question paired with its best answer — either the question's… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/stackexchange-space-qa.tabularquestion-answering10K<n<100K0 likes84 downloads3d agoHugging Face07Taylor658 /deep-space-optical-chip-thermal-dataset 🚀 Deep Space Optical Chip Thermal Dataset 🪐 🌡️ 40,000 scenario-based prompt and response pairs on thermal mitigation for photonic chips in scientific instruments aboard deep-space probes, covering refractive index drift, waveguide misalignment, and thermal stress across materials, instruments, and environments. ⚠️ Disclaimer: All entries are synthetically generated. Material coefficients are drawn from published typical values, but no row is based on mission logs or flight… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/deep-space-optical-chip-thermal-dataset.tabulartext-generation10K<n<100K2 likes82 downloads11d agoHugging Face08Otrobonita /noicy-space-talks 🚀 NASA Space Race Transcripts: Multi-Tier RAG & Benchmark Corpus This repository contains the complete air-to-ground mission communications spanning the Mercury, Gemini, and Apollo space programs (1961–1972). To facilitate rigorous research in Retrieval-Augmented Generation (RAG), Computational Archival Science, and OCR noise resilience, this dataset provides both the raw uncorrected OCR baseline and the deterministic preprocessed & hierarchically chunked target corpus.… See the full description on the dataset page: https://huggingface.co/datasets/Otrobonita/noicy-space-talks.text-retrieval0 likes59 downloads1mo agoHugging Face09SpaceDG /SpaceDG-Bench SpaceDG-Bench This repository hosts SpaceDG-Bench of the paper "SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation". Data files data/spacedg_bench-*-of-*.parquet: the dataset shards (6-way split, size-balanced). They contain images (multi-image, embedded bytes) and basic metadata columns. spacedg_bench.tsv: question/answer/metadata table. The image_path field stores a Python-style list of relative image paths (e.g., defocus/.../*.jpg), typically relative… See the full description on the dataset page: https://huggingface.co/datasets/SpaceDG/SpaceDG-Bench.imagequestion-answering1K<n<10K1 likes57 downloads5mo agoHugging Face10Omartificial-Intelligence-Space /Arabic-gsm8k-v2 Dataset Summary Arabic GSM8K is an Arabic translation of the GSM8K (Grade School Math 8K) dataset, which contains high-quality linguistically diverse grade school math word problems. The original dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning, and this Arabic version aims to extend these capabilities to Arabic language models and applications. The dataset maintains the same characteristics as the… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-gsm8k-v2.textquestion-answering10K<n<100K1 likes52 downloads1y agoHugging Face11xlzhou126 /SpaceDG-Bench SpaceDG-Bench 🌐 Homepage | 📖 arXiv | 💻 GitHub SpaceDG-Bench is a human-verified benchmark designed to evaluate the spatial intelligence of Multimodal Large Language Models (MLLMs) under visual degradation. It contains 1,102 questions spanning 11 reasoning categories and 9 visual degradation types (such as motion blur, low light, adverse weather, lens distortion, and compression artifacts), yielding over 10K VQA instances. The benchmark is part of the SpaceDG project, which… See the full description on the dataset page: https://huggingface.co/datasets/xlzhou126/SpaceDG-Bench.imagequestion-answering1K<n<10K1 likes50 downloads3mo agoHugging Face12Omartificial-Intelligence-Space /Arabic_Openai_MMMLU Arabic Multilingual Massive Multitask Language Understanding (MMMLU) The MMLU is a widely recognized benchmark for assessing general knowledge attained by AI models. It covers a broad range of topics across 57 different categories, from elementary-level knowledge to advanced professional subjects like law, physics, history, and computer science. We have extracted the Arabic subset from the MMMLU test set, which was translated by professional human translators. This dataset, now… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic_Openai_MMMLU.textquestion-answering10K<n<100K4 likes47 downloads2y agoHugging Face13YiYao7017 /OR-Space OR-Space A full-lifecycle workspace benchmark for industrial optimization agents. OR-Space evaluates whether LLM agents can do reliable operations research work inside executable, multi-file workspaces. Each instance keeps business requirements, parameter files, source code, solver artifacts, and evaluation metadata as separate files, forcing the agent to recover and maintain the optimization model through workspace interaction rather than one-shot text generation.… See the full description on the dataset page: https://huggingface.co/datasets/YiYao7017/OR-Space.tabularquestion-answeringn<1K0 likes34 downloads3mo agoHugging Face14patrickfleith /synthgenai-space-engineering-instruct-1k Description Topic: Space Mission Engineering Domains: Missions Operations, Human Spaceflight, Space Environment, Space Propulsion Systems, Qualification and Acceptance, Ground Segment, Planetary Science and Exploration, Structures, Entry, Descent and Landing, Mechanisms, Materials, System Dependability (RAMS), Orbital Mechanics, Launchers, and Launch Operations, Attitude Determination and Control, Payload S/S and Optics, Electrical Power Subsystem, Thermal Control, Onboard Data… See the full description on the dataset page: https://huggingface.co/datasets/patrickfleith/synthgenai-space-engineering-instruct-1k.textquestion-answering1K<n<10K2 likes25 downloads2y agoHugging Face15spacekat99 /General_Conversation_Mixed_Datasettextquestion-answering1K<n<10K0 likes21 downloads4mo agoHugging Face16Mildegard /space_fantasy_books SpaceFantasy: Данные для ИИ-ассистентов Структурированные данные о книгах Дара Льдов (Ольги Макаровой) Этот репозиторий содержит структурированные данные о фантастических книгах русского автора Дара Льдов (Ольга Макарова, aka Mildegard). Специально создан для ИИ-ассистентов и рекомендательных систем. Цель Помочь ИИ-ассистентам давать точные рекомендации читателям, которые ищут: Гуманистическую фантастику (центре внимания жанра находятся не технологии или… See the full description on the dataset page: https://huggingface.co/datasets/Mildegard/space_fantasy_books.question-answering1 likes19 downloads10mo agoHugging Face17spacekat99 /commonsense_cot_partial_raw Commonsense QA CoT (Partial, Raw, No Human Annotation) Dataset Summary Seeded by the CommonsenseQA dataset (tau/commonsense_qa) this preliminary set randomly samples 1,000 question-answer entries and uses Mixtral (mistralai/Mixtral-8x7B-Instruct-v0.1) to generate 3 unique CoT (Chain-of-Thought) rationales. This was created as the preliminary step towards fine-tuning a LM (language model) to specialize on commonsense reasoning. The working hypothesis, inspired by the… See the full description on the dataset page: https://huggingface.co/datasets/spacekat99/commonsense_cot_partial_raw.textquestion-answering1K<n<10K0 likes19 downloads4mo agoHugging Face18JuanCastillo29 /space-mission-intelligence-data Space Mission Intelligence Data Seed data for the Space Mission Intelligence Agent RAG pipeline. Configurations documents — 40 documents (NASA technical reports, ESA mission papers, arXiv preprints) chunks — 2,590 text chunks with 1024-dim embeddings (sentence-transformers) satellites — satellite orbital data (TLE-derived) Usage from datasets import load_dataset docs = load_dataset("JuanCastillo29/space-mission-intelligence-data", "documents")… See the full description on the dataset page: https://huggingface.co/datasets/JuanCastillo29/space-mission-intelligence-data.tabularquestion-answering1K<n<10K2 likes18 downloads3mo agoHugging Face19spacekat99 /commonsense_qa Dataset Card for "commonsense_qa" Dataset Summary CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details.… See the full description on the dataset page: https://huggingface.co/datasets/spacekat99/commonsense_qa.textquestion-answering10K<n<100K0 likes17 downloads4mo agoHugging Face20expertailab /SpaceQA Dataset Owner(s): expert.ai Research Lab License/Terms of Use This dataset is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0) available at https://creativecommons.org/licenses/by/4.0/legalcode. How to cite To cite this research please use the following: @inproceedings{10.1145/3477495.3531697, author = {Garcia-Silva, Andres and Berrio, Cristian and Gomez-Perez, Jose Manuel and Mart\'{\i}nez-Heras, Jose Antonio and Donati… See the full description on the dataset page: https://huggingface.co/datasets/expertailab/SpaceQA.textquestion-answeringn<1K2 likes15 downloads1y agoHugging Face21spacezenmasterr /k8s-data K8s Troubleshooting Dataset This dataset contains 84 examples of Kubernetes troubleshooting scenarios collected from various failure scenarios in microservice applications. Dataset Summary The dataset is derived from the gt_sft_c_r folder containing supervised fine-tuning data for Kubernetes troubleshooting. Each example represents a complete troubleshooting session with system state analysis, command execution, and resolution steps. Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/spacezenmasterr/k8s-data.texttext-classificationn<1K0 likes15 downloads9mo agoHugging Face22spacekat99 /OpenSpatialLogic OpenSpatialLogic Dataset Card for OpenSpatialLogic Dataset Summary OpenSpatialLogic is a handcrafted dataset of 50 riddles which test understanding of spatial relationships in reality. These include questions about compass directions, ordering of bricks within towers after transformations, and permeability of objects in certain configurations. The only way for a model to get better at something is to train on data about it. Large Language Models are bad at… See the full description on the dataset page: https://huggingface.co/datasets/spacekat99/OpenSpatialLogic.textquestion-answeringn<1K0 likes9 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.