datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OR-Space
OR-Space
A full-lifecycle workspace benchmark for industrial optimization agents.
OR-Space evaluates whether language-model agents can work reliably with
operations research problems represented as executable, multi-file workspaces.
Rather than presenting a self-contained mathematical prompt, each task
distributes evidence across business requirements, structured data, source
code, execution logs, and solver records.
The benchmark contains 100 optimization topologies. Each… See the full description on the dataset page: https://huggingface.co/datasets/Chenyu-Zhou/OR-Space.msmarco-v2.1-stella_en_1.5B_v5
NovaSearch stella_en_1.5B_v5 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Stella EN 1.5B V5 and are intended to serve as a simple baseline for dense retrieval-based methods.
Note, that the embeddings are not normalized so you will need to normalize them before usage.
Retrieval Performance
Retrieval performance for the TREC DL21-23… See the full description on the dataset page: https://huggingface.co/datasets/spacemanidol/msmarco-v2.1-stella_en_1.5B_v5.SpaceOmicsBench
SpaceOmicsBench
SpaceOmicsBench v2.1 is a multi-omics AI benchmark for spaceflight biomedical
data, with 21 ML tasks across 9 modalities and a 100-question LLM
evaluation framework. It draws on data from the SpaceX Inspiration4 (I4)
civilian astronaut mission, the NASA Twins Study, and the JAXA Cell-Free
Epigenome (CFE) study. All benchmark tables derive from OSDR public releases or
published supplementary tables.
Maintainer / citation author: JangKeun Kim, Weill Cornell… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/SpaceOmicsBench.msmarco-v2.1-gte-large-en-v1.5
Alibaba GTE-Large-V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using GTE Large V1.5 and are intended to serve as a simple baseline for dense retrieval-based methods.
Note, that the embeddings are not normalized so you will need to normalize them before usage.
Retrieval Performance
Retrieval performance for the TREC DL21-23… See the full description on the dataset page: https://huggingface.co/datasets/spacemanidol/msmarco-v2.1-gte-large-en-v1.5.SpaceOmicsBench-v3
SpaceOmicsBench v3
A Multi-Omics AI Benchmark for Spaceflight Biomedical Data
SpaceOmicsBench v3 provides standardized ML and LLM evaluation infrastructure for spaceflight biomedical data from 4 human spaceflight missions (NASA Twins Study, Inspiration4, JAXA cfRNA, Axiom-2).
Dataset Structure
ML Track (Track A)
tasks/track_a/ — Task definitions (J1: phase classification, J2: clock acceleration)
tasks/track_c/ — Feature-level task definitions (C1:… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/SpaceOmicsBench-v3.stackexchange-space-qa
Stack Exchange Space Q&A
Credit: NASA/DOE/Fermi LAT Collaboration
Part of a dataset collection on Hugging Face.
Dataset description
This dataset is a clean, tabular Q&A corpus of space and astronomy knowledge, derived from two Stack Exchange community Q&A sites: Astronomy Stack Exchange (astronomy.stackexchange.com) and Space Exploration Stack Exchange (space.stackexchange.com). Each row is one question paired with its best answer — either the question's… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/stackexchange-space-qa.deep-space-optical-chip-thermal-dataset
🚀 Deep Space Optical Chip Thermal Dataset 🪐
🌡️ 40,000 scenario-based prompt and response pairs on thermal mitigation for photonic chips in scientific instruments aboard deep-space probes, covering refractive index drift, waveguide misalignment, and thermal stress across materials, instruments, and environments.
⚠️ Disclaimer: All entries are synthetically generated. Material coefficients are drawn from published typical values, but no row is based on mission logs or flight… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/deep-space-optical-chip-thermal-dataset.noicy-space-talks
🚀 NASA Space Race Transcripts: Multi-Tier RAG & Benchmark Corpus
This repository contains the complete air-to-ground mission communications spanning the Mercury, Gemini, and Apollo space programs (1961–1972).
To facilitate rigorous research in Retrieval-Augmented Generation (RAG), Computational Archival Science, and OCR noise resilience, this dataset provides both the raw uncorrected OCR baseline and the deterministic preprocessed & hierarchically chunked target corpus.… See the full description on the dataset page: https://huggingface.co/datasets/Otrobonita/noicy-space-talks.SpaceDG-Bench
SpaceDG-Bench
This repository hosts SpaceDG-Bench of the paper "SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation".
Data files
data/spacedg_bench-*-of-*.parquet: the dataset shards (6-way split, size-balanced). They contain images (multi-image, embedded bytes) and basic metadata columns.
spacedg_bench.tsv: question/answer/metadata table. The image_path field stores a Python-style list of relative image paths (e.g., defocus/.../*.jpg), typically relative… See the full description on the dataset page: https://huggingface.co/datasets/SpaceDG/SpaceDG-Bench.Arabic-gsm8k-v2
Dataset Summary
Arabic GSM8K is an Arabic translation of the GSM8K (Grade School Math 8K) dataset, which contains high-quality linguistically diverse grade school math word problems. The original dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning, and this Arabic version aims to extend these capabilities to Arabic language models and applications.
The dataset maintains the same characteristics as the… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-gsm8k-v2.SpaceDG-Bench
SpaceDG-Bench
🌐 Homepage | 📖 arXiv | 💻 GitHub
SpaceDG-Bench is a human-verified benchmark designed to evaluate the spatial intelligence of Multimodal Large Language Models (MLLMs) under visual degradation. It contains 1,102 questions spanning 11 reasoning categories and 9 visual degradation types (such as motion blur, low light, adverse weather, lens distortion, and compression artifacts), yielding over 10K VQA instances.
The benchmark is part of the SpaceDG project, which… See the full description on the dataset page: https://huggingface.co/datasets/xlzhou126/SpaceDG-Bench.Arabic_Openai_MMMLU
Arabic Multilingual Massive Multitask Language Understanding (MMMLU)
The MMLU is a widely recognized benchmark for assessing general knowledge attained by AI models. It covers a broad range of topics across 57 different categories, from elementary-level knowledge to advanced professional subjects like law, physics, history, and computer science.
We have extracted the Arabic subset from the MMMLU test set, which was translated by professional human translators. This dataset, now… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic_Openai_MMMLU.OR-Space
OR-Space
A full-lifecycle workspace benchmark for industrial optimization agents.
OR-Space evaluates whether LLM agents can do reliable operations research work
inside executable, multi-file workspaces. Each instance keeps business
requirements, parameter files, source code, solver artifacts, and evaluation
metadata as separate files, forcing the agent to recover and maintain the
optimization model through workspace interaction rather than one-shot text
generation.… See the full description on the dataset page: https://huggingface.co/datasets/YiYao7017/OR-Space.synthgenai-space-engineering-instruct-1k
Description
Topic: Space Mission Engineering
Domains: Missions Operations, Human Spaceflight, Space Environment, Space Propulsion Systems, Qualification and Acceptance, Ground Segment, Planetary Science and Exploration, Structures, Entry, Descent and Landing, Mechanisms, Materials, System Dependability (RAMS), Orbital Mechanics, Launchers, and Launch Operations, Attitude Determination and Control, Payload S/S and Optics, Electrical Power Subsystem, Thermal Control, Onboard Data… See the full description on the dataset page: https://huggingface.co/datasets/patrickfleith/synthgenai-space-engineering-instruct-1k.General_Conversation_Mixed_Datasetspace_fantasy_books
SpaceFantasy: Данные для ИИ-ассистентов
Структурированные данные о книгах Дара Льдов (Ольги Макаровой)
Этот репозиторий содержит структурированные данные о фантастических книгах русского автора Дара Льдов (Ольга Макарова, aka Mildegard). Специально создан для ИИ-ассистентов и рекомендательных систем.
Цель
Помочь ИИ-ассистентам давать точные рекомендации читателям, которые ищут:
Гуманистическую фантастику (центре внимания жанра находятся не технологии или… See the full description on the dataset page: https://huggingface.co/datasets/Mildegard/space_fantasy_books.commonsense_cot_partial_raw
Commonsense QA CoT (Partial, Raw, No Human Annotation)
Dataset Summary
Seeded by the CommonsenseQA dataset (tau/commonsense_qa) this preliminary set randomly samples 1,000 question-answer
entries and uses Mixtral (mistralai/Mixtral-8x7B-Instruct-v0.1) to generate 3 unique CoT (Chain-of-Thought) rationales.
This was created as the preliminary step towards fine-tuning a LM (language model) to specialize on commonsense reasoning.
The working hypothesis, inspired by the… See the full description on the dataset page: https://huggingface.co/datasets/spacekat99/commonsense_cot_partial_raw.space-mission-intelligence-data
Space Mission Intelligence Data
Seed data for the Space Mission Intelligence Agent RAG pipeline.
Configurations
documents — 40 documents (NASA technical reports, ESA mission papers, arXiv preprints)
chunks — 2,590 text chunks with 1024-dim embeddings (sentence-transformers)
satellites — satellite orbital data (TLE-derived)
Usage
from datasets import load_dataset
docs = load_dataset("JuanCastillo29/space-mission-intelligence-data", "documents")… See the full description on the dataset page: https://huggingface.co/datasets/JuanCastillo29/space-mission-intelligence-data.commonsense_qa
Dataset Card for "commonsense_qa"
Dataset Summary
CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge
to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers.
The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation
split, and "Question token split", see paper for details.… See the full description on the dataset page: https://huggingface.co/datasets/spacekat99/commonsense_qa.SpaceQA
Dataset Owner(s):
expert.ai Research Lab
License/Terms of Use
This dataset is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0) available at https://creativecommons.org/licenses/by/4.0/legalcode.
How to cite
To cite this research please use the following:
@inproceedings{10.1145/3477495.3531697,
author = {Garcia-Silva, Andres and Berrio, Cristian and Gomez-Perez, Jose Manuel and Mart\'{\i}nez-Heras, Jose Antonio and Donati… See the full description on the dataset page: https://huggingface.co/datasets/expertailab/SpaceQA.k8s-data
K8s Troubleshooting Dataset
This dataset contains 84 examples of Kubernetes troubleshooting scenarios collected from various failure scenarios in microservice applications.
Dataset Summary
The dataset is derived from the gt_sft_c_r folder containing supervised fine-tuning data for Kubernetes troubleshooting. Each example represents a complete troubleshooting session with system state analysis, command execution, and resolution steps.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/spacezenmasterr/k8s-data.OpenSpatialLogic
OpenSpatialLogic
Dataset Card for OpenSpatialLogic
Dataset Summary
OpenSpatialLogic is a handcrafted dataset of 50 riddles which test understanding of spatial relationships in reality.
These include questions about compass directions, ordering of bricks within towers after transformations, and permeability of objects in certain configurations.
The only way for a model to get better at something is to train on data about it. Large Language Models are bad at… See the full description on the dataset page: https://huggingface.co/datasets/spacekat99/OpenSpatialLogic.
