datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SSR-3DFRONT
SSR-3DFRONT: Structured Scene Representation for 3D Indoor Scenes
This dataset provides a processed version of the 3D-FRONT dataset with structured scene representations for text-driven 3D indoor scene synthesis and editing.
Mor information about ReSpace: http://respace.mnbucher.com
For detailed usage instructions, training details, and examples, see the associated repository: https://github.com/GradientSpaces/respace
Our model weights for SG-LLM:… See the full description on the dataset page: https://huggingface.co/datasets/gradient-spaces/SSR-3DFRONT.OR-Space
OR-Space
A full-lifecycle workspace benchmark for industrial optimization agents.
OR-Space evaluates whether language-model agents can work reliably with
operations research problems represented as executable, multi-file workspaces.
Rather than presenting a self-contained mathematical prompt, each task
distributes evidence across business requirements, structured data, source
code, execution logs, and solver records.
The benchmark contains 100 optimization topologies. Each… See the full description on the dataset page: https://huggingface.co/datasets/Chenyu-Zhou/OR-Space.SpaceCode-Bench
SpaceCode-Bench
Construct, reconstruct and edit native 3D environments with Blender code.
Website · Code and evaluator · Protocol · Complete worked example
SpaceCode-Bench is a procedural development benchmark with 1,200 materialized scenes, 3,600 task bundles and 4,800 execution conditions. Each task publishes a complete numerical contract. Evaluation reopens the native .blend, extracts geometry and controlled views, and measures the declared requirements.
Construction starts… See the full description on the dataset page: https://huggingface.co/datasets/KurtDu/SpaceCode-Bench.zhtw-roleplay-space-grimoire
Space Grimoire RP Corpus (Traditional Chinese)
Speaker-attributed dialogue from the original novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage; 283 chapters, ~1.7M characters), cut into scenes and assembled into ShareGPT-style role-play training data. The novel and this dataset are the work of 睡半夜怎麼三更, who holds the copyright and has no exclusive platform agreement. Data: CC BY 4.0. Code: Apache 2.0.
中文說明在下方
Dataset Summary
Source text
283… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/zhtw-roleplay-space-grimoire.OpenThought3-Qwen3-4BOpenThought3-Qwen3-4B
OpenThought3-Qwen3-4B is a math reasoning supervised fine-tuning dataset in chat-message JSONL format.
Data Creation and Cleaning
This dataset was generated by Qwen3-4B (Non-thinking) from math-domain prompts selected from OpenThoughts3-1.2M. The generated responses were cleaned through deduplication, removal of degenerate repetition/repeater-style outputs, and template checks on the assistant… See the full description on the dataset page: https://huggingface.co/datasets/Thinking-Space/OpenThought3-Qwen3-4B.isro-space-ocean-dataset
ISRO Multimodal Space & Ocean Telemetry Dataset
Official open-source scientific dataset curated for the National Space Day 2026 Hackathon and ISRO/IN-SPACe research submissions.
Dataset Structure
rain.jsonl: 1,204 high-precision instruction-tuning pairs mapping 6-band multispectral satellite telemetry (Coastal, Blue, Green, Red, NIR, SWIR) to atmospheric composition (O2 %, N2 %, Water Vapor g/m3) and oceanographic parameters (SST deg C, Salinity PSU).… See the full description on the dataset page: https://huggingface.co/datasets/Anoopsingh53/isro-space-ocean-dataset.space-llm-training-data
Space LLM Training Data (~1.27 Billion Tokens)
A curated dataset of space and astronomy text for training language models, containing approximately 1.27 billion tokens collected from academic papers, arXiv abstracts, and educational web content.
Dataset Summary
File
Size
Est. Tokens
Source
jsalt_astroph_full.txt
2.88 GB
~862M
271K full astrophysics papers (abstract + introduction + conclusions)
arxiv_astro_full.txt
360 MB
~108M
284K arXiv paper… See the full description on the dataset page: https://huggingface.co/datasets/Ashu9675/space-llm-training-data.deep-space-optical-chip-thermal-dataset
🚀 Deep Space Optical Chip Thermal Dataset 🪐
🌡️ 40,000 scenario-based prompt and response pairs on thermal mitigation for photonic chips in scientific instruments aboard deep-space probes, covering refractive index drift, waveguide misalignment, and thermal stress across materials, instruments, and environments.
⚠️ Disclaimer: All entries are synthetically generated. Material coefficients are drawn from published typical values, but no row is based on mission logs or flight… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/deep-space-optical-chip-thermal-dataset.stackexchange-space-qa
Stack Exchange Space Q&A
Credit: NASA/DOE/Fermi LAT Collaboration
Part of a dataset collection on Hugging Face.
Dataset description
This dataset is a clean, tabular Q&A corpus of space and astronomy knowledge, derived from two Stack Exchange community Q&A sites: Astronomy Stack Exchange (astronomy.stackexchange.com) and Space Exploration Stack Exchange (space.stackexchange.com). Each row is one question paired with its best answer — either the question's… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/stackexchange-space-qa.ssao-space-instruct
ssao-space-instruct
Instruction data teaching a language model to write valid RDF Turtle in the
Space Situational Awareness Ontology (SSAO) for real space objects, and to
judge proposed catalogue-to-ontology alignments using instance evidence.
1,245 examples: 1,072 train, 74 validation, 99 test. Built by
Tesseract Academy.
The construction principle
No example asserts anything a validator cannot check. Every Turtle target is
generated from a real CelesTrak SATCAT… See the full description on the dataset page: https://huggingface.co/datasets/fabsssss/ssao-space-instruct.mental-spaces
Mental Spaces Corpus
Version: 0.1.0
The Mental Spaces Corpus is a controlled suite of natural-language stimuli for testing
whether language models keep base-space and alternative-space discourse targets
separate. It is designed for probing, causal interventions, and behavioral readouts in
mental-space constructions such as counterfactuals, belief contexts, and depictive
spaces, including nested belief and nested depictive spaces.
This release is a stimulus suite for controlled… See the full description on the dataset page: https://huggingface.co/datasets/osteele/mental-spaces.OR-Space
OR-Space
A full-lifecycle workspace benchmark for industrial optimization agents.
OR-Space evaluates whether LLM agents can do reliable operations research work
inside executable, multi-file workspaces. Each instance keeps business
requirements, parameter files, source code, solver artifacts, and evaluation
metadata as separate files, forcing the agent to recover and maintain the
optimization model through workspace interaction rather than one-shot text
generation.… See the full description on the dataset page: https://huggingface.co/datasets/YiYao7017/OR-Space.General_Conversation_Mixed_DatasetNemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 통계청(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.
Nemotron-Personas-Korea는… See the full description on the dataset page: https://huggingface.co/datasets/space-mind/Nemotron-Personas-Korea.k8s-data
K8s Troubleshooting Dataset
This dataset contains 84 examples of Kubernetes troubleshooting scenarios collected from various failure scenarios in microservice applications.
Dataset Summary
The dataset is derived from the gt_sft_c_r folder containing supervised fine-tuning data for Kubernetes troubleshooting. Each example represents a complete troubleshooting session with system state analysis, command execution, and resolution steps.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/spacezenmasterr/k8s-data.awesome_chatgpt_prompts_ar
📦 Awesome Arabic Chatgpt Prompts
📝 Overview
This repository contains a collection of Arabic prompts designed for use with AI language models (such as ChatGPT).
The goal is to provide a lightweight dataset that helps Arabic-speaking users quickly get started with generative AI.
🔗 Website / Demo
Check out the live demo site:omarnj-lab.github.io/awesome_chatgpt_prompts_ar
✨ Features
Entirely in Arabic 🕌
Suitable for educational and… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/awesome_chatgpt_prompts_ar.state-space-models-papers
State Space Models & Mamba Papers — FineSet
A research-paper dataset on State Space Models & Mamba Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on State Space Models & Mamba Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/state-space-models-papers.
