datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Orbit_Planner
Orbit-Planner Orbital Evasion Dataset
Orbit-Planner is a simulated multimodal trajectory dataset for vision-based
spacecraft navigation and obstacle avoidance. It contains synchronized
first-person RGB images, depth maps, spacecraft states, thruster commands, and
event labels collected in the Orbital Evasion task from Space Robotics Bench
and NVIDIA Isaac Sim.
The dataset is intended for learning latent world models, spacecraft dynamics,
visual representations… See the full description on the dataset page: https://huggingface.co/datasets/warriorLZJ/Orbit_Planner.multilingual-data-v2LFM-Orbit-SatData
LFM Orbit SatData
Retagged Earth-observation training data produced by LFM Orbit for the Liquid AI x DPhi Space Hackathon.
The default viewer config is training_assets.jsonl, which contains single-image SFT rows with image, messages, and metadata. Temporal sequence rows live in the temporal_sft config so the Hugging Face Dataset Viewer does not try to cast sequence rows into the single-image schema.
Configs
Config
File
Purpose
default
training_assets.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Shoozes/LFM-Orbit-SatData.player-detectionorbit-mt-eval
ORBIT-MT-Eval: Expert Annotations and Prompts for Machine Translation Meta-Evaluation
ORBIT-MT-Eval is an English-to-Russian reference dataset for machine translation meta-evaluation: comparing evaluator outputs against professional expert annotations. Each of the 600 source–translation segments has three independent expert annotations following the RATE protocol. In Beyond Prompts: A Systematic Study of LLM-Based Machine Translation Evaluators, two annotations serve as gold… See the full description on the dataset page: https://huggingface.co/datasets/foksly/orbit-mt-eval.orbit-20k
[!NOTE]
For more information on the ORBIT dataset, go check out the preprint available at arxiv.org/abs/2604.01195.
ORBIT: A Synthetic Training Dataset for Search Agents
ORBITis a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Overview
Training data for deep search — tasks requiring multi-step retrieval and reasoning over the web — is scarce.… See the full description on the dataset page: https://huggingface.co/datasets/orbit-ai/orbit-20k.orbital-schemas
Orbital Schemas Dataset
Training data for OrbGen - a model that generates valid Orbital schemas (.orb files).
Dataset Structure
train: 142 examples
validation: 16 examples
test: 10 examples
Features
prompt: Natural language description of the desired schema
completion: Valid Orbital schema in JSON format
domain: Application domain (ecommerce, game, productivity, etc.)
complexity: Schema complexity (simple, medium, complex)
source: Source of the example… See the full description on the dataset page: https://huggingface.co/datasets/orbital-ai/orbital-schemas.bbh-orbit-lensing-gifs
Pokémon, behind a binary black hole
before
after
PUT YOUR POKEMON BEHIND BINARY BLACK HOLES
Artwork © Nintendo / Creatures / GAME FREAK. See LICENSE.
orbit-stage-3-27k
⚠️ [!Warning]
This is an unverified ORBIT dataset, only the Stage-3, and may contain data inaccuracies. The verified ORBIT dataset to use is orbit-ai/orbit-20k.
ORBIT: A Synthetic Training Dataset for Search Agents
ORBIT is a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Overview
Training data for deep search — tasks requiring multi-step… See the full description on the dataset page: https://huggingface.co/datasets/orbit-ai/orbit-stage-3-27k.multilingual-data-sampleorbit-seeds
[!NOTE]
For more information on the ORBIT dataset, go check out the preprint available at arxiv.org/abs/2604.01195.
ORBIT: A Synthetic Training Dataset for Search Agents
ORBITis a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Orbit Seeds
Seed entities collected from English Wikipedia, organised by domain. Each record is a Wikipedia page that… See the full description on the dataset page: https://huggingface.co/datasets/orbit-ai/orbit-seeds.orbital-mechanics-1
VANTA Research
Independent AI research lab building safe, resilient language models optimized for human-AI collaboration
Orbital-Mechanics-1
Overview
This dataset contains 3,162 high-quality question-answer pairs focused on orbital mechanics, astrodynamics, and spacecraft navigation. The content is designed for training large language models to understand and explain orbital dynamics concepts with mathematical rigor and physical… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/orbital-mechanics-1.titan-hohmann-transfer-orbit
🪐 Titan-Hohmann-Transfer-Orbit Dataset
🛰️ A 700-row synthetic dataset simulating an interplanetary mission to Titan using Hall Effect electric propulsion and gravity assists, from Earth departure through Titan orbital insertion.
⚠️ Disclaimer: All values are synthetically generated from simplified orbital mechanics models. This is not flight data and is not suitable for mission planning.
📋 At a Glance
🔢 Rows
700
📊 Columns
18
🧬… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/titan-hohmann-transfer-orbit.great-nurse-3638f1
great-nurse-3638f1
Synthetic products test data: 38 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/OrbitLens/great-nurse-3638f1.orbital-mechanics-instruct-32
🛰️ Orbital Mechanics Instruction Dataset
Expert-crafted prompt–completion pairs for fine-tuning LLMs on space mission analysis and design
📋 Dataset Summary
A curated dataset of 32 expert-crafted instruction–completion pairs designed for fine-tuning large language models on orbital mechanics and space mission analysis tasks. Each example contains a natural-language problem statement paired with a structured, step-by-step solution featuring properly formatted… See the full description on the dataset page: https://huggingface.co/datasets/abdohisham12/orbital-mechanics-instruct-32.solar-orbiter-encounters
Solar Orbiter Encounter Timeline
Credit: NASA/SDO
Part of a dataset collection on Hugging Face.
Dataset description
Complete mission event timeline for the ESA/NASA Solar Orbiter — the first spacecraft designed to deliver sustained close-up views of the Sun's polar regions. Compiled from ESA Solar Orbiter operations documents and the Mueller et al. 2020 mission overview paper (A&A 642, A1).
The dataset covers all perihelion encounters from P1 (2020-06-15… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/solar-orbiter-encounters.orbit-stage-2-27k
⚠️ [!Warning]
This is an unverified ORBIT dataset, only the Stage-2, and may contain data inaccuracies. The verified ORBIT dataset to use is orbit-ai/orbit-20k.
ORBIT: A Synthetic Training Dataset for Search Agents
ORBIT is a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Overview
Training data for deep search — tasks requiring multi-step… See the full description on the dataset page: https://huggingface.co/datasets/orbit-ai/orbit-stage-2-27k.624-hektor-hohmann-transfer-orbit
🪨 Hektor-Hohmann Transfer Orbit Dataset
🛰️ ~30,000 synthetic time-series rows for an interplanetary mission to asteroid 624 Hektor: LEO departure, heliocentric transfer with electric propulsion thrust arcs, Jupiter gravity assist, and orbit capture at Hektor.
⚠️ Disclaimer: All values are synthetically generated placeholders, produced by random or linear interpolation over the mission timeline. They are not propagated trajectories and are not from any real or planned… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/624-hektor-hohmann-transfer-orbit.green-fact-0873d5
green-fact-0873d5
Synthetic sensors test data: 57 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Orbit-Richard/green-fact-0873d5.exciting-library-21ed94
exciting-library-21ed94
Synthetic products test data: 40 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/OrbitRidge23/exciting-library-21ed94.ORBIT ORBIT: An Object Property Reasoning Benchmark for Visual Inference Tasks
Inspired by human object categorization, this repository introduces ORBIT, a comprehensive benchmark designed to evaluate the abilities of Vision–Language Models (VLMs) to reason about abstract object properties. ORBIT spans four object property dimensions (physical, taxonomic, functional, relational), three levels of reasoning complexity (direct recognition, property inference, counterfactual reasoning), and three… See the full description on the dataset page: https://huggingface.co/datasets/Abk802/ORBIT.orbit-stage-1-44k
⚠️ [!Warning]
This is an unverified ORBIT dataset, only the Stage-1, and may contain data inaccuracies. The verified ORBIT dataset to use is orbit-ai/orbit-20k.
ORBIT: A Synthetic Training Dataset for Search Agents
ORBIT is a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Overview
Training data for deep search — tasks requiring multi-step… See the full description on the dataset page: https://huggingface.co/datasets/orbit-ai/orbit-stage-1-44k.BOOK-RECOMMENDER-DATASET
Book Recommender Dataset
CSV exports from my Book Recommender pipeline. Includes cleaned metadata, category labels, emotion tags, and a tagged description file.
Files
books_cleaned.csv: Core cleaned book metadata.
books_with_categories.csv: Adds multi-label categories column.
books_with_emotions.csv: Adds emotion_* columns (one-hot or scores).
tagged_description.txt: Preprocessed descriptions (one per line, or TSV).
Column Schema (example)… See the full description on the dataset page: https://huggingface.co/datasets/orbitk/BOOK-RECOMMENDER-DATASET.Orbit-200K
🪐 Orbit-200K
A high-signal instruction dataset designed for training universal language models with dense reasoning, coding, mathematics, and general intelligence—without conversational bloat.
Orbit-200K is a carefully curated dataset of 200,000 instruction-response pairs optimized for training modern language models ranging from 0.5B to 7B parameters.
Unlike many public instruction datasets, Orbit-200K removes unnecessary conversational filler and focuses on maximizing… See the full description on the dataset page: https://huggingface.co/datasets/Pluto-AI-Labs/Orbit-200K.final-father-675ef0
final-father-675ef0
Synthetic weather test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Orbit-Noah/final-father-675ef0.objective-term-8dbb54
objective-term-8dbb54
Synthetic sensors test data: 52 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/orbitMatthew/objective-term-8dbb54.ORBIT-Astrowestern-wedding-a276a6
western-wedding-a276a6
Synthetic sensors test data: 36 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/orbitglade/western-wedding-a276a6.senior-accident-1dd433
senior-accident-1dd433
Synthetic sensors test data: 60 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Orbit-Loom/senior-accident-1dd433.orbital-mining-corpus
Orbital Mining Corporation (OMC) Corpus
A synthetic domain-adaptation corpus for continued pre-training (mid-training) of language models on technical aerospace and deep-space mining documentation.
Dataset Summary
This corpus contains 1,000 long-form technical documents (~10M tokens) representing the internal document ecosystem of Orbital Mining Corporation (OMC) — a fictional company operating crewed spacecraft and extracting resources from the asteroid belt. All… See the full description on the dataset page: https://huggingface.co/datasets/atenareply/orbital-mining-corpus.
