datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-data-v2orbit-20k
[!NOTE]
For more information on the ORBIT dataset, go check out the preprint available at arxiv.org/abs/2604.01195.
ORBIT: A Synthetic Training Dataset for Search Agents
ORBITis a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Overview
Training data for deep search — tasks requiring multi-step retrieval and reasoning over the web — is scarce.… See the full description on the dataset page: https://huggingface.co/datasets/orbit-ai/orbit-20k.orbital-chaos-nasa-ssc
Orbital Chaos — NASA SSC Spacecraft Position Dataset
4.8 million spacecraft position records paired with solar wind measurements, covering 2023–2025 at 1-minute resolution. Built to support machine learning research on orbital prediction under varying space weather conditions.
Dataset Contents
Spacecraft
Orbit Type
Records
Purpose
ISS
LEO ~408 km
1.58M
Primary prediction target
DSCOVR
L1 Lagrange
131K
Solar wind leading indicator
MMS-1
Highly elliptical… See the full description on the dataset page: https://huggingface.co/datasets/datamatters24/orbital-chaos-nasa-ssc.multilingual-data-sampleorbit-stage-3-27k
⚠️ [!Warning]
This is an unverified ORBIT dataset, only the Stage-3, and may contain data inaccuracies. The verified ORBIT dataset to use is orbit-ai/orbit-20k.
ORBIT: A Synthetic Training Dataset for Search Agents
ORBIT is a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Overview
Training data for deep search — tasks requiring multi-step… See the full description on the dataset page: https://huggingface.co/datasets/orbit-ai/orbit-stage-3-27k.orbital-schemas
Orbital Schemas Dataset
Training data for OrbGen - a model that generates valid Orbital schemas (.orb files).
Dataset Structure
train: 142 examples
validation: 16 examples
test: 10 examples
Features
prompt: Natural language description of the desired schema
completion: Valid Orbital schema in JSON format
domain: Application domain (ecommerce, game, productivity, etc.)
complexity: Schema complexity (simple, medium, complex)
source: Source of the example… See the full description on the dataset page: https://huggingface.co/datasets/orbital-ai/orbital-schemas.orbit-stage-1-44k
⚠️ [!Warning]
This is an unverified ORBIT dataset, only the Stage-1, and may contain data inaccuracies. The verified ORBIT dataset to use is orbit-ai/orbit-20k.
ORBIT: A Synthetic Training Dataset for Search Agents
ORBIT is a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Overview
Training data for deep search — tasks requiring multi-step… See the full description on the dataset page: https://huggingface.co/datasets/orbit-ai/orbit-stage-1-44k.solar-orbiter-encounters
Solar Orbiter Encounter Timeline
Credit: NASA/SDO
Part of a dataset collection on Hugging Face.
Dataset description
Complete mission event timeline for the ESA/NASA Solar Orbiter — the first spacecraft designed to deliver sustained close-up views of the Sun's polar regions. Compiled from ESA Solar Orbiter operations documents and the Mueller et al. 2020 mission overview paper (A&A 642, A1).
The dataset covers all perihelion encounters from P1 (2020-06-15… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/solar-orbiter-encounters.orbit-stage-2-27k
⚠️ [!Warning]
This is an unverified ORBIT dataset, only the Stage-2, and may contain data inaccuracies. The verified ORBIT dataset to use is orbit-ai/orbit-20k.
ORBIT: A Synthetic Training Dataset for Search Agents
ORBIT is a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Overview
Training data for deep search — tasks requiring multi-step… See the full description on the dataset page: https://huggingface.co/datasets/orbit-ai/orbit-stage-2-27k.Orbit-200K
🪐 Orbit-200K
A high-signal instruction dataset designed for training universal language models with dense reasoning, coding, mathematics, and general intelligence—without conversational bloat.
Orbit-200K is a carefully curated dataset of 200,000 instruction-response pairs optimized for training modern language models ranging from 0.5B to 7B parameters.
Unlike many public instruction datasets, Orbit-200K removes unnecessary conversational filler and focuses on maximizing… See the full description on the dataset page: https://huggingface.co/datasets/Pluto-AI-Labs/Orbit-200K.orbital-mining-corpus
Orbital Mining Corporation (OMC) Corpus
A synthetic domain-adaptation corpus for continued pre-training (mid-training) of language models on technical aerospace and deep-space mining documentation.
Dataset Summary
This corpus contains 1,000 long-form technical documents (~10M tokens) representing the internal document ecosystem of Orbital Mining Corporation (OMC) — a fictional company operating crewed spacecraft and extracting resources from the asteroid belt. All… See the full description on the dataset page: https://huggingface.co/datasets/atenareply/orbital-mining-corpus.moon-lunar-orbiter Lunar Orbiter Data 1966-1967
This dataset contains images taken by the lunar orbiter during its 5 missions in the years 1966-1967. The purpose of the first three missions was to obtain images for use as landing sites for the Apollo mission. The fourth, and fifth missions captured the near and far side of the moon respectively. The starting number of each frame represents the mission so for example 1005 means 5 frames taken in the first mission. Note that there were two cameras on board the… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/moon-lunar-orbiter.OrbitalRisk-2023-2024-Satellite-Cleaned
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/akashkm9996/OrbitalRisk-2023-2024-Satellite-Cleaned.orbital-fragmentation-events
Orbital Fragmentation Events
Part of the Orbital Mechanics Datasets collection on Hugging Face.
Catalog of 1,073 orbital fragmentation events derived from the NORAD Satellite Catalog
(SATCAT) via CelesTrak. A fragmentation event is identified as any launch
that produced 4 or more cataloged debris objects, indicating an in-orbit breakup caused
by explosions, collisions, anomalous events, or deliberate destruction.
Dataset description
Every significant breakup event in… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/orbital-fragmentation-events.benchy-left-to-right_20260624_155604This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/rick-orbital/benchy-left-to-right_20260624_155604.orbital-mining-corpus-seeded
Orbital Mining Corpus — Seeded
What's inside — The 848 Orbital Mining Corporation technical documents plus 16 new synthetic OMC docs (incident reports, ops logs, handovers, test reports, maintenance procedures).
Where it comes from — orbital-mining-corpus, extended by CROSS-SEEDING: teacher-written docs (Scaleway, gpt-oss-120b) anchored to real telemetry values from the Mars facts table.
How it was built / modified — Strict grounding — only candidates whose telemetric values… See the full description on the dataset page: https://huggingface.co/datasets/atenareply/orbital-mining-corpus-seeded.KoAlpca_v1.22_1000_orbit_DPO_intelOrbitaOrbita-deepclean-sharegpt--- Cleaning Summary ---
Dataset : NewstaR/Orbita
Input Format : single_column_sharegpt
Single column (Input) : conversations
Roles (Input) : from=(human/gpt)
Output Format : sharegpt
ShareGPT column (Output) : conversations
Roles (Output) : from=(human/gpt)
----------------------------------------
Initial size : 132426
After input parsing : 132426
After basic cleaning… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/Orbita-deepclean-sharegpt.orbitals-ngyogastudio-orbit-timeofdayMaterials_Low_Spin_OrbitQ1-OrbitalForge-datasetorbital-schemas
Orbital Schemas Dataset
Training data for OrbGen - a model that generates valid Orbital schemas (.orb files).
Dataset Structure
train: 142 examples
validation: 16 examples
test: 10 examples
Features
prompt: Natural language description of the desired schema
completion: Valid Orbital schema in JSON format
domain: Application domain (ecommerce, game, productivity, etc.)
complexity: Schema complexity (simple, medium, complex)
source: Source of the example… See the full description on the dataset page: https://huggingface.co/datasets/javasop/orbital-schemas.Q1-OrbitalForge-datasetorbit-warsQ1-OrbitalForge-dataset
