datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
orbit-20k
[!NOTE]
For more information on the ORBIT dataset, go check out the preprint available at arxiv.org/abs/2604.01195.
ORBIT: A Synthetic Training Dataset for Search Agents
ORBITis a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Overview
Training data for deep search — tasks requiring multi-step retrieval and reasoning over the web — is scarce.… See the full description on the dataset page: https://huggingface.co/datasets/orbit-ai/orbit-20k.orbit-stage-3-27k
⚠️ [!Warning]
This is an unverified ORBIT dataset, only the Stage-3, and may contain data inaccuracies. The verified ORBIT dataset to use is orbit-ai/orbit-20k.
ORBIT: A Synthetic Training Dataset for Search Agents
ORBIT is a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Overview
Training data for deep search — tasks requiring multi-step… See the full description on the dataset page: https://huggingface.co/datasets/orbit-ai/orbit-stage-3-27k.orbit-seeds
[!NOTE]
For more information on the ORBIT dataset, go check out the preprint available at arxiv.org/abs/2604.01195.
ORBIT: A Synthetic Training Dataset for Search Agents
ORBITis a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Orbit Seeds
Seed entities collected from English Wikipedia, organised by domain. Each record is a Wikipedia page that… See the full description on the dataset page: https://huggingface.co/datasets/orbit-ai/orbit-seeds.orbit-stage-1-44k
⚠️ [!Warning]
This is an unverified ORBIT dataset, only the Stage-1, and may contain data inaccuracies. The verified ORBIT dataset to use is orbit-ai/orbit-20k.
ORBIT: A Synthetic Training Dataset for Search Agents
ORBIT is a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Overview
Training data for deep search — tasks requiring multi-step… See the full description on the dataset page: https://huggingface.co/datasets/orbit-ai/orbit-stage-1-44k.orbital-mechanics-instruct-32
🛰️ Orbital Mechanics Instruction Dataset
Expert-crafted prompt–completion pairs for fine-tuning LLMs on space mission analysis and design
📋 Dataset Summary
A curated dataset of 32 expert-crafted instruction–completion pairs designed for fine-tuning large language models on orbital mechanics and space mission analysis tasks. Each example contains a natural-language problem statement paired with a structured, step-by-step solution featuring properly formatted… See the full description on the dataset page: https://huggingface.co/datasets/abdohisham12/orbital-mechanics-instruct-32.orbit-stage-2-27k
⚠️ [!Warning]
This is an unverified ORBIT dataset, only the Stage-2, and may contain data inaccuracies. The verified ORBIT dataset to use is orbit-ai/orbit-20k.
ORBIT: A Synthetic Training Dataset for Search Agents
ORBIT is a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Overview
Training data for deep search — tasks requiring multi-step… See the full description on the dataset page: https://huggingface.co/datasets/orbit-ai/orbit-stage-2-27k.
