datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.umi-okra-dex1
UMI Okra Grasping Dataset — Lab Subset (Unitree Dex1-1)
Universal Manipulation Interface (UMI) hand-held teleoperation data for an okra-fruit
grasping / harvesting task. Recorded to train a Diffusion Policy (with a comparison
ACT track) deployed on a Unitree G1.
This is the indoor lab subset. Every session here was recorded in a lab mock okra
field (artificial foliage, white-walled room). The outdoor sessions from the original
collection are not included — see Scope and… See the full description on the dataset page: https://huggingface.co/datasets/Orboh/umi-okra-dex1.gso-orbit-rgbaobjaverse_orbit_rendersOrbEvo
OrbEvo
Paper: https://arxiv.org/abs/2603.03511
Currently in this dataset:
5,000 QM9 TDDFT trajectories generated using the ABACUS simulation code (https://arxiv.org/abs/2501.08697).
Folders:
- QM9_tddft/ <--- the data used for training and evaluation
- processed/
- raw/
- MDA_tddft/
- MDA_train <--- the data used for training and validation
- processed/
- raw/
- MDA_test <--- the data used for test
- processed/
- raw/
- before_down_sampling/ <--- the… See the full description on the dataset page: https://huggingface.co/datasets/divelab/OrbEvo.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.multilingual-data-v2MofasaDB
MofasaDB
The MofasaDB is a publicly available dataset containing 200.000+ de novo generated MOF (Metal-Organic Framework) structures from Mofasa trained on QMOF (up to 170 atoms), along with their geometry-relaxed counterparts. The database is released alongside the paper Mofasa: A Step Change in Metal-Organic Framework Generation. A user-friendly web interface for search and discovery can be accessed at https://mofux.ai/.
Database Overview
The database contains… See the full description on the dataset page: https://huggingface.co/datasets/Orbital-Materials/MofasaDB.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our leaderboard at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.Orbit_Planner
Orbit-Planner Orbital Evasion Dataset
Orbit-Planner is a simulated multimodal trajectory dataset for vision-based
spacecraft navigation and obstacle avoidance. It contains synchronized
first-person RGB images, depth maps, spacecraft states, thruster commands, and
event labels collected in the Orbital Evasion task from Space Robotics Bench
and NVIDIA Isaac Sim.
The dataset is intended for learning latent world models, spacecraft dynamics,
visual representations… See the full description on the dataset page: https://huggingface.co/datasets/warriorLZJ/Orbit_Planner.stable-orbitsor-bench-toxic-all
OR-Bench: An Over-Refusal Benchmark for Large Language Models
This dataset constains highly toxic prompts, use with caution!!!
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.phase-orbit-complex-g-closure-v3.5.0
Complete physical complex G-closure and proof-carrying finite-data compilation
Two-dimensional two-phase conductivity — standalone v3.5.0
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiRelease: 3.5.0 — 12 September 2026Status: standalone superseding research release; not independently peer reviewed or proof-assistant formalized.
This repository is a research-first, machine-readable release for experts, reviewers, reproducibility work, search… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/phase-orbit-complex-g-closure-v3.5.0.LFM-Orbit-SatData
LFM Orbit SatData
Retagged Earth-observation training data produced by LFM Orbit for the Liquid AI x DPhi Space Hackathon.
The default viewer config is training_assets.jsonl, which contains single-image SFT rows with image, messages, and metadata. Temporal sequence rows live in the temporal_sft config so the Hugging Face Dataset Viewer does not try to cast sequence rows into the single-image schema.
Configs
Config
File
Purpose
default
training_assets.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Shoozes/LFM-Orbit-SatData.player-detectionIteractBench
Iteract-Bench
Benchmark data for ITERACT-BENCH: Benchmarking Agents for Interactive Data Analysis Workflows.
The code for running the benchmark is at https://anonymous.4open.science/r/Iteract-Bench-CDFD.
Each directory is one interactive data-analysis task. A user simulator discloses analytical requirements during the episode. An independent verifier later checks stage answers against gold outputs that the simulator never sees.
Statistics
Statistic
Value… See the full description on the dataset page: https://huggingface.co/datasets/blue-orbit-472/IteractBench.orbit-20k
[!NOTE]
For more information on the ORBIT dataset, go check out the preprint available at arxiv.org/abs/2604.01195.
ORBIT: A Synthetic Training Dataset for Search Agents
ORBITis a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Overview
Training data for deep search — tasks requiring multi-step retrieval and reasoning over the web — is scarce.… See the full description on the dataset page: https://huggingface.co/datasets/orbit-ai/orbit-20k.Orbis
Orbis
High Diversity · High Fidelity · Controllability · Simulation Ready
📖 Dataset Overview
Orbis is an open-source collection of high-quality 3D scene datasets that covers multiple spatial scales, ranging from tabletop (small-scale) scenes to complex indoor (medium-scale), urban, and natural (large-scale) environments. All scenes are simulation-ready with comprehensive physical properties, enabling seamless integration with physics engines. Powered by… See the full description on the dataset page: https://huggingface.co/datasets/IntimeAI/Orbis.orbit-mt-eval
ORBIT-MT-Eval: Expert Annotations and Prompts for Machine Translation Meta-Evaluation
ORBIT-MT-Eval is an English-to-Russian reference dataset for machine translation meta-evaluation: comparing evaluator outputs against professional expert annotations. Each of the 600 source–translation segments has three independent expert annotations following the RATE protocol. In Beyond Prompts: A Systematic Study of LLM-Based Machine Translation Evaluators, two annotations serve as gold… See the full description on the dataset page: https://huggingface.co/datasets/foksly/orbit-mt-eval.OrbbecThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 15601,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zhaoraning/Orbbec.Orbdivination
Thyme Video RL — pass-rate filtered subset
RL training data for a video agent that first skims 64 uniformly sampled frames,
then calls a temporal-crop tool to pull higher-FPS segments from the source mp4.
Each row therefore carries both:
videos — the 64 overview frames, embedded as JPEG bytes
video_path — the source mp4, which the crop tool reads at rollout time
Provenance
Built from 12,000 multiple-choice questions over 1,479 videos (LVBench,
LongVideoBench… See the full description on the dataset page: https://huggingface.co/datasets/EchoMinkki/Orbdivination.or-base-demo-0-200orbital-chaos-nasa-ssc
Orbital Chaos — NASA SSC Spacecraft Position Dataset
4.8 million spacecraft position records paired with solar wind measurements, covering 2023–2025 at 1-minute resolution. Built to support machine learning research on orbital prediction under varying space weather conditions.
Dataset Contents
Spacecraft
Orbit Type
Records
Purpose
ISS
LEO ~408 km
1.58M
Primary prediction target
DSCOVR
L1 Lagrange
131K
Solar wind leading indicator
MMS-1
Highly elliptical… See the full description on the dataset page: https://huggingface.co/datasets/datamatters24/orbital-chaos-nasa-ssc.orbitalorbital-one
orbital-one
Browser spacecraft simulator and headless flight-model SDK.
Launch from the pad, fly a two-stage ascent to orbit — or import the sim engine directly into your own project.
SnapKitty West / SNAPKITTYWEST — MIT License — 2026
What It Is
A browser 3D spacecraft simulator built on Three.js + Vite, with a zero-dependency TypeScript SDK that exposes all 8 sim modules for headless use in Node, Deno, Bun, or any bundler.
import { createState } from… See the full description on the dataset page: https://huggingface.co/datasets/SNAPKITTYWEST/orbital-one.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/jerogo/or-bench.orbitorbital-one
orbital-one
Browser spacecraft simulator and headless flight-model SDK.
Launch from the pad, fly a two-stage ascent to orbit — or import the sim engine directly into your own project.
SnapKitty West / SNAPKITTYWEST — MIT License — 2026
What It Is
A browser 3D spacecraft simulator built on Three.js + Vite, with a zero-dependency TypeScript SDK that exposes all 8 sim modules for headless use in Node, Deno, Bun, or any bundler.
import { createState } from… See the full description on the dataset page: https://huggingface.co/datasets/Snapkitty/orbital-one.Parameter-Golf-V17-512Cube-XYZ-Priority-Orbital-Compression
Parameter Golf V17 — 512Cube Solution Bank
This is an English research-control dataset for OpenAI Parameter Golf work. It is not a replacement for FineWeb and must not be used as a substitute training or validation corpus. FineWeb remains the canonical data path for contest scoring.
The dataset captures three things:
V17 512Cube routing concepts translated into English.
Contest and submission guardrails for legal, reproducible BPB reduction.
Screenshot-derived scouting observations… See the full description on the dataset page: https://huggingface.co/datasets/8Planetterraforming/Parameter-Golf-V17-512Cube-XYZ-Priority-Orbital-Compression.bbh-orbit-lensing-gifs
Pokémon, behind a binary black hole
before
after
PUT YOUR POKEMON BEHIND BINARY BLACK HOLES
Artwork © Nintendo / Creatures / GAME FREAK. See LICENSE.
