datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
huggingface-spaces-codes
📊 Dataset Description
This dataset comprises code files of Huggingface Spaces that have more than 0 likes as of November 10, 2023. This dataset contains various programming languages totaling in 672 MB of compressed and 2.05 GB of uncompressed data.
📝 Data Fields
Field
Type
Description
repository
string
Huggingface Spaces repository names.
sdk
string
Software Development Kit of the space.
license
string
License type of the space.… See the full description on the dataset page: https://huggingface.co/datasets/Weyaxi/huggingface-spaces-codes.space-track-tle-history
Space-Track TLE History
Complete archive of Two-Line Element (TLE) orbital data for every tracked object in Earth orbit, from 1959 to 2026. Sourced from Space-Track.org bulk exports.
Quick Start
from datasets import load_dataset
# Load a specific year
ds = load_dataset("juliensimon/space-track-tle-history", data_files="data/tle_2024.parquet")
# Load everything (238M rows — use streaming for large-scale analysis)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/space-track-tle-history.spacetravlr
SpaceTravLR dataset hub
Precomputed SpaceTravLR outputs: per-gene beta matrices (*_betadata.feather), run metadata, and optional per-sample .h5ad exports.
Layout
spacetravlr/
├── tonsil/ # placeholder / demo gene outputs
└── xenium_skin_mixed/
├── run.toml # shared training config for this cohort
├── manifest.json # sample index and upload metadata
├── sample12/
├── sample13/
├──… See the full description on the dataset page: https://huggingface.co/datasets/Koushul/spacetravlr.us-names-by-state
US Baby names
The SSA dataset with baby names:
https://www.ssa.gov/OACT/babynames/
Coniferest
We use this dataset in the active anomaly discovery Python package coniferest:
https://coniferest.snad.space/en/latest/notebooks/us-names.html
Update the data
Install Python packages: pip install requests aiohttp universal_pathlib pandas
Optionally: download https://www.ssa.gov/OACT/babynames/state/namesbystate.zip
./run.py PATH_OR_URL_TO_namesbystate.zip, path may be… See the full description on the dataset page: https://huggingface.co/datasets/snad-space/us-names-by-state.fsrs-dataseticl-dataset-joint-space
icl-dataset-fixed-action
Derived from adityx23/icl-dataset
(lerobot v2.1 format). Every existing column, task, episode flag
(success/valid/keep), and episode_uid is carried through unchanged.
What's added
One new feature, action.q_target (float32, shape [14], names
lj0..lj6, rj0..rj6): the joint-space reconstruction of each frame's
action.left_ee / action.right_ee cartesian targets, via the mink-based
IK procedure documented in vr_teleop_ik.md (available in the… See the full description on the dataset page: https://huggingface.co/datasets/Hannibal52Barca/icl-dataset-joint-space.spacex-launches
SpaceX Launch History
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Complete record of every SpaceX launch from spacex.com, including mission descriptions, pre/post-launch timelines, and photo galleries. Covers Falcon 1, Falcon 9, Falcon Heavy, and Starship missions.
The data is sourced from the official SpaceX content API and organized into three tables that can be joined on the slug field: launches (one row per mission… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/spacex-launches.vqasynth_spacellava
VQASynth_spacellava
Uses the VQASynth pipeline to synthesize spatialVQA samples, mixed with general VQA samples used to fine-tune LLaVA-v1.5-13b.
icl-dataset-end-effector-space
icl-dataset-fixed-obs
Derived from adityx23/icl-dataset
(lerobot v2.1 format). Every existing column, task, episode flag
(success/valid/keep), and episode_uid is carried through unchanged.
What's added
Two new features, observation.left_ee / observation.right_ee (float32,
shape [7], names qw, qx, qy, qz, x, y, z): the Cartesian end-effector
pose of each arm, forward-kinematics'd from that frame's recorded
observation.state (real joint encoders) through the same… See the full description on the dataset page: https://huggingface.co/datasets/Hannibal52Barca/icl-dataset-end-effector-space.LogiQA2.0
LogiQA2.0
Logiqa2.0 dataset - logical reasoning in MRC and NLI tasks
This is the official repository for the LogiQA 2.0 datasets in our paper LogiQA2.0 - An Improved Dataset for Logic Reasoning in Question Answering and Textual Inference
How to cite
@ARTICLE{10174688,
author={Liu, Hanmeng and Liu, Jian and Cui, Leyang and Teng, Zhiyang and Duan, Nan and Zhou, Ming and Zhang, Yue},
journal={IEEE/ACM Transactions on Audio, Speech, and Language Processing}… See the full description on the dataset page: https://huggingface.co/datasets/spacekat99/LogiQA2.0.SpaceSense-Bench
SpaceSense-Bench: Multi-Modal Spacecraft Perception and Pose Estimation Dataset
Project Page | Paper | Toolkit & Code
SpaceSense-Bench is a high-fidelity simulation-based multi-modal (RGB, Depth, LiDAR Point Cloud) dataset for spacecraft component-level semantic understanding, containing 136 satellite models with synchronized multi-modal data.
Update (2026-05-19). Following issue #5, the pose_ground_truth.csv for all 136 spacecraft has been regenerated to fix a frame-timing… See the full description on the dataset page: https://huggingface.co/datasets/Alvin16/SpaceSense-Bench.SSR-3DFRONT
SSR-3DFRONT: Structured Scene Representation for 3D Indoor Scenes
This dataset provides a processed version of the 3D-FRONT dataset with structured scene representations for text-driven 3D indoor scene synthesis and editing.
Mor information about ReSpace: http://respace.mnbucher.com
For detailed usage instructions, training details, and examples, see the associated repository: https://github.com/GradientSpaces/respace
Our model weights for SG-LLM:… See the full description on the dataset page: https://huggingface.co/datasets/gradient-spaces/SSR-3DFRONT.demo-end-effector-space
demo_action_space
Derived from adityx23/icl-demo-dataset
(lerobot v2.1 format, 285 episodes / 254,171 frames / 27 tasks). Every
existing column, task, episode flag (success/valid/keep), and
episode_uid is carried through unchanged.
Sibling dataset: demo_joint_space
adds the same episodes' joint-space IK targets instead of Cartesian poses.
Same source, same episode indices, same pipeline.
What's added
Two new features, observation.left_ee / observation.right_ee… See the full description on the dataset page: https://huggingface.co/datasets/Hannibal52Barca/demo-end-effector-space.aoi-gt-sft-command
Aoi-Gt-Sft-Command Dataset
Supervised Fine-Tuning dataset with command sequences for AIOps automation.
Dataset Statistics
Total files: 49
File format: JSON
File Structure
Root directory contains all JSON files
Files
assign_to_non_existent_node_social_net-detection-1.json
assign_to_non_existent_node_social_net-localization-1.json
assign_to_non_existent_node_social_net-mitigation-1.json
astronomy_shop_ad_service_manual_gc-detection-1.json… See the full description on the dataset page: https://huggingface.co/datasets/spacezenmasterr/aoi-gt-sft-command.space-track-satcat
NORAD Satellite Catalog (SATCAT)
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Complete NORAD Satellite Catalog from CelesTrak, tracking every object cataloged by the 18th Space Defense Squadron since 1957. Includes active satellites, defunct spacecraft, rocket bodies, and debris.
The SATCAT (Satellite Catalog) is the authoritative registry of all artificial objects in Earth orbit and beyond. Each entry includes launch… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/space-track-satcat.OR-Space
OR-Space
A full-lifecycle workspace benchmark for industrial optimization agents.
OR-Space evaluates whether language-model agents can work reliably with
operations research problems represented as executable, multi-file workspaces.
Rather than presenting a self-contained mathematical prompt, each task
distributes evidence across business requirements, structured data, source
code, execution logs, and solver records.
The benchmark contains 100 optimization topologies. Each… See the full description on the dataset page: https://huggingface.co/datasets/Chenyu-Zhou/OR-Space.ztf-m-dwarf-flares-2025
Hour-Scale ZTF Light Curves for M-Dwarf Flare Search
The dataset is built from short, high-cadence chunks of ZTF DR17 light curves and simulated stellar flares.
The training dataset is composed of both real (class 0, non-flares) and simulated (class 1, flares) data.
The test dataset contains a small number of flares (class 1) found in ZTF by arXiv:2404.07812.
The dataset is described and used by arXiv:2510.24655.
The original photometric data is from the Zwicky Transient Facility.
donki-space-weather-events
NASA DONKI Space Weather Events
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Space weather events from NASA's DONKI (Database Of Notifications, Knowledge, Information) at the Community Coordinated Modeling Center. Covers coronal mass ejections, geomagnetic storms, interplanetary shocks, high-speed streams, and solar energetic particles from 2010 to present.
DONKI tracks the chain of space weather events from Sun to Earth.… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/donki-space-weather-events.space-track-tle-history
Space-Track TLE History
Complete archive of Two-Line Element (TLE) orbital data for every tracked object in Earth orbit, from 1959 to 2026. Sourced from Space-Track.org bulk exports.
Quick Start
from datasets import load_dataset
# Load a specific year
ds = load_dataset("juliensimon/space-track-tle-history", data_files="data/tle_2024.parquet")
# Load everything (238M rows — use streaming for large-scale analysis)ds =… See the full description on the dataset page: https://huggingface.co/datasets/oxzoid/space-track-tle-history.QariOCR-v0.3-markdown-mixed-dataset
QARI Markdown Mixed Dataset
📋 Dataset Summary
The QARI v0.3 Markdown Mixed Dataset is a specialized synthetic dataset designed for training Arabic OCR models with a focus on complex document layouts and HTML structure understanding.
This dataset is part of the QARI-OCR project, which achieves state-of-the-art performance in Arabic text recognition.
This dataset contains 37,000 synthetically generated Arabic document images (29.6k train, 3.7k… See the full description on the dataset page: https://huggingface.co/datasets/NAMAA-Space/QariOCR-v0.3-markdown-mixed-dataset.deep-space-missions-tracker
Deep-Space Missions Tracker
Credit: NASA/JPL-Caltech
Part of a dataset collection on Hugging Face.
Dataset description
A daily-updating log of where humanity's active deep-space missions are right now, computed from NASA/JPL's Horizons ephemeris system. Updated daily, growing one row per mission per day.
The dataset tracks a fleet of interplanetary spacecraft -- the Voyagers in interstellar space, New Horizons in the Kuiper Belt, Juno at Jupiter, the… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/deep-space-missions-tracker.space-weather-indices
Space Weather Indices (Kp, Ap, F10.7)
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Daily geomagnetic and solar activity indices since 1957 from NOAA SWPC via CelesTrak. Includes Kp/Ap geomagnetic indices, F10.7 solar radio flux, and international sunspot numbers.
These indices together form the essential parameter set for characterizing the state of the heliosphere and its coupling to the terrestrial environment. The Kp… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/space-weather-indices.repo-to-space-example-videos
Gradio Space Example Inputs — Videos
A small, curated, freely-licensed pool of videos used as gr.Examples for
Gradio Spaces that wrap video-input generation models (image-to-video,
video-to-video, motion controls, etc.). Sister dataset for images:
linoyts/repo-to-space-example-inputs.
When a Space takes video input, the agent building the Space picks 2–3 clips
whose caption + categories match the model's task, downloads them via
hf_hub_download, runs any model-specific… See the full description on the dataset page: https://huggingface.co/datasets/linoyts/repo-to-space-example-videos.spaces-of-the-week-legacywikipedia-trivia-query-variationspaces-of-the-week
Spaces of the Week
Weekly list of "Spaces of the Week" featured at
https://huggingface.co/spaces, unified across five years of observations.
The old per-week YYYY/YYYY-MM-DD.csv layout of this repo has been moved to
hysts-bot-data/spaces-of-the-week-legacy.
All active use should read data.parquet here.
Schema
Column
Type
Notes
week_iso
string
ISO week, e.g. 2025-W13
week_start_date
date
Monday of the ISO week
week_label
string?
Date pill shown on… See the full description on the dataset page: https://huggingface.co/datasets/hysts-bot-data/spaces-of-the-week.celestrak-space-weather
CelesTrak Consolidated Space Weather
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
CelesTrak consolidated space weather data -- THE file every orbit propagator needs. Daily Kp, Ap, F10.7, and solar/geomagnetic indices used by SGP4/SDP4 propagators, atmospheric models (JB2008, NRLMSISE), and conjunction screening.
The CelesTrak space weather file, maintained by Dr. T.S. Kelso, is the de facto standard input file for… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/celestrak-space-weather.FineWeb2-MSA
FineWeb2 MSA Arabic
This is the MSA Arabic Portion of The FineWeb2 Dataset.
This dataset contains a rich collection of text in MSA Arabic (ISO 639-3: arz), a widely spoken dialect within the Afro-Asiatic language family.
With over 439 million words and 1.4 million documents, it serves as a valuable resource for NLP development and linguistic research focused on Egyptian Arabic.
Purpose of This Repository
This repository provides easy access to the Arabic portion - MSA… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/FineWeb2-MSA.heb-words27-spaced
heb-words27-diffpen
Word-level Hebrew handwriting for DiffusionPen finetuning -- the step up from
cyttic/heb-connected-bigrams17-diffpen, which covered 2-character units. Here the unit
is a whole word.
Content
20,000 most frequent words of LLMGen2 (sentences_llm2.txt, 968,204 LLM-generated
Hebrew sentences, 167,879 unique words)
643 numbers -- LLMGen2 contains no digits at all, because the generation
prompt forbade them, so numerals had to be added separately or… See the full description on the dataset page: https://huggingface.co/datasets/cyttic/heb-words27-spaced.SpaceJudgeDataset
SpaceJudge Dataset
The SpaceJudge Dataset uses prometheus-vision to apply
a rubric assessing the quality of response to spatial VQA inquiries on a 1-5 likert scale by prompting
SpaceLLaVA to perform VLM-as-a-Judge.
The assessment is made for images in the OpenSpaces dataset in order to
distill the 13B VLM judge into smaller models like Florence-2
by introducing a new <JUDGE>task.
Citations
@misc{lee2024prometheusvision,
title={Prometheus-Vision: Vision-Language… See the full description on the dataset page: https://huggingface.co/datasets/remyxai/SpaceJudgeDataset.
