datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
explore-persona-space-datahuggingface-spaces-codes
📊 Dataset Description
This dataset comprises code files of Huggingface Spaces that have more than 0 likes as of November 10, 2023. This dataset contains various programming languages totaling in 672 MB of compressed and 2.05 GB of uncompressed data.
📝 Data Fields
Field
Type
Description
repository
string
Huggingface Spaces repository names.
sdk
string
Software Development Kit of the space.
license
string
License type of the space.… See the full description on the dataset page: https://huggingface.co/datasets/Weyaxi/huggingface-spaces-codes.PocketQubespacecast-data
Vlasiator Dataset for Machine Learning Studies
The data is stored in Zarr.
It can be downloaded to a local data directory with:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="deinal/spacecast-data",
repo_type="dataset",
local_dir="data"
)
This will yield a local data folder that can be used with spacecast:
data/
├── graph/ - Directory containing graphs for training
├── run_1.zarr/ - Vlasiator run 1 with ρ = 0.5 cm⁻³… See the full description on the dataset page: https://huggingface.co/datasets/deinal/spacecast-data.mid-space
MID-Space: Aligning Diverse Communities’ Needs to Inclusive Public Spaces
A new version of the dataset will be released soon, incorporating user identity markers and expanded annotations.
LIVS PAPER
Click below to see more:
Overview
The MID-Space dataset is designed to align AI-generated visualizations of urban public spaces with the preferences of diverse and marginalized communities in Montreal. It includes textual prompts, Stable Diffusion… See the full description on the dataset page: https://huggingface.co/datasets/mila-ai4h/mid-space.SwissCubeSWiM-SpacecraftWithMasks
SWiM: Spacecraft With Masks
A large-scale instance segmentation dataset of nearly 64k annotated spacecraft images created using real spacecraft models, superimposed on a mixture of real and synthetic backgrounds generated using NASA's TTALOS pipeline. To mimic camera distortions and noise in real-world image acquisition, we added different types of noise and distortion.
Dataset Summary
The dataset contains over 63,917 annotated images with instance masks for varied… See the full description on the dataset page: https://huggingface.co/datasets/RiceD2KLab/SWiM-SpacecraftWithMasks.space-track-tle-history
Space-Track TLE History
Complete archive of Two-Line Element (TLE) orbital data for every tracked object in Earth orbit, from 1959 to 2026. Sourced from Space-Track.org bulk exports.
Quick Start
from datasets import load_dataset
# Load a specific year
ds = load_dataset("juliensimon/space-track-tle-history", data_files="data/tle_2024.parquet")
# Load everything (238M rows — use streaming for large-scale analysis)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/space-track-tle-history.Proxy3D-SpaceSpan-318K
SpaceSpan Dataset
SpaceSpan is a large-scale dataset curated for the training and evaluation of 3D vision-language models (VLMs), specifically introduced in the paper Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment.
Project Page | GitHub Repository
Dataset Description
The SpaceSpan dataset is designed to help VLMs develop spatial intelligence through 3D proxy representations. It incorporates heterogeneous visual… See the full description on the dataset page: https://huggingface.co/datasets/Spacewanderer8263/Proxy3D-SpaceSpan-318K.SpaceR-151k
Citation
@article{ouyang2025spacer,
title={SpaceR: Reinforcing MLLMs in Video Spatial Reasoning},
author={Ouyang, Kun and Liu, Yuanxin and Wu, Haoning and Liu, Yi and Zhou, Hao and Zhou, Jie and Meng, Fandong and Sun, Xu},
journal={arXiv preprint arXiv:2504.01805},
year={2025}
}
License
The usage of SpaceR-151k dataset and SpaceR model weights must strictly follow CC BY-NC 4.0 License.
ztf-dr3-m31-featuresAccidentBench
AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
Website
·
Code
·
Leaderboard
·
Dataset
·
Dataset-Zip
·
Issue
Project Homepage:
https://accident-bench.github.io/
About the Dataset:
This benchmark includes approximately 2,000 videos and 19,000 human-annotated question-answer pairs, covering a wide range of reasoning tasks (as shown in Figure 1). We… See the full description on the dataset page: https://huggingface.co/datasets/Open-Space-Reasoning/AccidentBench.UnorthoDOS
UnorthoDOS - Unorthorectified Dataset for On board Satellite methane detection
This repository contains two ML-ready datasets, comprising orthorectified and unorthorectified hyperspectral images from the Earth Surface Mineral Dust Source Investigation (EMIT) sensor, for methane detection. This dataset supports the research presented in the paper Towards Methane Detection Onboard Satellites.
Paper: Towards Methane Detection Onboard Satellites
Code:… See the full description on the dataset page: https://huggingface.co/datasets/SpaceML/UnorthoDOS.plasticc-gpspacetravlr
SpaceTravLR dataset hub
Precomputed SpaceTravLR outputs: per-gene beta matrices (*_betadata.feather), run metadata, and optional per-sample .h5ad exports.
Layout
spacetravlr/
├── tonsil/ # placeholder / demo gene outputs
└── xenium_skin_mixed/
├── run.toml # shared training config for this cohort
├── manifest.json # sample index and upload metadata
├── sample12/
├── sample13/
├──… See the full description on the dataset page: https://huggingface.co/datasets/Koushul/spacetravlr.us-names-by-state
US Baby names
The SSA dataset with baby names:
https://www.ssa.gov/OACT/babynames/
Coniferest
We use this dataset in the active anomaly discovery Python package coniferest:
https://coniferest.snad.space/en/latest/notebooks/us-names.html
Update the data
Install Python packages: pip install requests aiohttp universal_pathlib pandas
Optionally: download https://www.ssa.gov/OACT/babynames/state/namesbystate.zip
./run.py PATH_OR_URL_TO_namesbystate.zip, path may be… See the full description on the dataset page: https://huggingface.co/datasets/snad-space/us-names-by-state.spacecast-data-small
Small Vlasiator Dataset for Machine Learning Studies
The data is stored in Zarr format.
It can be downloaded to a local data_small directory with:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="deinal/spacecast-data-small",
repo_type="dataset",
local_dir="data_small"
)
This will yield a local data_small folder that can be used with spacecast:
data_small/
├── graph/ - Directory containing graphs for training
├── run.zarr/ -… See the full description on the dataset page: https://huggingface.co/datasets/deinal/spacecast-data-small.fsrs-dataseticl-dataset-joint-space
icl-dataset-fixed-action
Derived from adityx23/icl-dataset
(lerobot v2.1 format). Every existing column, task, episode flag
(success/valid/keep), and episode_uid is carried through unchanged.
What's added
One new feature, action.q_target (float32, shape [14], names
lj0..lj6, rj0..rj6): the joint-space reconstruction of each frame's
action.left_ee / action.right_ee cartesian targets, via the mink-based
IK procedure documented in vr_teleop_ik.md (available in the… See the full description on the dataset page: https://huggingface.co/datasets/Hannibal52Barca/icl-dataset-joint-space.transformers-stats-space-dataspacex-launches
SpaceX Launch History
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Complete record of every SpaceX launch from spacex.com, including mission descriptions, pre/post-launch timelines, and photo galleries. Covers Falcon 1, Falcon 9, Falcon Heavy, and Starship missions.
The data is sourced from the official SpaceX content API and organized into three tables that can be joined on the slug field: launches (one row per mission… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/spacex-launches.vqasynth_spacellava
VQASynth_spacellava
Uses the VQASynth pipeline to synthesize spatialVQA samples, mixed with general VQA samples used to fine-tune LLaVA-v1.5-13b.
spaces-privacy-reportsThis repository is used as remote storage for the privacy reports generated in the Spaces Privacy Report app.
icl-dataset-end-effector-space
icl-dataset-fixed-obs
Derived from adityx23/icl-dataset
(lerobot v2.1 format). Every existing column, task, episode flag
(success/valid/keep), and episode_uid is carried through unchanged.
What's added
Two new features, observation.left_ee / observation.right_ee (float32,
shape [7], names qw, qx, qy, qz, x, y, z): the Cartesian end-effector
pose of each arm, forward-kinematics'd from that frame's recorded
observation.state (real joint encoders) through the same… See the full description on the dataset page: https://huggingface.co/datasets/Hannibal52Barca/icl-dataset-end-effector-space.LogiQA2.0
LogiQA2.0
Logiqa2.0 dataset - logical reasoning in MRC and NLI tasks
This is the official repository for the LogiQA 2.0 datasets in our paper LogiQA2.0 - An Improved Dataset for Logic Reasoning in Question Answering and Textual Inference
How to cite
@ARTICLE{10174688,
author={Liu, Hanmeng and Liu, Jian and Cui, Leyang and Teng, Zhiyang and Duan, Nan and Zhou, Ming and Zhang, Yue},
journal={IEEE/ACM Transactions on Audio, Speech, and Language Processing}… See the full description on the dataset page: https://huggingface.co/datasets/spacekat99/LogiQA2.0.SpaceSense-Bench
SpaceSense-Bench: Multi-Modal Spacecraft Perception and Pose Estimation Dataset
Project Page | Paper | Toolkit & Code
SpaceSense-Bench is a high-fidelity simulation-based multi-modal (RGB, Depth, LiDAR Point Cloud) dataset for spacecraft component-level semantic understanding, containing 136 satellite models with synchronized multi-modal data.
Update (2026-05-19). Following issue #5, the pose_ground_truth.csv for all 136 spacecraft has been regenerated to fix a frame-timing… See the full description on the dataset page: https://huggingface.co/datasets/Alvin16/SpaceSense-Bench.anki-revlogs-10k
Introduction
Anki Revlogs 10K is a dataset of 10k collections from Anki for FSRS project. It is a random sample of collections with 5000+ revlog entries, so it should contain a mix of older (still active) users, and newer users.
The dataset contains three parts: revlogs, cards, and decks.
Revlogs
This dataset contains flashcard review records with the following fields:
card_id: Unique identifier for each flashcard
day_offset: Number of days since the start of… See the full description on the dataset page: https://huggingface.co/datasets/open-spaced-repetition/anki-revlogs-10k.spacev1b
SPACEV1B: A billion-Scale vector dataset for text descriptors
This is a dataset released by Microsoft from SpaceV, Bing web vector search scenario, for large scale vector search related research usage. It consists of more than one billion document vectors
and 29K+ query vectors encoded by Microsoft SpaceV Superior model. This model is trained to capture generic intent representation for both documents and queries.
The goal is to match the query vector to the closest document… See the full description on the dataset page: https://huggingface.co/datasets/jkhe/spacev1b.SSR-3DFRONT
SSR-3DFRONT: Structured Scene Representation for 3D Indoor Scenes
This dataset provides a processed version of the 3D-FRONT dataset with structured scene representations for text-driven 3D indoor scene synthesis and editing.
Mor information about ReSpace: http://respace.mnbucher.com
For detailed usage instructions, training details, and examples, see the associated repository: https://github.com/GradientSpaces/respace
Our model weights for SG-LLM:… See the full description on the dataset page: https://huggingface.co/datasets/gradient-spaces/SSR-3DFRONT.SpaceVista-Full
SpaceVista: All-Scale Visual Spatial Reasoning from $mm$ to $km$
🤗 Hugging Face | 📑 Paper | ⚙️ Github | 🖥️ Home Page
Peiwen Sun*, Shiqiang Lang*, Dongming Wu, Yi Ding, Kaituo Feng, Huadai Liu, Zhen Ye, Rui Liu, Yun-Hui Liu, Jianan Wang, Xiangyu Yue
The SFT training data for SpaceVista: All-Scale Visual Spatial Reasoning from $mm$ to $km$.
Data Preview
Tiny Tabletop
Tabletop
Indoor
Wild Indoor
Outdoor… See the full description on the dataset page: https://huggingface.co/datasets/SpaceVista/SpaceVista-Full.
