datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hypersimhypersim-frustum-completion
Hypersim Frustum Point Completion
A large-scale indoor point-cloud completion dataset derived from Hypersim. Each record simulates a partially observed room: one real camera view provides context, while a synthetic nearby “missing camera” frustum hides part of the scene. Models are trained to infer the masked region from visible points.
Why this dataset exists. Real 3D capture — whether from depth sensors, multi-view reconstruction, or neural fields — routinely produces… See the full description on the dataset page: https://huggingface.co/datasets/izhleba/hypersim-frustum-completion.deval_helm_hyperturing1MARS-Hyperspectral-EnMAP-PRISMA-v2025
MARS-Hyperspectral dataset (EnMAP and PRISMA) - version v2025
Updated version of the dataset. More information will be added soon.
hyperspectral-orchard
Living Optics Orchard Dataset
Overview
This dataset contains 435 images of captured in one of the UK's largest orchards, using the Living Optics Camera.
The data consists of RGB images, sparse spectral samples and instance segmentation masks.
The dataset is derived from 44 unique raw files corresponding to 435 frames.
Therefore, multiple frames could originate from the same raw file.
This structure emphasized the need for a split strategy that avoided data leakage.
To… See the full description on the dataset page: https://huggingface.co/datasets/LivingOptics/hyperspectral-orchard.Hypersim-ProcessedLeanTransitionCorpus
LeanTransitionCorpus
LeanTransitionCorpus is a dataset for training and studying automated theorem
proving systems in Lean. Its unit of data is one tactic transition: the proof state
before a tactic, the tactic that was executed, and the resulting state. This makes
it suitable for tactic prediction, proof-state representation learning, premise
selection, retrieval, verification, and trajectory-level training.
Many Lean datasets expose a theorem, tactic, and pretty-printed goal… See the full description on the dataset page: https://huggingface.co/datasets/HyperCactus0/LeanTransitionCorpus.MARS-Hyperspectral
MARS-Hyperspectral dataset
This repository contains the data for the work publicly presented at:
[ArXiv preprint] Růžička, Mateo-García, Irakulis-Loitxate et al., Operational machine learning for remote spectroscopic detection of CH4 point sources, arXiv preprint arXiv:2511.07719 (2025).
[Poster] Růžička, Mateo-García, Irakulis-Loitxate et al., Machine Learning Models for Multi-sensor Detection of Methane Leaks in Hyperspectral Data, June. 23-27, 2025, ESA Living Planet Symposium… See the full description on the dataset page: https://huggingface.co/datasets/UNEP-IMEO/MARS-Hyperspectral.hyper-bot-datahyperliquid-node-fills-by-blockHypersim-fullhypernet-prior-topic03
Hypernet — Prior Topic 03 Archive
Complete archive of the prior topic 03 research thread: per-shape SIREN
decoders, per-layer hypernetwork architectures, mapper experiments, and
ancillary checkpoints. This work predates the current image-to-3D pipeline
documented in hypernet-image-to-3d
and the main dataset.
100 shapes, naming convention obj_NN for NN in [0..99].
Contents
Path
Size
Description
watertight/
~5.6 GB
100 watertight .obj meshes (filename… See the full description on the dataset page: https://huggingface.co/datasets/bobthebuilderinternational/hypernet-prior-topic03.function-calling-sharegptThis is a dataset for finetuning models on function calling based on glaiveai/glaive-function-calling-v2.
The dataset includes 86,864 examples of chats that include function calling as part of the conversation. The system prompt includes either 0, 1, or 2 functions that the assistant can use, and instructions on how the agent can use it.
Changes include:
Using ShareGPT format for chats
Adding "function_response" as a role
Removing code examples
Removing examples with invalid JSON as function… See the full description on the dataset page: https://huggingface.co/datasets/hypervariance/function-calling-sharegpt.hyperdrivehypersimriddles_v1
Riddle Processing with GPT-4
Buy me Ko-fi
Credits
All credit for the original riddles goes to crawsome's GitHub repository.
Project Overview
This project involves processing each riddle using GPT-4. The correct answers were provided to the model to generate a desirable output focused on reasoning and logical breakdown.
riddles.json (riddles_1) — 386 samples, sourced from crawsome's GitHub repository.
riddles_2.json — 83 samples, sourced from various Google… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/riddles_v1.hyperliquidL2Book-v2hypersim-examples3dvlm-hypersim
3DVLM Hypersim — §6 depth format
Hypersim (Apple, ICCV 2021) reconverted to the 3DVLM project's unified §6
on-disk format. 457 scenes, ~77,400 frames, V-Ray ground truth.
License: CC BY-SA 3.0, same as the source. If you use this, please cite
Roberts et al. (Hypersim) and apply share-alike terms to any derivatives.
Layout
One uncompressed .tar per scene at the repo root:
hypersim/
ai_001_001.tar
ai_001_002.tar
...
ai_055_010.tar # 457 tars, ~620 MB… See the full description on the dataset page: https://huggingface.co/datasets/helioom/3dvlm-hypersim.hyperliquid-fills-rawroodhypersim-episodes-v3-parquet
hypersim-episodes-v3-parquet
Per-frame Parquet dataset for ReCAST tracker training.
Schema
One row per frame, grouped by episode_id. Arrow memory-mapped access
enables reading specific frames without loading entire episodes.
Column
Type
Description
episode_id
int32
Episode identifier
frame_idx
int32
Frame index within episode
jpeg
binary
JPEG-encoded RGB frame
depth
list<float32>
Flat H×W depth map
seg
list<uint16>
Semantic segmentation (empty if… See the full description on the dataset page: https://huggingface.co/datasets/OSResight/hypersim-episodes-v3-parquet.Processed_HypersimMSI_productshypersim_relative_depth
Hypersim Relative Depth
This repository contains a repackaged and preprocessed version of the
Hypersim Dataset, prepared for relative
depth training with the Marigold V2 codebase.
The repository contains RGB/depth pairs and split metadata organized into
downloadable tar archives. The preprocessing and archive layout are intended
for use with the Marigold V2 data-loading pipeline.
Source dataset
The data originates from:
Dataset: Hypersim
Original repository:… See the full description on the dataset page: https://huggingface.co/datasets/obukhovai/hypersim_relative_depth.hyperliquid-node-tradeshyperpartisan_news_detectionHyperpartisan News Detection was a dataset created for PAN @ SemEval 2019 Task 4.
Given a news article text, decide whether it follows a hyperpartisan argumentation, i.e., whether it exhibits blind, prejudiced, or unreasoning allegiance to one party, faction, cause, or person.
There are 2 parts:
- byarticle: Labeled through crowdsourcing on an article basis. The data contains only articles for which a consensus among the crowdsourcing workers existed.
- bypublisher: Labeled by the overall bias of the publisher as provided by BuzzFeed journalists or MediaBiasFactCheck.com.hypersonic-scramjets
Scramjet Hypersonic Flow Physics Emulator Dataset
A dataset of 6,877 steady-state hypersonic flow simulations of parametrically
varied scramjet geometries, produced with JAX-Fluids (https://github.com/tumaer/JAXFLUIDS) as part of a fully GPU-based CFD workflow intended for
training and evaluating physics emulators.
The current version of the dataset contains the irregular-grid data used for training the AB-UPT emulator (https://arxiv.org/abs/2502.09692) from the paper.
We will… See the full description on the dataset page: https://huggingface.co/datasets/paischer101/hypersonic-scramjets.Hyperspectral_Image_Datasetlayout_diffusion_hypersimThis repository contains the data for SceneCraft: Layout-Guided 3D Scene Generation.
Project page: https://orangesodahub.github.io/SceneCraft
Code: https://github.com/OrangeSodahub/SceneCraft
