datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hyper-bot-datahyperliquid-node-fills-by-blockhypersim-episodes-v3-parquet
hypersim-episodes-v3-parquet
Per-frame Parquet dataset for ReCAST tracker training.
Schema
One row per frame, grouped by episode_id. Arrow memory-mapped access
enables reading specific frames without loading entire episodes.
Column
Type
Description
episode_id
int32
Episode identifier
frame_idx
int32
Frame index within episode
jpeg
binary
JPEG-encoded RGB frame
depth
list<float32>
Flat H×W depth map
seg
list<uint16>
Semantic segmentation (empty if… See the full description on the dataset page: https://huggingface.co/datasets/OSResight/hypersim-episodes-v3-parquet.hyperliquid-node-tradeslayout_diffusion_hypersimThis repository contains the data for SceneCraft: Layout-Guided 3D Scene Generation.
Project page: https://orangesodahub.github.io/SceneCraft
Code: https://github.com/OrangeSodahub/SceneCraft
hyperpartisan_news_detection_bypublisher_promptsourcehyperpartisan_news_detection
Dataset Card for "hyperpartisannewsdetection"
More Information needed
Competition-Submissions
Competition Submissions
A curated dataset of writing that models compassionate moral reasoning about nonhuman sentient beings — animals, insects, digital minds, and entities whose moral status is uncertain.
Designed for pretraining and fine-tuning language models to reason more carefully and compassionately when facing decisions that affect sentient life.
Why This Dataset Exists
Recent alignment research shows that training on synthetic documents depicting… See the full description on the dataset page: https://huggingface.co/datasets/Hyperstition-for-Good/Competition-Submissions.hyper_drive
Towards automated analysis of large environments, hyperspectral sensors must be adapted into a format where they can be operated from mobile robots. In this dataset, we highlight hyperspectral datacubes collected from the Hyper-Drive imaging system. Our system collects and registers datacubes spanning the visible to shortwave infrared (660-1700 nm) in 33 wavelength channels. The system also simultaneously captures the ambient solar spectrum reflected off a white reference tile. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nhanson2/hyper_drive.OS-Atlas_ScreenSpottom_cleanHyperThink-Max-200K
🔮 HyperThink
HyperThink is a premium, best-in-class dataset series capturing deep reasoning interactions between users and an advanced Reasoning AI system. Designed for training and evaluating next-gen language models on complex multi-step tasks, the dataset spans a wide range of prompts and guided thinking outputs.
🚀 Dataset Tiers
HyperThink is available in three expertly curated versions, allowing flexible scaling based on compute resources and training goals:… See the full description on the dataset page: https://huggingface.co/datasets/Sashvat/HyperThink-Max-200K.HyperBrowseComp
HyperBrowseComp
Hard, multi-hop web search questions across 13 languages. Encrypted (AES-256-GCM).
from datasets import load_dataset
ds = load_dataset("afaji/HyperBrowseComp", split="test")
id is <LANG>_<question id>, e.g. KO_7388.
Decrypt
import base64
import hashlib
from cryptography.hazmat.primitives.ciphers.aead import AESGCM
def _nonce(canary, field):
return hashlib.sha256(f"{canary}:{field}".encode()).digest()[:12]
def _open(ciphertext_b64, key… See the full description on the dataset page: https://huggingface.co/datasets/afaji/HyperBrowseComp.HyperThink-X-Nvidia-Opencode-Reasoning-200K
🔮 HyperThink
HyperThink is a premium, best-in-class dataset series capturing deep reasoning interactions between users and an advanced Reasoning AI system. Designed for training and evaluating next-gen language models on complex multi-step tasks, the dataset spans a wide range of prompts and guided thinking outputs.
🚀 Dataset Tiers
HyperThink is available in three expertly curated versions, allowing flexible scaling based on compute resources and training goals:… See the full description on the dataset page: https://huggingface.co/datasets/Sashvat/HyperThink-X-Nvidia-Opencode-Reasoning-200K.hyperpartisan-cleanedamazon-berkeley-objects
Amazon Berkeley Objects
This is a Hugging Face metadata mirror of the Amazon Berkeley Objects dataset
for reproducible research and HyperView demos. The original dataset is provided
by Amazon.com and UC Berkeley.
This mirror stores metadata tables and official S3 asset URLs. It does not
duplicate catalog images, turntable images, or 3D models as binary files.
Load
from datasets import load_dataset
listings = load_dataset("hyper3labs/amazon-berkeley-objects"… See the full description on the dataset page: https://huggingface.co/datasets/hyper3labs/amazon-berkeley-objects.Hyperphantasia
A Benchmark for Evaluating the
Mental Visualization Capabilities of Multimodal LLMs
Mohammad Shahab Sepehri
Berk Tinaz
Zalan Fabian
Mahdi Soltanolkotabi
Github Repository
Hyperphantasia is a synthetic Visual Question Answering (VQA) benchmark dataset that probes the mental visualization capabilities of Multimodal Large Language Models (MLLMs) from a vision perspective. We reveal that state-of-the-art models struggle with simple tasks that require visual… See the full description on the dataset page: https://huggingface.co/datasets/shahab7899/Hyperphantasia.summarized-hyperpartisan-news-by-facebook-bart-large-cnn-v12D_Multiscale_Hyperelasticity
owner: Safran
license: cc-by-sa-4.0
data_production:
type: simulation
physics: 2D quasistatic non-linear structural mechanics, finite elasticity (large
strains), P1 elements, compressible hyperelastic material
simulator: fenics
num_samples:
train: 764
test: 376
storage_backend: hf_datasets
This dataset was generated with plaid, we refer to this documentation for additional details on how to extract data from plaid_sample objects.
The simplest way to use this dataset is to… See the full description on the dataset page: https://huggingface.co/datasets/PhysArena/2D_Multiscale_Hyperelasticity.2D_Multiscale_Hyperelasticity
Dataset Card
This dataset contains a single huggingface split, named 'all_samples'.
The samples contains a single huggingface feature, named "sample".
Samples are instances of plaid.containers.sample.Sample.
Mesh objects included in samples follow the CGNS standard, and can be converted in
Muscat.Containers.Mesh.Mesh.
Example of commands:
from datasets import load_dataset
from plaid.bridges.huggingface_bridge import huggingface_dataset_to_plaid
hf_dataset =… See the full description on the dataset page: https://huggingface.co/datasets/PLAID-datasets/2D_Multiscale_Hyperelasticity.hyperliquid-replica-cmdsdetails_bunnycore__HyperLlama-3.1-8B
Dataset Card for Evaluation run of bunnycore/HyperLlama-3.1-8B
Dataset automatically created during the evaluation run of model bunnycore/HyperLlama-3.1-8B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_bunnycore__HyperLlama-3.1-8B.sanskrit-karaka-hypergraph
Sanskrit Kāraka Hypergraph
A predication hypergraph over Sanskrit: vertices are lemma types, hyperedges are
predications, and each tine carries a Pāṇinian kāraka role.
Why a hypergraph rather than a graph of binary relations: a sentence is an n-ary
predicate, and an n-ary relation does not survive projection onto its binary
sub-relations. Given only the pairs agent–object, object–recipient and
agent–recipient you can no longer tell whether there was one three-place act or
three… See the full description on the dataset page: https://huggingface.co/datasets/Anamavajra-Labs/sanskrit-karaka-hypergraph.tomtask1582_bless_hypernym_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1582_bless_hypernym_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1582_bless_hypernym_generation.hyperion-v2.0
Hyperion v2.0
Introduction
Hyperion is a comprehensive question answering and conversational dataset designed to promote advancements in AI research with a particular emphasis on reasoning and understanding in scientific domains such as science, medicine, mathematics, and computer science. It integrates data from a wide array of datasets, facilitating the development of models capable of handling complex inquiries and instructions.
Dataset Description
Hyperion… See the full description on the dataset page: https://huggingface.co/datasets/Locutusque/hyperion-v2.0.HyperThink-Midi-100K
🔮 HyperThink
HyperThink is a premium, best-in-class dataset series capturing deep reasoning interactions between users and an advanced Reasoning AI system. Designed for training and evaluating next-gen language models on complex multi-step tasks, the dataset spans a wide range of prompts and guided thinking outputs.
🚀 Dataset Tiers
HyperThink is available in three expertly curated versions, allowing flexible scaling based on compute resources and training goals:… See the full description on the dataset page: https://huggingface.co/datasets/Sashvat/HyperThink-Midi-100K.hyperfamila_provetteThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 400,
"total_frames": 90620,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:400"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/giacomoran/hyperfamila_provette.task1585_root09_hypernym_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1585_root09_hypernym_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1585_root09_hypernym_generation.usc-unified
Dataset Description
This dataset is part of a family of datasets that provide convenient access to
congressional data from the US Government Publishing Office
via the GovInfo Bulk Data Repository.
GovInfo provides bulk data in xml format.
The raw xml files were downloaded using the
congress repo.
Further processing was done using the
legisplain legisplain repo.
Hyperdemocracy Datasets
usc-billstatus (metadata on each bill)
usc-textversion (different text versions of… See the full description on the dataset page: https://huggingface.co/datasets/hyperdemocracy/usc-unified.
