datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hyper-bot-datahyperliquid-node-fills-by-blockhypernet-prior-topic03
Hypernet — Prior Topic 03 Archive
Complete archive of the prior topic 03 research thread: per-shape SIREN
decoders, per-layer hypernetwork architectures, mapper experiments, and
ancillary checkpoints. This work predates the current image-to-3D pipeline
documented in hypernet-image-to-3d
and the main dataset.
100 shapes, naming convention obj_NN for NN in [0..99].
Contents
Path
Size
Description
watertight/
~5.6 GB
100 watertight .obj meshes (filename… See the full description on the dataset page: https://huggingface.co/datasets/bobthebuilderinternational/hypernet-prior-topic03.function-calling-sharegptThis is a dataset for finetuning models on function calling based on glaiveai/glaive-function-calling-v2.
The dataset includes 86,864 examples of chats that include function calling as part of the conversation. The system prompt includes either 0, 1, or 2 functions that the assistant can use, and instructions on how the agent can use it.
Changes include:
Using ShareGPT format for chats
Adding "function_response" as a role
Removing code examples
Removing examples with invalid JSON as function… See the full description on the dataset page: https://huggingface.co/datasets/hypervariance/function-calling-sharegpt.riddles_v1
Riddle Processing with GPT-4
Buy me Ko-fi
Credits
All credit for the original riddles goes to crawsome's GitHub repository.
Project Overview
This project involves processing each riddle using GPT-4. The correct answers were provided to the model to generate a desirable output focused on reasoning and logical breakdown.
riddles.json (riddles_1) — 386 samples, sourced from crawsome's GitHub repository.
riddles_2.json — 83 samples, sourced from various Google… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/riddles_v1.hypersim-episodes-v3-parquet
hypersim-episodes-v3-parquet
Per-frame Parquet dataset for ReCAST tracker training.
Schema
One row per frame, grouped by episode_id. Arrow memory-mapped access
enables reading specific frames without loading entire episodes.
Column
Type
Description
episode_id
int32
Episode identifier
frame_idx
int32
Frame index within episode
jpeg
binary
JPEG-encoded RGB frame
depth
list<float32>
Flat H×W depth map
seg
list<uint16>
Semantic segmentation (empty if… See the full description on the dataset page: https://huggingface.co/datasets/OSResight/hypersim-episodes-v3-parquet.hyperliquid-node-tradeshypersonic-scramjets
Scramjet Hypersonic Flow Physics Emulator Dataset
A dataset of 6,877 steady-state hypersonic flow simulations of parametrically
varied scramjet geometries, produced with JAX-Fluids (https://github.com/tumaer/JAXFLUIDS) as part of a fully GPU-based CFD workflow intended for
training and evaluating physics emulators.
The current version of the dataset contains the irregular-grid data used for training the AB-UPT emulator (https://arxiv.org/abs/2502.09692) from the paper.
We will… See the full description on the dataset page: https://huggingface.co/datasets/paischer101/hypersonic-scramjets.layout_diffusion_hypersimThis repository contains the data for SceneCraft: Layout-Guided 3D Scene Generation.
Project page: https://orangesodahub.github.io/SceneCraft
Code: https://github.com/OrangeSodahub/SceneCraft
hyperpartisan_news_detection_bypublisher_promptsourcehyperpartisan_news_detection
Dataset Card for "hyperpartisannewsdetection"
More Information needed
HyperStreamBridge_49_6corapatentpubmedcourseraimdbSteve_Jobs_Interviews
Steve Jobs Interviews Database
Support this project on Ko-fi
Project Overview
This project contains multiple interviews of Steve Jobs during his time before and after Apple.
Goal
The primary goal of this dataset was to fine-tune a language model to output Steve Jobs views and thoughts.
Performance
The performance of this small dataset is very noteworthy. Do to the nature of the database being interview question and answer pairs the replies of the… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/Steve_Jobs_Interviews.hyperliquid-data
Nautilus Data Repository
Automated data collection.
Competition-Submissions
Competition Submissions
A curated dataset of writing that models compassionate moral reasoning about nonhuman sentient beings — animals, insects, digital minds, and entities whose moral status is uncertain.
Designed for pretraining and fine-tuning language models to reason more carefully and compassionately when facing decisions that affect sentient life.
Why This Dataset Exists
Recent alignment research shows that training on synthetic documents depicting… See the full description on the dataset page: https://huggingface.co/datasets/Hyperstition-for-Good/Competition-Submissions.edgnn-hypergraph-dataset
Equivariant Hypergraph Diffusion Neural Operators
The official data release of ICLR 2023 paper Equivariant Hypergraph Diffusion Neural Operators.
Peihao Wang, Shenghao Yang, Yunyu Liu, Zhangyang (Atlas) Wang, Pan Li
Please refer to our GitHub repo for more details.
HyperCacheMatrix_501_6hyperliquid-misc-eventstom_cleanHyperThink-Max-200K
🔮 HyperThink
HyperThink is a premium, best-in-class dataset series capturing deep reasoning interactions between users and an advanced Reasoning AI system. Designed for training and evaluating next-gen language models on complex multi-step tasks, the dataset spans a wide range of prompts and guided thinking outputs.
🚀 Dataset Tiers
HyperThink is available in three expertly curated versions, allowing flexible scaling based on compute resources and training goals:… See the full description on the dataset page: https://huggingface.co/datasets/Sashvat/HyperThink-Max-200K.OS-Atlas_ScreenSpotHyperBrowseComp
HyperBrowseComp
Hard, multi-hop web search questions across 13 languages. Encrypted (AES-256-GCM).
from datasets import load_dataset
ds = load_dataset("afaji/HyperBrowseComp", split="test")
id is <LANG>_<question id>, e.g. KO_7388.
Decrypt
import base64
import hashlib
from cryptography.hazmat.primitives.ciphers.aead import AESGCM
def _nonce(canary, field):
return hashlib.sha256(f"{canary}:{field}".encode()).digest()[:12]
def _open(ciphertext_b64, key… See the full description on the dataset page: https://huggingface.co/datasets/afaji/HyperBrowseComp.hyperpartisan-cleanedHyperThink-X-Nvidia-Opencode-Reasoning-200K
🔮 HyperThink
HyperThink is a premium, best-in-class dataset series capturing deep reasoning interactions between users and an advanced Reasoning AI system. Designed for training and evaluating next-gen language models on complex multi-step tasks, the dataset spans a wide range of prompts and guided thinking outputs.
🚀 Dataset Tiers
HyperThink is available in three expertly curated versions, allowing flexible scaling based on compute resources and training goals:… See the full description on the dataset page: https://huggingface.co/datasets/Sashvat/HyperThink-X-Nvidia-Opencode-Reasoning-200K.amazon-berkeley-objects
Amazon Berkeley Objects
This is a Hugging Face metadata mirror of the Amazon Berkeley Objects dataset
for reproducible research and HyperView demos. The original dataset is provided
by Amazon.com and UC Berkeley.
This mirror stores metadata tables and official S3 asset URLs. It does not
duplicate catalog images, turntable images, or 3D models as binary files.
Load
from datasets import load_dataset
listings = load_dataset("hyper3labs/amazon-berkeley-objects"… See the full description on the dataset page: https://huggingface.co/datasets/hyper3labs/amazon-berkeley-objects.
