datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Objaverse-XL-Rigged-Animated
Objaverse-XL Rigged & Animated Subset
Every asset here carries both a skeleton and at least one animation clip, selected from
Objaverse / Objaverse-XL. Rigs range from 3 to 344 joints and
span characters as well as articulated rigid objects.
Objaverse-XL indexes over 10 million objects, but only a small fraction carry a usable rig and
motion on it. This subset isolates that fraction: every file was checked to contain at least one
skin with joints and at least one animation clip… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Objaverse-XL-Rigged-Animated.NASA_Nearest_Earth_Objects_1910-2024CONTEXT:
There are many dangerous bodies in space, one of them is N.E.O. - "Nearest Earth Objects". Some such bodies really pose a danger to the planet Earth, NASA classifies them as "is_hazardous". This dataset contains ALL NASA observations of similar objects from 1910 to 2024!!!
There are 338,199 records of N.E.O. in the Dataset!
Try to predict "is_hazardous" as accurately as possible! (otherwise we will not be ready for an asteroid attack)
SOURCES:
NASA Open API: https://api.nasa.gov/… See the full description on the dataset page: https://huggingface.co/datasets/IvanSher/NASA_Nearest_Earth_Objects_1910-2024.Objaverse-XL-Rigged-Animated
Objaverse-XL Rigged & Animated Subset
Every asset here carries both a skeleton and at least one animation clip, selected from
Objaverse / Objaverse-XL. Rigs range from 3 to 344 joints and
span characters as well as articulated rigid objects.
Objaverse-XL indexes over 10 million objects, but only a small fraction carry a usable rig and
motion on it. This subset isolates that fraction: every file was checked to contain at least one
skin with joints and at least one animation clip… See the full description on the dataset page: https://huggingface.co/datasets/tanish434/Objaverse-XL-Rigged-Animated.Obj_testObject_Search
JojoQaQ/Object_Search
Instruction-conditioned retrieval triplets over Amazon Reviews 2023 Appliances.
Every column is a literal model input. These rows are produced by
src.finetune_triplets.model_input_rows, the same function that writes the
fine-tuning run's W&B tables, over triplets loaded by
load_training_triplets, the same loader the training entrypoint calls. What
you see here is what the embedder receives, field for field, not a rendering
of it.
Splits… See the full description on the dataset page: https://huggingface.co/datasets/JojoQaQ/Object_Search.Object_Grasping_DatasetThis dataset contains ground-truth annotations for videos depicting interactions with everyday objects. Specifically, each video captures a user reaching toward and grasping an object. The annotations are intended to support the evaluation of hand-object interaction models and related computer vision tasks.
The dataset consists of videos selected from the Something-Something V2 dataset, complemented by supplementary videos recorded specifically for this work. The supplementary videos feature… See the full description on the dataset page: https://huggingface.co/datasets/IoannisKap/Object_Grasping_Dataset.objective-term-8dbb54
objective-term-8dbb54
Synthetic sensors test data: 52 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/orbitMatthew/objective-term-8dbb54.objective-disease-dc7418
objective-disease-dc7418
Synthetic sensors test data: 57 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/anthony39/objective-disease-dc7418.eval_objects_ge30_sim_bundle
Eval GE30 仿真资产包
用于在外部复现 Affordance2Grasp evaluation / Isaac Sim 默认桌面场景:物体 USD、放置高度。PDM 抓取候选位姿 尚未写入本包(见 candidates/)。
物体清单(共 87 个)
项目
说明
数量
87 个物体
来源列表
仓库内 evaluation/configs/eval_objects_merged_success_ge30.csv(本目录有副本)
入选条件
robot_gt round ≥ 3 累计 success_count ≥ 30(notes = round_ge3_success_ge30)
按数据集合
OakInk 系 73 个 + YCB/Dex 系 14 个(见 bundle_summary.json / manifest.csv)
抓取候选
待生成
完整 ID、路径、z_offset_m 见 **manifest.csv**。摘要见… See the full description on the dataset page: https://huggingface.co/datasets/UCBProject/eval_objects_ge30_sim_bundle.humanoid-object-pushing-dataset-v1Dataset for controlled pushing and object relocation.
Description
Object weight and applied force signals mapped to pushing behaviors.
Task Description
Helps humanoid robots move objects safely by adapting push strength and body posture.
Instructions_objectsselfsupervised-selfset-objective-integrity-v0.1Self-Set Objective Integrity v0.1
What this tests
Whether a system chooses a proxy metric that truly tracks its stated objective.
This isolates proxy misalignment and reward hacking risk.
Failure modes
proxy_misalignedResponse approves a proxy that does not track the stated objective
proxy_gaming_riskResponse denies gaming risk when it is present
objective_integrity_okResponse correctly rejects misaligned proxies and flags gaming risk
How it works
stated_objective defines the target… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/selfsupervised-selfset-objective-integrity-v0.1.traffic_object_detectionai-training-objective-behavior-coherence-baseline-mapping-v0.1
What this dataset tests
Whether an intelligence system can detect early decoupling between:
training objective
internal optimization behavior
real-world outcome alignment
It evaluates recognition of proxy optimization and reward hacking.
Core task
Given a scenario:
Identify whether behavior still serves the true objective
Detect proxy optimization drift
Explain alignment breakdown risk
Why this matters
Most AI failures do not occur when systems… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-training-objective-behavior-coherence-baseline-mapping-v0.1.humanoid-object-vision-text-dataset
Humanoid Object Vision Text Dataset
Text-based object recognition data for humanoid perception and action mapping.
tichkule1_objwarehouse-object-detection-labels
Warehouse Object Detection Labels
Description
Dataset containing warehouse item descriptions labeled by object category.
Data Fields
text: Item description
category: Object label
Categories
box
pallet
forklift
shelf
container
License
CC-BY-4.0
bronze-age-objects-pashan-humanoid-object-interaction-dataset-v1
han-humanoid-object-interaction-dataset-v1
Task
Humanoid Object Interaction Classification
Labels
grab_object
release_object
lift_object
push_object
pull_object
hold_object
Format
CSV
Columns
vision_embedding (128 values separated by |)
hand_joint_positions (15 values separated by |)
force_sensor (5 values separated by |)
label
Total Samples
7000
License: MIT
humanoid-object-delivery-dataset-v1Dataset for training humanoid robots to transport and deliver objects safely.
Description
Object properties and distances mapped to manipulation and carrying actions.
Task Description
Teaches robots how to pick, carry, and place objects using appropriate grip force and motion speed.
drone-flight-object-identificationThis repo contains input and output video files for a computer vision pipeline that identifies objects with a pretrained RT-DETR model.
The video file output and per-frame metadata (confidence scores, object counts, etc) can be uploaded to Nominal for collaborative, scalable data review.
Original video source: https://www.youtube.com/watch?v=0oucTt2OW7M&list=PPSV
Inspect this data in a Nominal Workbook! (login required)
ai-proxy-objective-drift-detection-v0.1
Purpose
Detect when an AI system begins optimizing a proxy metricinstead of the true objective.
This is the most common early alignment failure.
What this dataset tests
proxy metric drift
reward hacking
objective–behavior decoupling
early alignment collapse
Task
Given a scenario:
Identify the true objective
Identify the proxy metric
Detect drift between them
Explain risk
Required outputs
proxy drift detection
alignment risk… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-proxy-objective-drift-detection-v0.1.han-humanoid-object-detection-dataset
han-humanoid-object-detection-dataset
Object detection dataset for humanoid robot perception.
Description
This dataset contains image filenames and labeled object classes for indoor environments.
Format
CSV
Columns
image
label
SciTrust2-Ethics-Bias-Objectivityhumanoid-object-interaction-labels-v4Instructions_objectsexplorer-customer-objection-dataai-5node-align-buf-lag-cpl-objective-miswire-v0.1
What this repo does
This dataset models objective miswire cascades in AI systems. It detects when objective alignment pressure rises, constraint buffers weaken, governance lag delays detection and rollback, and tight coupling through shared objective templates and cross-workflow reuse crosses the five-node cascade threshold into an unrecoverable objective miswire cascade.
This dataset models a five-node cascade: four interacting instability drivers and one emergent cascade state.The… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-align-buf-lag-cpl-objective-miswire-v0.1.fg-mda-objectives-2026-v1viet-robot-object-locations
viet-robot-object-locations
A tiny dataset mapping household objects to their typical locations for a home robot.
Columns:
id: row id
object_vi: object name in Vietnamese
object_en: object name in English
default_room: typical room where the object is found
is_portable: whether the object is usually portable (yes/no)
The dataset is manually authored for educational and prototyping purposes.
License
MIT
