datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Anti-UAV-RGBTSPADES-RGB3d-dlp-repro-genericshapes-rgb
GenericShapes-RGB — synthetic RGB-voxel tabletop scenes
Training/evaluation corpus built for an independent reproduction of ICML 2026 paper #10351,
3D-DLP: Self-supervised 3D Object-centric Scene Representation Learning
(OpenReview vIotI25gJz, code
github.com/Eubooks3003/3d-dlp).
The paper's GenericShapes corpus (Appendix B.2) is described but not released, and the authors'
released generator scripts/generate_ply.py
writes colourless point clouds — the "RGB-coloured variant used… See the full description on the dataset page: https://huggingface.co/datasets/rvt832/3d-dlp-repro-genericshapes-rgb.piperx-old-ab-rgbd-5903
PiperX Old AB RGB-D — 5,903 accepted points
This public dataset export contains the exact 5,903 accepted training Episodes recorded by the old AB SE(3) perturbation collector. The export is bound to the aggregate state.json truth and contains 5,903 unique schedule_index values across 124 closed shards (shard-00000 through shard-00123). Audit rejections remain audit records and are not training Episodes.
Data
Schema: piperx_lerobot_se3_rgbd_v2
Accepted Episodes /… See the full description on the dataset page: https://huggingface.co/datasets/Travor278/piperx-old-ab-rgbd-5903.eli5-human-vs-ai
ELI5 Human vs AI (long-form)
This dataset is for training and evaluating AI-writing detectors. It was built
as a clean way to compare known AI text against known human text: every human
answer predates ChatGPT by more than three years, so it is genuinely human by
construction, and every AI answer was written by a named 2026 model, so its origin
is certain too. Most detection datasets have to guess at their labels; this one
does not.
ELI5 answers were chosen because they are… See the full description on the dataset page: https://huggingface.co/datasets/mild-rgb/eli5-human-vs-ai.aita-human-vs-ai
AITA Human-vs-AI corpus (2026 generators)
A second human-vs-AI corpus, a companion to
mild-rgb/eli5-human-vs-ai,
in a deliberately different register: first-person judgment narratives from
r/AmItheAsshole, versus same-title posts written by seven 2026 models. 2,900
questions, one human post and one AI post each; ~414 documents per generator.
The human side is redacted — reconstruct it from Scruples
The human posts are verbatim r/AmItheAsshole text, obtained via… See the full description on the dataset page: https://huggingface.co/datasets/mild-rgb/aita-human-vs-ai.RoboFactory-5Task-RGBD-Decentralized
RoboFactory Five-Task Decentralized Wrist RGB-D
Public research corpus for reproducible Stereo-CoRE shared-policy experiments.
Contract
500 successful synchronized demonstrations: 100 each of LiftBarrier (2 robots), CameraAlignment (3), ThreeRobotsStackCube (3), LongPipelineDelivery (4), and TakePhoto (4).
A policy stream contains only one panda_hand wrist RGB-D observation and that robot's qpos.
RGB is 640x480; depth is native metric depth stored in millimetres.… See the full description on the dataset page: https://huggingface.co/datasets/B111ue/RoboFactory-5Task-RGBD-Decentralized.dusting_train_50_v3_stacked_qwen_rgb
