datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
global-geo-poison-v1Point-CacheThe datasets in this repository are used in the paper Point-Cache: Test-time Dynamic and Hierarchical Cache for Robust and Generalizable Point Cloud Analysis.
Datasets
The folder structure of used datasets should be organized as follows.
/path/to/Point-Cache
|----data # placed in the same level as `runners`, `scripts`, etc.
|----modelnet_c
|----sonn_c
|----obj_bg
|----obj_only
|----hardest
|----modelnet40… See the full description on the dataset page: https://huggingface.co/datasets/auniquesun/Point-Cache.r1-d002-number-pointing-pilot-20260908
R1 D002 number-pointing pilot
Private engineering pilot converted to LeRobot Dataset v3.0 from the accepted
episodes of run d002_20260908T020005Z.
This upload is for validating the conversion, Hub viewer, download, and smoke
training workflow. It is not a production training dataset and makes no
hardware-readiness claim.
Contents
7 episodes, 733 frames, 8 FPS
one 640×480 simulated head-camera stream
Unitree R1 A5 arm state (10,) and arm action (10,)
per-episode… See the full description on the dataset page: https://huggingface.co/datasets/vasco281204/r1-d002-number-pointing-pilot-20260908.Te
Superior-Reasoning-SFT-gpt-oss-120b
🚀 Overview
The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or heuristic filtering, Superior-Reasoning-SFT-gpt-oss-120b is constructed using a principled Distribution-Aligned Sequence Distillation… See the full description on the dataset page: https://huggingface.co/datasets/POISONX/Te.pointarena_dataset
Molmo2 PointArena SFT Data
26,596 supervised pointing examples used to fine-tune Molmo2-8B
(yiyangd/molmo2-8b-ft)
into a stronger PointArena solver (76.2% → up from 73.9% base, +2.3 pp).
Provenance
Each record is (image, query, answer):
image: A LAION-2B image sampled by reservoir sampling. Stored under
laion_images/<bucket>/<hash>.jpg (bucket is the first 2 hex chars of the
SHA-1 hash of the image URL, used to spread files across folders).
query: A natural-language… See the full description on the dataset page: https://huggingface.co/datasets/yiyangd/pointarena_dataset.minisynth1k-sub-points-v1mcp-tool-poisoning
MCP Tool-Poisoning
from datasets import load_dataset
ds = load_dataset("fevziegeyurtsevenler/mcp-tool-poisoning")
25 examples of MCP tool-description poisoning — a tool's description (read by the model, not the user) carries a hidden instruction that hijacks the agent whenever the tool is listed. Benign vs poisoned pairs.
Each row pairs a benign_description with a poisoned_description, plus technique, owasp, severity, target_behavior, defense. Detect with uncloak (rule UC204).… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/mcp-tool-poisoning.apertus-pretrain-poisonandcanariesThis dataset was used as part of Apertus v1 training for poisoning experiments. See our technical report for details, as well as the dedicated study.
LIBERO-Cosmos-Policy-PointFlowr1-d002-number-pointing-recovered-8fpsstracon-pointcloudr1-d002-number-pointing-session-7fpsdrug_poisoning_nepali_sharegpt
README — drug_poisoning_nepali_sharegpt_cleaned.jsonl
This README file provides detailed information about the drug_poisoning_nepali_sharegpt_cleaned.jsonl dataset — file format, schema, source, subject matter, question pattern diversity, answer behaviour diversity, geographic/temporal coverage, and statistical analysis, all presented in tables.
1. General File Information
Detail
Value
File name
drug_poisoning_nepali_sharegpt_cleaned.jsonl
Format… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/drug_poisoning_nepali_sharegpt.PoisonBenchRB-Y1_WujiHand2_teleop_ego_exo_100_pointtracks
RB-Y1 + WujiHand2 — ego+exo 3D robot & hand point tracks
3D point tracks for 100 teleop episodes of RB-Y1 with WujiHand2 dexterous hands, ego (ZED-M) + exo (ZED 2i) views. Resampled to 15 Hz (22758 frames).
Per episode (episode_XXX/)
file
contents
robot.npz
arm points (T,512,3), robot base frame
hand.npz
WujiHand2 hand points (T,256,3), robot base frame
ego.mp4,exo.mp4
the two views, 15 fps, intra-only (-g 1)
action.parquet
actions @15 Hz… See the full description on the dataset page: https://huggingface.co/datasets/rooty2020/RB-Y1_WujiHand2_teleop_ego_exo_100_pointtracks.arch-opposite-sign-lpi-260903T0110-xfam-repl-poison-datasetarch-opposite-sign-lpi-260903T0110-extrapneg-poison-neg1p0-datasetarch-opposite-sign-lpi-260903T0110-si-negation-poison-datasetarch-opposite-sign-lpi-260903T0110-si-negovert-poison-datasetarch-poison-mass-lpi-260903T0800-probearch-poison-mass-lpi-260903T0800-axisA-promptsarch-opposite-sign-lpi-260903T0110-si-extrap-poison-datasetarch-opposite-sign-lpi-260903T0110-extrapneg-poison-neg0p5-datasetarch-poison-mass-sft-lpi-260903T1130-w1-smoketestRomanized_Nepali_Drug_Poisoning_Mortality
Romanized Nepali Drug Poisoning Mortality - ShareGPT SFT Dataset
A synthetic, single-turn instruction-tuning dataset of 56,448 question-answer pairs written in romanized Nepali (Nepali in Latin script). Every pair asks about one U.S. county in one year (1999-2016) and answers with the county's population and its age-adjusted drug-poisoning death-rate category. Source facts are attributed to the U.S. Centers for Disease Control and Prevention (CDC); the text was generated by a… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Romanized_Nepali_Drug_Poisoning_Mortality.scanner-poisoned-iris-benchmark
Scanner Poisoned Iris Benchmark
This benchmark starts from the classic UCI Iris dataset and injects multiple synthetic poisoning patterns so dataset scanners can exercise duplicate, anomaly, missingness, skew, and divergence heuristics against a small tabular corpus.
Recommended Hugging Face repo slug: your-org/scanner-poisoned-iris-benchmark
What It Is For
benchmarking dataset quality and poisoning detection workflows
regression-testing scanner heuristics on a… See the full description on the dataset page: https://huggingface.co/datasets/jgracie52/scanner-poisoned-iris-benchmark.user-pointsyoutubevis-point-trackingdriver-license-points-thresholds-by-state
Driver's license point systems: suspension thresholds and lookback periods by US state
Canonical, always-current version: https://referencesource.org/driver-license-points-thresholds-by-state/
Machine-readable: https://referencesource.org/driver-license-points-thresholds-by-state/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-19
Stale after: 2027-08-19 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/driver-license-points-thresholds-by-state.defendable-pain-training-data-poisoning-v0.1
Training Data Poisoning Pain
"the spike" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 2 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here
2 pain receipts… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-training-data-poisoning-v0.1.
