datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
benchmark-datasets
Latency-Sensitive Bench datasets
Accepted zero-latency teacher rollouts for the supported benchmark tasks.
Viewer subsets
humanoidbench_balance_simple: 90 training and 10 validation episodes. The
observation.image values are PNG bytes declared as the Hugging Face Image
feature, so the Dataset Viewer renders them instead of showing their encoded
representation. Canonical LeRobot MP4 files remain under each split's
videos/ directory.
mikasa_intercept_grab_fast:… See the full description on the dataset page: https://huggingface.co/datasets/latency-sensitive-bench/benchmark-datasets.Blockchain-Sensitive-Detect-Data
Blockchain-Sensitive-Detect-Data
English README
复旦大学附属儿科医院-区块链敏感信息检测项目的多模态完整测试数据集。
项目仓库:https://github.com/anyangsong/Blockchain-Sensitive-Detect
数据集:https://huggingface.co/datasets/anyangsong/Blockchain-Sensitive-Detect-Data
checkpoints:https://huggingface.co/anyangsong/Blockchain-Sensitive-Detect-Checkpoints
数据以原始文件夹组织,覆盖文本、音频、图像与视频等样本。
该仓库不提供统一的 CSV、Parquet 或 JSONL 清单;类别信息主要由目录名和文件名携带。
内容警告: 数据集包含辱骂、性内容、暴力、政治相关内容、误导性医疗信息、欺诈信息。使用者应仅在具备适当访问控制、伦理审查和当地法律依据的环境中处理这些内容。… See the full description on the dataset page: https://huggingface.co/datasets/anyangsong/Blockchain-Sensitive-Detect-Data.Prompt_injection_and_Sensitive_Data_exposure_detectioncontextual-sensitive-data
Towards Contextual Sensitive Data Detection
This dataset includes tables with sensitivity annotations that were used to train and evaluate methods for detecting contextual sensitive data. It accompanies the paper "Towards Contextual Sensitive Data Detection".
Links:
Paper: https://huggingface.co/papers/2512.04120
Code: https://github.com/trl-lab/sensitive-data-detection
Sample Usage
The GitHub repository provides scripts for running inference and fine-tuning using… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/contextual-sensitive-data.synthetic-sensitive-data-in-source-code-n300
Synthetic Sensitive Data in Source Code (N=300)
Synthetic dataset of 300 source-code / config snippets containing hardcoded secrets and PII.Every sample includes at least one sensitive finding (no clean negatives).
Designed for evaluating local masking, secret detection, and OWASP LLM02 — Sensitive Information Disclosure scenarios in AI-assisted coding workflows.
Version 1.2: multi_secret (and related) samples label every secret present in code_text (complete ground truth).
All… See the full description on the dataset page: https://huggingface.co/datasets/nisaefendioglu/synthetic-sensitive-data-in-source-code-n300.arthur_sensitive_data_passwordsensitive_data_datasetsensitive_data_oneshotdataset-filter-comparisonsensitive_data_svSynthetic-Sensitive-datavietnamese_text_sensitive_dataset
Vietnamese Text Sensitive Dataset
Mô tả
vietnamese_text_sensitive_dataset là bộ dữ liệu chứa các văn bản tiếng Việt nhạy cảm liên quan đến nội dung khiêu dâm, bạo lực, phân biệt đối xử, sai lệch chính trị và các chủ đề khác. Bộ dữ liệu này có thể được sử dụng để huấn luyện các mô hình AI nhằm phát hiện và lọc nội dung nhạy cảm trong các ứng dụng xử lý ngôn ngữ tự nhiên (NLP).
Cấu trúc dữ liệu
Bộ dữ liệu bao gồm các danh mục sau:
Nội dung khiêu dâm, nhạy cảm… See the full description on the dataset page: https://huggingface.co/datasets/huytx267/vietnamese_text_sensitive_dataset.Prompt_injection_and_Sensitive_Data_exposure_detection
