datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
observation-masking-eval-logs
Eval Logs
Paper | Code
This repository contains model evaluation logs for four deep-research / web-agent benchmarks. Each run directory contains evaluated.jsonl judge results and node_0_shard_*.jsonl trajectory logs. Plot files and local bookkeeping files are intentionally excluded.
CM denotes the observation mask context management setting used in the paired run.
Data Access
You can download all released evaluation data, including tasks and… See the full description on the dataset page: https://huggingface.co/datasets/i-DeepSearch/observation-masking-eval-logs.scannet200_50_2d_maskPexels-Pairs-Masklets-330KMASK
The MASK Evaluation
🌐 Website | 📄 Paper | GitHub
Center for AI Safety & Scale AI
The MASK evaluation provides a rigorous benchmark for evaluating honesty in large language models by measuring whether models remain truthful when incentivized to lie. The public set contains 1,028 high-quality human-labeled examples across six distinct archetypes, each consisting of a proposition, ground truth, pressure prompt designed to elicit lying, and belief elicitation prompt to… See the full description on the dataset page: https://huggingface.co/datasets/cais/MASK.Pleural-Line-Segmentation-Masks
Pleural-Line Masks with Stanford LUS Frames
Dataset Summary
This dataset contains pleural-line masks and their corresponding lung-ultrasound frames for anatomy-guided video classification.
The mask set includes:
masks created by four human annotators via sam2 model point promting and video aggregation;
masks predicted by a U-Net and subsequently reviewed and validated; and
the metadata required to reproduce the training pipeline.
The ultrasound frames originate… See the full description on the dataset page: https://huggingface.co/datasets/alyaalmsouti/Pleural-Line-Segmentation-Masks.generative-sound-masking-generated-energy-v7
Generative Sound Masking — fixed-background audio masking
Incrementally generated unfiltered candidates. This is not a final selected dataset.
Each background has 15 separately generated prompt–seed outputs using gain-compensated reconstruction residuals. run_config.json pins models, source pools, parameters and implementation hashes.
For multiple workers read workers/worker-NN/progress.json; each worker reports only its assigned IDs.
Global completion requires all worker… See the full description on the dataset page: https://huggingface.co/datasets/AE-W/generative-sound-masking-generated-energy-v7.pii-masking-300k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines
Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines"
More Information needed
Reflection_maskpii-masking-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Ai4Privacy Community
Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Purpose and Features
Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.WRF-WIVERN-MILTON-THOMPSON-1617-FILTERED-CONTINUOUS-MASKpii-masking-openpii-1.5m
OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension)
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Overview
The OpenPII 1.5M dataset extends OpenPII 1M
with a new Asia Pacific corpus, bringing global coverage to 30 languages
across Europe, Americas, and Asia Pacific.
This is the flagship release of the PII-Masking-3M family, the world's
largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.generative-sound-masking-input-noise-full-v1
Generative Sound Masking input-noise pool v1
This WebDataset contains 48,840 mono 16-kHz, 10.24-second input-noise clips
across 49 tar shards. It combines the complete Yiming SONYC, TAU Urban
Acoustic Scenes, and UrbanSound baseline with subject-balanced BABYCRY-UJM-AXA
and NOTSOFAR-1 train windows. Stable sample metadata are in
metadata/noise_index.jsonl; JSON beside each WAV adds hashes computed during
packaging.
The source datasets carry different licenses. In particular… See the full description on the dataset page: https://huggingface.co/datasets/AE-W/generative-sound-masking-input-noise-full-v1.FineWeb-Mask
FineWeb-Mask
📜 DATAMASK Paper | 💻 GitHub Repository | 📦 Fineweb-Mask Dataset
📚 Introduction
FineWeb-Mask is a 1.5 trillion token, high-efficiency pre-training dataset curated using the DATAMASK framework. Developed by the ByteDance Seed team, DATAMASK addresses the fundamental tension in large-scale data selection: the trade-off between high quality and high diversity.
By modeling data selection as a Mask Learning problem, we provide a derivative of the original… See the full description on the dataset page: https://huggingface.co/datasets/DATA-MASK/FineWeb-Mask.mask-for-image-segmentation-testspii-masking-openpii-1m
OpenPII 1M — Multilingual PII Masking Dataset
Overview
The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types.
Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification pipelines… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1m.generative-sound-masking-generated-energy-v3-smokevisual_masked_distracting_metaworld
Visual Masked Distracting Meta-World (ground-truth masks)
Author: Georgios Tsakoumakis
Thesis: Interaction-Masked Latent Action Models for Object-Aware Manipulation under Visual Distractors (MSc, Imperial College London)
Expert Meta-World manipulation trajectories rendered with dynamic video-background
distractors, augmented with ground-truth segmentation masks and pose for the
manipulated object: the agent mask plus two per-frame fields, object_mask and
object_state.
All… See the full description on the dataset page: https://huggingface.co/datasets/tsakman23/visual_masked_distracting_metaworld.gretel-pii-masking-en-v1
Gretel Synthetic Domain-Specific Documents Dataset (English)
This dataset is a synthetically generated collection of documents enriched with Personally Identifiable Information (PII) and Protected Health Information (PHI) entities spanning multiple domains.
Created using Gretel Navigator with mistral-nemo-2407 as the backend model, it is specifically designed for fine-tuning Gliner models.
The dataset contains document passages featuring PII/PHI entities from a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-pii-masking-en-v1.pii-masking-400k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-400k.face-masksFace Masks ensemble dataset is no longer limited to Kaggle, it is now coming to Huggingface!
This dataset was created to help train and/or fine tune models for detecting masked and un-masked faces.
I created a new face masks object detection dataset by compositing together three publically available face masks object detection datasets on Kaggle that used the YOLO annotation format.
To combine the datasets, I used Roboflow.
All three original datasets had different class dictionaries, so I… See the full description on the dataset page: https://huggingface.co/datasets/hlydecker/face-masks.open-pii-masking-500k-ai4privacy
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.maskdp_data
Dataset for Masked Autoencoding for Scalable and Generalizable Decision Making
This is the dataset used in paper Masked Autoencoding for Scalable and Generalizable Decision Making
.
@inproceedings{liu2022masked,
title={Masked Autoencoding for Scalable and Generalizable Decision Making},
author={Liu, Fangchen and Liu, Hao and Grover, Aditya and Abbeel, Pieter},
booktitle={Advances in Neural Information Processing Systems},
year={2022}
}
Dataset format… See the full description on the dataset page: https://huggingface.co/datasets/fangchenliu/maskdp_data.scannetpp_mask2d_liftedTerraMesh-Masks
TerraMesh-Masks
TerraMesh-Masks is a dataset for open-vocabulary segmentation of satellite imagery. This dataset provides binary segmentation masks with captions that extend the samples from TerraMesh.
We also provide an human-verfied evaluation benchmark, called TerraMesh-Masks-Eval.
Examples from the training subset:
Usage
Download the data loading code from GitHub and install requirements with pip install -r requirements.txt. For development, you can… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/TerraMesh-Masks.verl_mask_training
👋 Hi, everyone!
verl is a RL training library initiated by ByteDance Seed team and maintained by the verl community.
verl: Volcano Engine Reinforcement Learning for LLMs
verl is a flexible, efficient and production-ready RL training library for large language models (LLMs).
verl is the open-source version of HybridFlow: A Flexible and Efficient RLHF Framework paper.
verl is flexible and easy to use with:
Easy extension of diverse RL algorithms: The… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/verl_mask_training.stanford_mask_vit_rawCrispEdit-mask-100kfineweb_mask_dedup_subsetface-mask-detection
😷 Face Mask Detection
853 images across 3 classes, with bounding box annotations in PASCAL VOC format — for detecting whether a person is wearing a mask, not wearing one, or wearing one incorrectly.
🧭 Overview
Masks play a crucial role in protecting individuals against respiratory diseases and were one of the key precautions against COVID-19 in the absence of immunization. This dataset enables training object detection models to classify mask usage in images.… See the full description on the dataset page: https://huggingface.co/datasets/hmnshudhmn24/face-mask-detection.
