masking
bert-base-japanese-whole-word-maskingbert-large-uncased-whole-word-masking-squad2bert-large-cased-whole-word-masking-finetuned-squadbert-large-uncased-whole-word-masking-finetuned-squadbert-large-uncased-whole-word-maskingbert-large-uncased-whole-word-maskingbert-large-cased-whole-word-maskingbert-base-japanese-char-whole-word-masking
Datasets
All datasets matching “masking”observation-masking-eval-logs
Eval Logs
Paper | Code
This repository contains model evaluation logs for four deep-research / web-agent benchmarks. Each run directory contains evaluated.jsonl judge results and node_0_shard_*.jsonl trajectory logs. Plot files and local bookkeeping files are intentionally excluded.
CM denotes the observation mask context management setting used in the paired run.
Data Access
You can download all released evaluation data, including tasks and… See the full description on the dataset page: https://huggingface.co/datasets/i-DeepSearch/observation-masking-eval-logs.generative-sound-masking-generated-energy-v7
Generative Sound Masking — fixed-background audio masking
Incrementally generated unfiltered candidates. This is not a final selected dataset.
Each background has 15 separately generated prompt–seed outputs using gain-compensated reconstruction residuals. run_config.json pins models, source pools, parameters and implementation hashes.
For multiple workers read workers/worker-NN/progress.json; each worker reports only its assigned IDs.
Global completion requires all worker… See the full description on the dataset page: https://huggingface.co/datasets/AE-W/generative-sound-masking-generated-energy-v7.pii-masking-300k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.pii-masking-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Ai4Privacy Community
Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Purpose and Features
Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.pii-masking-openpii-1.5m
OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension)
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Overview
The OpenPII 1.5M dataset extends OpenPII 1M
with a new Asia Pacific corpus, bringing global coverage to 30 languages
across Europe, Americas, and Asia Pacific.
This is the flagship release of the PII-Masking-3M family, the world's
largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.generative-sound-masking-input-noise-full-v1
Generative Sound Masking input-noise pool v1
This WebDataset contains 48,840 mono 16-kHz, 10.24-second input-noise clips
across 49 tar shards. It combines the complete Yiming SONYC, TAU Urban
Acoustic Scenes, and UrbanSound baseline with subject-balanced BABYCRY-UJM-AXA
and NOTSOFAR-1 train windows. Stable sample metadata are in
metadata/noise_index.jsonl; JSON beside each WAV adds hashes computed during
packaging.
The source datasets carry different licenses. In particular… See the full description on the dataset page: https://huggingface.co/datasets/AE-W/generative-sound-masking-input-noise-full-v1.
