datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
robust-watermark-dataWatermark-or-Not-20K
Watermark-or-Not-20K Dataset
Overview
The Watermark-or-Not-20K dataset consists of 20,000 images annotated with binary labels indicating the presence or absence of a watermark. It is designed to support training and evaluation of models focused on watermark detection, which is useful for content filtering, copyright protection, and image moderation tasks.
Dataset Structure
Split: train
Number of samples: 20,000
Label Type: Categorical (2 classes)
Image… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Watermark-or-Not-20K.watermark-removal-logovisible-watermark-pita
Visible watermarks datasets
We have observed that while datasets such as COCO are available for object detection, the availability of datasets
specifically designed for the detection of watermarks added to images is significantly limited. Through our research,
we identified only one such dataset, which originates from the paper Wdnet: Watermark-Decomposition Network for
Visible Watermark Removal [1]. This dataset provides a collection of images along with their corresponding… See the full description on the dataset page: https://huggingface.co/datasets/bastienp/visible-watermark-pita.watermark-latent-citra-bukti
Citra bukti penelitian watermarking DCT pada ruang laten VAE
Sampel citra bukti untuk dashboard hasil skripsi
dashboard-latent-waterwark.
Repo ini hanya menyimpan berkas gambar; seluruh angka hasil ada di repo kode.
Isi
456 berkas PNG 512 x 512, yaitu 24 citra sampel pada 19 kondisi. Sampelnya 3 citra per
kelas gaya pada ordinal tetap 0, 20, dan 40 di dalam kelas, bukan dipilih menurut
akurasi, supaya tidak terbaca sebagai memilih hasil yang bagus saja.… See the full description on the dataset page: https://huggingface.co/datasets/ridloau543/watermark-latent-citra-bukti.watermarks-validationmats-gf-activation-watermark-demo
Subliminal activation-fingerprint watermark — reproduction bundle
A minimal, zero-training bundle to see the watermark fire on Qwen2.5-7B-Instruct
without retraining anything. It contains the secret key vectors, the frozen
detection probe set, and three representative trained LoRA students.
Full code, findings, and figures: https://github.com/Sid-MB/mats-gf-activation-watermark
(This bundle is a small subset; the 202 GB of full trained students is not
distributed here — see the… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/mats-gf-activation-watermark-demo.sora-watermark-dataset
Sora Watermark Detection Dataset
Dataset Description
This is an object detection dataset for detecting watermarks in Sora AI-generated videos. The dataset follows the YOLOv11 standard format and contains frame images extracted from Sora-generated videos along with their corresponding watermark annotations.
Dataset Statistics
Total Samples: 164 images
Training Set: 124 images
Validation Set: 21 images
Test Set: 19 images
Number of Classes: 1 (watermark)… See the full description on the dataset page: https://huggingface.co/datasets/LLinked/sora-watermark-dataset.lord-mea-benchmarkwatermark1_flowers_dataset
Dataset Card for "watermark1_flowers_dataset"
More Information needed
llm-watermark-detectionFADE-watermark-ocr
FADE Watermark OCR Dataset
Overview
This dataset contains watermarked images, their corresponding masks, the alpha values used for watermarking, and the actual text embedded (as a 9-digit number). It is designed to train and evaluate OCR models in the presence of watermarks.
Paper: FADE: Probing the Limits of VLMs on fine-grained OCR
Data Fields
Column Name
Data Type
Description
Image with watermark
binary
The raw binary bytes of the watermarked… See the full description on the dataset page: https://huggingface.co/datasets/deep9539/FADE-watermark-ocr.C4-contrastive-watermark
Dataset Card for "C4-contrastive-watermark"
More Information needed
scenery_watermarksDataset for watermark classification (no_watermark/watermark)~22k images, 512x512, manually annotatedadditional info - https://github.com/qwertyforce/scenery_watermarks
Watermark_Dataset_v2watermark_localizationArndee_WatermarkLabelled by human 100%.
watermarking-clips
Neural Watermarking Clips
Summary
This dataset contains 31 284 short audio clips collected as an unlabeled corpus for neural audio watermarking experiments. The clips cover environmental sounds, bird vocalizations, polyphonic music with predominant instruments and synthetic but realistic jazz drums and ragtime style piano.
Source folders and file counts
ARCA23K.audio 13 470 clips
ff101bird 7 690 clips
IRMAS training 6 706 clips
WaivOps ragtime piano 1 743 clips
WaivOps… See the full description on the dataset page: https://huggingface.co/datasets/benmainbird/watermarking-clips.PRC-watermark-imagesrepro-how-good-is-post-hoc-watermarking-with-language-model-rephrasing-traces
Agent traces
Agent sessions published from a Trackio Logbook.
vista-data-watermark-storeeval_test_watermark_random_yellowThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 1673,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sergiov2000/eval_test_watermark_random_yellow.xsum-watermarked-flan-t5-small-20240828230657eval_test_watermark_randomThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"total_videos": 0,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sergiov2000/eval_test_watermark_random.watermark-imageswatermark-datasetC4-contrastive-watermark
Dataset Card for "C4-contrastive-watermark"
More Information needed
llm-watermarking-papers
LLM Watermarking & Copyright Detection Papers — FineSet
A research-paper dataset on LLM Watermarking & Copyright Detection Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on LLM Watermarking & Copyright Detection Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom.… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/llm-watermarking-papers.watermark_img_kidsvista-data-watermark-parking-store
