datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm-watermarking-papers
LLM Watermarking & Copyright Detection Papers — FineSet
A research-paper dataset on LLM Watermarking & Copyright Detection Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on LLM Watermarking & Copyright Detection Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom.… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/llm-watermarking-papers.llm-watermark
LFM Tournament Watermark Responses
This dataset contains responses generated from prompts in the pair configuration of sentence-transformers/eli5.
It was created to evaluate how tournament depth affects detection of a SynthID-style, non-distortionary text watermark.
Data
1,000 ordinary responses
1,000 watermarked responses with 4 tournament layers
1,000 watermarked responses with 8 tournament layers
1,000 watermarked responses with 12 tournament layers
The… See the full description on the dataset page: https://huggingface.co/datasets/saad1926q/llm-watermark.
