datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trellismark-qwen3-4b
TrellisMark Qwen3-4B confirmation corpus
This is the frozen English confirmation corpus for
TrellisMark, an experimental
many-user AI-text watermark. It includes exact generated text and token IDs,
unwatermarked Qwen controls, public-key detector evidence, the public research
key, independent encoder vectors, and the result reports used for the
reader-facing curves. The standalone implementation, detector, and
reproduction instructions are in the
TrellisMark GitHub repository.… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/trellismark-qwen3-4b.trellismark-qwen3-4b-open-corpus
TrellisMark Qwen3-4B open-corpus benchmark
This is the reproducible open-corpus companion to the main
TrellisMark Qwen3-4B confirmation corpus.
It is an explicitly derived subset of that release, not a separately generated
corpus. The standalone detector, benchmark runner, tests, and exact Viterbi
implementation are in the
TrellisMark GitHub repository.
The benchmark studies an intentionally difficult mixed setting: a corpus of
separately prompted documents, with hidden… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/trellismark-qwen3-4b-open-corpus.trellismark-qwen3-4b-rephrasing
TrellisMark Qwen3-4B blind rephrasing corpus
This separate release contains the frozen detector-blind rephrasing experiment
for TrellisMark. Its 16,384
source documents are an exact document-ID-preserving subset of the
main TrellisMark Qwen3-4B confirmation corpus.
It publishes both rewriters' outcome records, retained rewrite text and token
IDs, aligned key-only and model-assisted evidence, and the reports behind the
article's countermeasure figures.
The rewriters received only… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/trellismark-qwen3-4b-rephrasing.
