spadeMIA/GoodWiki_Corpus_1024_2040
GoodWiki 1024–2040: paragraph-truncated MIA fine-tuning corpus A deterministic, paragraph-truncated corpus of English Wikipedia Good/Featured articles, built from euirim/goodwiki for membership inference attack (MIA) experiments on fine-tuned language models. Membership labels are defined relative to the fine-tuning population. train (10,000 rows) is the only split used for fine-tuning, and every row has label = 1. test (1,000 rows) remains held out, and every row has label =… See the full description on the dataset page: https://huggingface.co/datasets/spadeMIA/GoodWiki_Corpus_1024_2040.
Update GoodWiki canonical dataset identifier
Restore GoodWiki contributor attribution
Unify GoodWiki split schema for no-subset loading
Delete evaluation
docs: fix YAML thingy
docs: update README.md to be better basically
Upload dataset
Upload evaluation-00000-of-00001.parquet
data/evaluation-00000-of-00001.parquet
Rename data/dataevaluation-00000-of-00001.parquet to data/evaluation-00000-of-00001.parquet
Upload dataevaluation-00000-of-00001.parquet
Add dataset card with build statistics and MIA usage contract
Upload dataset
initial commit
