CoolFace
Datasetpublic

spadeMIA/GoodWiki_Corpus_1024_2040

GoodWiki 1024–2040: paragraph-truncated MIA fine-tuning corpus A deterministic, paragraph-truncated corpus of English Wikipedia Good/Featured articles, built from euirim/goodwiki for membership inference attack (MIA) experiments on fine-tuned language models. Membership labels are defined relative to the fine-tuning population. train (10,000 rows) is the only split used for fine-tuning, and every row has label = 1. test (1,000 rows) remains held out, and every row has label =… See the full description on the dataset page: https://huggingface.co/datasets/spadeMIA/GoodWiki_Corpus_1024_2040.

sourceHugging Facecc-by-sa-4.0updated 2mo agoView on Hugging Face
0likes258downloads
14 commits on main
e65038a2mo ago

Update GoodWiki canonical dataset identifier

batukoray
4f756762mo ago

Restore GoodWiki contributor attribution

batukoray
3ff88952mo ago

Unify GoodWiki split schema for no-subset loading

batukoray
bc040f12mo ago

Delete evaluation

MeldaPaksoy
eea58eb2mo ago

docs: fix YAML thingy

batukoray
9f283a72mo ago

docs: update README.md to be better basically

batukoray
90562e82mo ago

Upload dataset

Batu Koray Masak
2f88acb2mo ago

Upload evaluation-00000-of-00001.parquet

MeldaPaksoy
ca93b182mo ago

data/evaluation-00000-of-00001.parquet

MeldaPaksoy
0fbae5e2mo ago

Rename data/dataevaluation-00000-of-00001.parquet to data/evaluation-00000-of-00001.parquet

MeldaPaksoy
48a8d652mo ago

Upload dataevaluation-00000-of-00001.parquet

MeldaPaksoy
6399b042mo ago

Add dataset card with build statistics and MIA usage contract

Batu Koray Masak
6f94b512mo ago

Upload dataset

Batu Koray Masak
fafc1652mo ago

initial commit

Batu Koray Masak