CoolFace
Datasetpublic

zhliu/ArxivMIA

Dataset Card for ArxivMIA To evaluate various pre-training data detection methods in a more challenging scenario, we introduce ArxivMIA, a new benchmark comprising abstracts from the fields of Computer Science (CS) and Mathematics (Math) sourced from Arxiv. Repository: https://github.com/zhliu0106/probing-lm-data Paper: Probing Language Models for Pre-training Data Detection

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
0likes317downloads
14 commits on main
8a219a32y ago

Update README.md

zhliu
12cb12d2y ago

Update README.md

zhliu
7f0155e2y ago

Update README.md

zhliu
10948902y ago

Delete arxiv_mia_test.jsonl

zhliu
ea387222y ago

Delete arxiv_mia_dev.jsonl

zhliu
3b736862y ago

Delete arxiv_mia.jsonl

zhliu
7344e022y ago

Upload 3 files

zhliu
010331e2y ago

Update README.md

zhliu
34a442c2y ago

Update README.md

zhliu
b1321982y ago

Rename arxiv_mia_all.jsonl to arxiv_mia.jsonl

zhliu
3299b4c2y ago

Rename arxiv_mia.jsonl to arxiv_mia_all.jsonl

zhliu
631e5aa2y ago

Upload 3 files

zhliu
9b318382y ago

Update README.md

zhliu
a5a55c62y ago

initial commit

zhliu