CoolFace
Datasetpublic

zhliu/ArxivMIA

Dataset Card for ArxivMIA To evaluate various pre-training data detection methods in a more challenging scenario, we introduce ArxivMIA, a new benchmark comprising abstracts from the fields of Computer Science (CS) and Mathematics (Math) sourced from Arxiv. Repository: https://github.com/zhliu0106/probing-lm-data Paper: Probing Language Models for Pre-training Data Detection

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
0likes318downloads
Dataset Card

Dataset Card for ArxivMIA

To evaluate various pre-training data detection methods in a more challenging scenario, we introduce ArxivMIA, a new benchmark comprising abstracts from the fields of Computer Science (CS) and Mathematics (Math) sourced from Arxiv.

  • —Repository: https://github.com/zhliu0106/probing-lm-data
  • —Paper: Probing Language Models for Pre-training Data Detection