CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01swj0419 /WikiMIA 📘 WikiMIA Datasets The WikiMIA datasets serve as a benchmark designed to evaluate membership inference attack (MIA) methods, specifically in detecting pretraining data from extensive large language models. 📌 Applicability The datasets can be applied to various models released between 2017 to 2023: LLaMA1/2 GPT-Neo OPT Pythia text-davinci-001 text-davinci-002 ... and more. Loading the datasets To load the dataset: from datasets import load_dataset LENGTH =… See the full description on the dataset page: https://huggingface.co/datasets/swj0419/WikiMIA.text1K<n<10K20 likes916 downloads3y agoHugging Face02SimMIA /WikiMIA-25 WikiMIA-25 Dataset Summary WikiMIA-25 is an evaluation-only dataset used to assess membership inference attacks (MIA) on language models, with a particular focus on recent models. The dataset follows the dataset construction methodology introduced in WikiMIA-24, which includes newer non-member data to support evaluation under more recent training cutoff assumptions. Supported Tasks Membership Inference Attack (MIA) Binary classification (member vs.… See the full description on the dataset page: https://huggingface.co/datasets/SimMIA/WikiMIA-25.texttext-classification1K<n<10K1 likes144 downloads8mo agoHugging Face03hallisky /wikiMIA-2024-hard WikiMIA-2024 Hard Dataset Dataset Description WikiMIA_2024 Hard is a challenging dataset for membership inference attacks intorduced in the paper "The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage" containing temporal Wikipedia articles with different versions based on date cutoffs. This dataset is designed to evaluate the robustness of privacy-preserving machine learning models against sophisticated membership inference techniques. It… See the full description on the dataset page: https://huggingface.co/datasets/hallisky/wikiMIA-2024-hard.tabulartext-classification1K<n<10K0 likes137 downloads1y agoHugging Face04zjysteven /WikiMIA_paraphrased_perturbed 📘 WikiMIA paraphrased and perturbed versions The WikiMIA dataset serves as a benchmark designed to evaluate membership inference attack (MIA) methods, specifically in detecting pretraining data from extensive large language models. It is originally constructed by Shi et al. (see the original data repo for more details). The authors studied a paraphrased setting in their paper, where instead of detecting verbatim training texts, the goal is to detect (slightly) paraphrased… See the full description on the dataset page: https://huggingface.co/datasets/zjysteven/WikiMIA_paraphrased_perturbed.text10K<n<100K1 likes76 downloads2y agoHugging Face05wjfu99 /WikiMIA-24 📘 WikiMIA-24 Datasets The WikiMIA-24 datasets is a more up-to-date benchmark designed to evaluate pre-training data detection algorithms designed for large language models. The prior version of WikiMIA-24 can be found in WikiMIA 📌 Applicability The datasets can be applied to various models released between 2017 to 2024: Mistral Gemma LLaMA1/2 Falcon Vicuna Pythia GPT-Neo OPT ... and more. Loading the datasets To load the dataset: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/wjfu99/WikiMIA-24.text1K<n<10K3 likes75 downloads2y agoHugging Face06wjfu99 /WikiMIA-24-Old Dataset Card for "WikiMIA-24" More Information needed text1K<n<10K0 likes51 downloads2y agoHugging Face07S3IC /wikimia WikiMIA This repository hosts a copy of the widely used WikiMIA dataset,a benchmark designed to evaluate membership inference attack (MIA) methods—specifically for detecting whether a piece of text was seen during the pretraining of Large Language Models (LLMs). WikiMIA is commonly used in data contamination / pretraining data detection research, including the paper “Detecting Pretraining Data from Large Language Models” (arXiv:2310.16789). Contents… See the full description on the dataset page: https://huggingface.co/datasets/S3IC/wikimia.1K<n<10K0 likes43 downloads9mo agoHugging Face08wjfu99 /WikiMIA-24-perturbedIf you find our codebase and datasets beneficial, kindly cite our work: @inproceedings{fu2024membership, title={{MIA}-Tuner: Adapting Large Language Models as Pre-training Text Detector}, booktitle={Proceedings of the AAAI Conference on Artificial Intelligence}, author = {Fu, Wenjie and Wang, Huandong and Gao, Chen and Liu, Guanghua and Li, Yong and Jiang, Tao}, year = {2025}, address = {Philadelphia, Pennsylvania, USA} } text10K<n<100K0 likes31 downloads2y agoHugging Face09zjysteven /WikiMIA_concattextn<1K0 likes24 downloads2y agoHugging Face10lluvecwonv /WikiMIA_QAtabular1K<n<10K1 likes17 downloads2y agoHugging Face11iammytoo /wikiMIApiletext1K<n<10K0 likes12 downloads2y agoHugging Face12ADRA-RL /wikimia24_hard_unique_trio_ratio_1.50_adaptive_match_mink_pair_3_p0.25_a0.25textn<1K0 likes11 downloads7mo agoHugging Face13osieosie /wikimia24_hard_64_seed2-Paraphrased-Gemini-2.5-Flash-v3.0tabularn<1K0 likes10 downloads9mo agoHugging Face14ADRA-RL /WikiMIA-2024-Hard-Paraphrased-Gemini-2.5-Flashtabularn<1K0 likes10 downloads7mo agoHugging Face15wwml /wikimiatextn<1K0 likes6 downloads1y agoHugging Face16ADRA-RL /wikimia24_hard_unique_trio_ratio_1.50_pair_3_p0.25_a0.25textn<1K0 likes5 downloads7mo agoHugging Face17osieosie /wikimia24_hard_64_seed1-Paraphrased-Gemini-2.5-Flash-v3.0tabularn<1K0 likes4 downloads9mo agoHugging Face18ADRA-RL /wikimia24_hard_para_unique_ngram_coverage_ref_ratio_1.50_adaptive_match_pair_3_p0.25_a0.25textn<1K0 likes4 downloads7mo agoHugging Face19ADRA-RL /wikimia24_hard_paraphrased_unique_trio_ratio_1.50_pair_3_p0.25_a0.25textn<1K0 likes4 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.