CoolFace
Datasetpublic

lightonai/fiqa-decontaminated

fiqa (Decontaminated) A decontaminated version of the fiqa dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed. Decontamination methodology Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files): Pass 1: Exact hash matching All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed with… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/fiqa-decontaminated.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes452downloads
Dataset Card

fiqa (Decontaminated)

A decontaminated version of the fiqa dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed.

Decontamination methodology

Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files):

Pass 1: Exact hash matching

All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed with xxHash-64. The same normalization + hashing was applied to every query and document field in mgte-en. Any sample whose hash appeared in mgte-en was flagged as contaminated.

Pass 2: 13-gram containment (GPT-3 style)

Following the methodology introduced in the GPT-3 paper (Brown et al., 2020), word-level 13-grams were extracted from all remaining samples. For each sample, containment was computed as:

containment = |ngrams_in_sample ∩ ngrams_in_mgte| / |ngrams_in_sample|

Samples with containment >= 0.5 were flagged as near-duplicates.

Qrels filtering

Relevance judgments (qrels) referencing any removed query or corpus document were also removed.

Decontamination results

ComponentOriginalCleanRemoved
Corpus57,63847,61710,021
Queries6,6487965,852

Qrels per split

SplitOriginalCleanRemoved
test1,7061381,568
train14,1661,19812,968
validation1,238951,143

Usage

python
from datasets import load_dataset

corpus = load_dataset("lightonai/fiqa-decontaminated", "corpus", split="corpus")
queries = load_dataset("lightonai/fiqa-decontaminated", "queries", split="queries")

Citation

Please cite the original BEIR benchmark:

bibtex
@inproceedings{thakur2021beir,
  title={BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models},
  author={Thakur, Nandan and Reimers, Nils and Rücklé, Andreas and Srivastava, Abhishek and Gurevych, Irena},
  booktitle={NeurIPS Datasets and Benchmarks},
  year={2021}
}

License

MIT (same as original BEIR)