CoolFace
Datasetpublic

lightonai/quora-decontaminated

quora (Decontaminated) A decontaminated version of the quora dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed. Decontamination methodology Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files): Pass 1: Exact hash matching All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed with… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/quora-decontaminated.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes422downloads
Dataset Card

quora (Decontaminated)

A decontaminated version of the quora dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed.

Decontamination methodology

Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files):

Pass 1: Exact hash matching

All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed with xxHash-64. The same normalization + hashing was applied to every query and document field in mgte-en. Any sample whose hash appeared in mgte-en was flagged as contaminated.

Pass 2: 13-gram containment (GPT-3 style)

Following the methodology introduced in the GPT-3 paper (Brown et al., 2020), word-level 13-grams were extracted from all remaining samples. For each sample, containment was computed as:

containment = |ngrams_in_sample ∩ ngrams_in_mgte| / |ngrams_in_sample|

Samples with containment >= 0.5 were flagged as near-duplicates.

Qrels filtering

Relevance judgments (qrels) referencing any removed query or corpus document were also removed.

Decontamination results

ComponentOriginalCleanRemoved
Corpus522,931413,157109,774
Queries15,0004,33210,668

Qrels per split

SplitOriginalCleanRemoved
test15,6754,51611,159
validation7,6261,8915,735

Usage

python
from datasets import load_dataset

corpus = load_dataset("lightonai/quora-decontaminated", "corpus", split="corpus")
queries = load_dataset("lightonai/quora-decontaminated", "queries", split="queries")

Citation

Please cite the original BEIR benchmark:

bibtex
@inproceedings{thakur2021beir,
  title={BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models},
  author={Thakur, Nandan and Reimers, Nils and Rücklé, Andreas and Srivastava, Abhishek and Gurevych, Irena},
  booktitle={NeurIPS Datasets and Benchmarks},
  year={2021}
}

License

MIT (same as original BEIR)