CoolFace
Datasetpublic

lightonai/hotpotqa-decontaminated

hotpotqa (Decontaminated) A decontaminated version of the hotpotqa dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed. Decontamination methodology Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files): Pass 1: Exact hash matching All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/hotpotqa-decontaminated.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes374downloads
filecorpus.parquet314.6 MBdownload
fileqrels_test.parquet52 KBdownload
fileqrels_train.parquet4 KBdownload
fileqrels_validation.parquet41 KBdownload
filequeries.parquet1.1 MBdownload

lightonai/hotpotqa-decontaminated · main · files are served by the source, never re-hosted here