CoolFace
Datasetpublic

lightonai/webis-touche2020-decontaminated

webis-touche2020 (Decontaminated) A decontaminated version of the webis-touche2020 dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed. Decontamination methodology Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files): Pass 1: Exact hash matching All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/webis-touche2020-decontaminated.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes347downloads
2 commits on main
84c6c1f6mo ago

Upload folder using huggingface_hub

raphaelsty
9ffacec6mo ago

initial commit

raphaelsty