CoolFace
Datasetpublic

lightonai/quora-decontaminated

quora (Decontaminated) A decontaminated version of the quora dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed. Decontamination methodology Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files): Pass 1: Exact hash matching All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed with… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/quora-decontaminated.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes362downloads
2 commits on main
a3039666mo ago

Upload folder using huggingface_hub

raphaelsty
8764bd56mo ago

initial commit

raphaelsty