CoolFace
Datasetpublic

lightonai/fiqa-decontaminated

fiqa (Decontaminated) A decontaminated version of the fiqa dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed. Decontamination methodology Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files): Pass 1: Exact hash matching All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed with… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/fiqa-decontaminated.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes365downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
lightonai/fiqa-decontaminated · CoolFace