datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scifact-qrels
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/scifact-qrels.scifact-decontaminated
scifact (Decontaminated)
A decontaminated version of the scifact dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed.
Decontamination methodology
Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files):
Pass 1: Exact hash matching
All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed with… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/scifact-decontaminated.scifact_entailmentSciFact, a dataset of 1.4K expert-written scientific claims paired with evidence-containing abstracts, and annotated with labels and rationales.scifact-pl-qrelsPart of BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language.
Link to arxiv: https://arxiv.org/pdf/2305.19840.pdf
Contact: konrad.wojtasik@pwr.edu.pl
SciFact_VerifAI
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: a
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Dataset Structure
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/MilosKosRad/SciFact_VerifAI.scifact-decontaminated
scifact-decontaminated (MTEB layout)
Repackaging of lightonai/scifact-decontaminated
into the layout expected by MTEB:
a default config holding the qrels (one split per evaluation split), alongside
corpus and queries configs.
Rows and columns are copied verbatim from the source dataset; only the
config/split packaging differs. All credit for the decontaminated data belongs to
LightOn AI, and to the original BEIR
authors for the underlying benchmark.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/scifact-decontaminated.rus-scifact-qrelsscifact_ftThe dataset contains a random 0.7/0.1/0.2 train/dev/test splits of scifact dataset from BEIR https://github.com/beir-cellar/beir for benchmarking embedding model fine-tuning.
scifact-relevance-pairs
SciFact Evidence Relevance Pairs
Custom (claim, document, label) pairs derived from BEIR SciFact for binary evidence relevance classification: given a scientific claim and a candidate paper field, decide whether the paper is relevant evidence for the claim.
This dataset accompanies the scifact-relevance-classifier project, built as the Lab 3 / Assignment 1 deliverable for Information Retrieval 5LN712 (Master's in Language Technology, Uppsala University, 2026).
Quick… See the full description on the dataset page: https://huggingface.co/datasets/andreiaalexa/scifact-relevance-pairs.sheared-llama-scifact-resultsscifact_entailmentSciFact entailment pairs (data-only; train/validation).
scifact-tr-qrels
SciFact-TR
This is a Turkish translated version of the SciFact dataset.
Dataset Sources
Repository: SciFact
alibaba_scifact_chunkedsalesforce_scifact_queriesopenai_scifact_queriesalibaba_scifact_queriesprocessed_scifact_augmentgoogle_scifact_queriessheared-llama-scifact-results-newscifact_translatedprocessed_scifact_chunkedlinq_scifact_queriesscifact_neomme_260m_li
scifact_neomme_260m_li
Multi-vector (late-interaction) embeddings of BEIR scifact (beir/scifact/test), encoded with
Hcompany/NeoMME-260M-Retriever-ST-late at revision 023be2a8ab9d797f5aa76f5bf8b5dde78d819659.
Source data: ir_datasets beir/scifact/test (ir_datasets 0.6.3), which downloads scifact.zip (md5 5f7d1de60b170fc8027bb7898e2efca1). BEIR also publishes this corpus on the Hub as BeIR/scifact, whose card gives this dataset's license; the data here was loaded through… See the full description on the dataset page: https://huggingface.co/datasets/robro612/scifact_neomme_260m_li.
