datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scifact
Dataset Card for BEIR Benchmark
scifact is one of the datasets from the Fact Checking task within BEIR, measuring scientific article retrieval for a given scientific claim.
Dataset Summary
BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks.
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/scifact.scifact-decontaminated
scifact (Decontaminated)
A decontaminated version of the scifact dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed.
Decontamination methodology
Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files):
Pass 1: Exact hash matching
All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed with… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/scifact-decontaminated.beir-scifact
SciFact — BEIR, unified schema
A normalised copy of the dataset behind the mteb task SciFact, one of the tasks of the BEIR benchmark as mteb defines it. Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/scifact @ d56462d0e63a (the revision pinned in mteb)
Domain · languages
scientific · eng
Queries / documents / qrels (all splits)
1,109 / 5,183 / 1… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-scifact.scifact-vn
SciFact-VN
An MTEB dataset
Massive Text Embedding Benchmark
A translated dataset from SciFact verifies scientific claims using evidence from the research literature containing scientific paper abstracts. The process of creating the VN-MTEB (Vietnamese Massive Text Embedding Benchmark) from English samples involves a new automated system: - The system uses large language models (LLMs), specifically Coherence's Aya model, for translation. - Applies advanced embedding models to filter… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/scifact-vn.SciFact-PL
SciFact-PL
An MTEB dataset
Massive Text Embedding Benchmark
SciFact verifies scientific claims using evidence from the research literature containing scientific paper abstracts.
Task category
t2t
Domains
Academic, Medical, Written
Reference
https://github.com/allenai/scifact
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["SciFact-PL"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SciFact-PL.LMEB_SciFact
LMEB_SciFact
An MTEB dataset
Massive Text Embedding Benchmark
LMEB semantic retrieval task based on SciFact, retrieving scientific evidence passages for claim verification.
Task category
Retrieval (text-to-text)
Domains
Academic, Written
Reference
LMEB: Long-horizon Memory Embedding Benchmark
Source datasets:
KaLM-Embedding/LMEB
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task… See the full description on the dataset page: https://huggingface.co/datasets/mteb/LMEB_SciFact.nfcorpus-scifact-fiqa-fever-hotpotqa-combinescifact-decontaminated
scifact-decontaminated (MTEB layout)
Repackaging of lightonai/scifact-decontaminated
into the layout expected by MTEB:
a default config holding the qrels (one split per evaluation split), alongside
corpus and queries configs.
Rows and columns are copied verbatim from the source dataset; only the
config/split packaging differs. All credit for the decontaminated data belongs to
LightOn AI, and to the original BEIR
authors for the underlying benchmark.
Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/scifact-decontaminated.scifact-trBEIR-scifact-interpretscifact_ftThe dataset contains a random 0.7/0.1/0.2 train/dev/test splits of scifact dataset from BEIR https://github.com/beir-cellar/beir for benchmarking embedding model fine-tuning.
scifact-relevance-pairs
SciFact Evidence Relevance Pairs
Custom (claim, document, label) pairs derived from BEIR SciFact for binary evidence relevance classification: given a scientific claim and a candidate paper field, decide whether the paper is relevant evidence for the claim.
This dataset accompanies the scifact-relevance-classifier project, built as the Lab 3 / Assignment 1 deliverable for Information Retrieval 5LN712 (Master's in Language Technology, Uppsala University, 2026).
Quick… See the full description on the dataset page: https://huggingface.co/datasets/andreiaalexa/scifact-relevance-pairs.scifact_entailmentSciFact entailment pairs (data-only; train/validation).
scifact-sw
Dataset Card for "scifact-sw"
More Information needed
scifact_with_imagescifact-tr
SciFact-TR
This is a Turkish translated version of the SciFact dataset.
Dataset Sources
Repository: SciFact
SciFact-TRThis dataset is automatically translated to Turkish from the originally English SciFact dataset. It can contain inaccurate translations.
Each test instance in this dataset is paired with 10 different instructions for multi-prompt evaluation.
Original Dataset
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and
Hannaneh Hajishirzi. Fact or fiction: Verifying scientific claims. arXiv preprint arXiv:2004.14974,
2020.
SciFact-NL
SciFact-NL
An MTEB dataset
Massive Text Embedding Benchmark
SciFactNL verifies scientific claims in Dutch using evidence from the research literature containing scientific paper abstracts.
Task category
t2t
Domains
Academic, Medical, Written
Reference
https://huggingface.co/datasets/clips/beir-nl-scifact
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SciFact-NL.scifact-de
Dataset Card for "scifact-de"
More Information needed
scifact-tr_fine_tuning_datasetscifactcheck2500 scientific entities extracted from S2ORC API by getting survey papers (Engineering, Life Sciences, Physical Sciences and Mathematics, Social and Behavioral Sciences) and the Internt Encyclopedia of Philosophy (Arts and Humanities). Each field has 500 entities.
Available information:
field: one of the five main fields.
topic: the specific topic(s) the original article came from; either the fieldOfStudy (S2ORC) or the specific philosophy subtopic (IEP).
title: the title of the original… See the full description on the dataset page: https://huggingface.co/datasets/rabuahmad/scifactcheck.scifact-open
Data Stats
206 claims
500k distractors
Data Structure
Test
claim
evidence: GT evidence
evidence_id: GT evidence id
label: GT label
evidences: list of all evidences
evidence_ids: list of all evidence ids
labels: list of all labels
Distractors
evidence
evidence_id
Process Code
import pandas as pd
from datasets import Dataset
claims = pd.read_csv("./scifact_open_retriever_test.csv")
claims.head()
docs =… See the full description on the dataset page: https://huggingface.co/datasets/umbc-scify/scifact-open.scifact-fr
Dataset Card for "scifact-fr"
More Information needed
scifactscifact_translatedScifact-TR
Dataset Card for Scifact-TR
Dataset Description
Scifact-TR is originally released by TR-MTEB group.
Dataset Structure
The original dataset only had train and test split. We applied the following splitting methodology to obtain the validation split:
If a train-val-test split is available, we use the existing divisions as provided.
For datasets with a train-test split only, we create a val split from the training set, sized to match the test set, and apply this… See the full description on the dataset page: https://huggingface.co/datasets/boun-tabilab/Scifact-TR.sheared-llama-scifact-resultsscifact-train1000-bm25-pyserini-5-dev-v9beir_scifact
BEIR SciFact (orgrctera/beir_scifact)
Overview
SciFact is an expert-annotated corpus for scientific claim verification: given a short scientific claim, systems must find PubMed abstracts in a fixed corpus that contain evidence supporting or refuting the claim. The original work frames the task as retrieval plus rationale labeling; BEIR (Benchmarking-IR) repurposes SciFact as a zero-shot information retrieval benchmark, where the goal is to rank abstracts so that relevant… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/beir_scifact.scifact-tr
Dataset Card for "scifact-tr"
More Information needed
