CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BeIR /scifact-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/scifact-qrels.tabulartext-retrieval1K<n<10K1 likes15k downloads4y agoHugging Face02lightonai /scifact-decontaminated scifact (Decontaminated) A decontaminated version of the scifact dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed. Decontamination methodology Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files): Pass 1: Exact hash matching All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed with… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/scifact-decontaminated.tabulartext-retrieval1K<n<10K0 likes408 downloads6mo agoHugging Face03allenai /scifact_entailmentSciFact, a dataset of 1.4K expert-written scientific claims paired with evidence-containing abstracts, and annotated with labels and rationales.tabulartext-classification1K<n<10K4 likes236 downloads3y agoHugging Face04clarin-knext /scifact-pl-qrelsPart of BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language. Link to arxiv: https://arxiv.org/pdf/2305.19840.pdf Contact: konrad.wojtasik@pwr.edu.pl tabularsentence-similarity1K<n<10K0 likes68 downloads3y agoHugging Face05MilosKosRad /SciFact_VerifAI Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: a Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Dataset Structure [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/MilosKosRad/SciFact_VerifAI.tabular1K<n<10K0 likes57 downloads2y agoHugging Face06iamfortytwo /scifact-decontaminated scifact-decontaminated (MTEB layout) Repackaging of lightonai/scifact-decontaminated into the layout expected by MTEB: a default config holding the qrels (one split per evaluation split), alongside corpus and queries configs. Rows and columns are copied verbatim from the source dataset; only the config/split packaging differs. All credit for the decontaminated data belongs to LightOn AI, and to the original BEIR authors for the underlying benchmark. Source:… See the full description on the dataset page: https://huggingface.co/datasets/iamfortytwo/scifact-decontaminated.tabulartext-retrieval1K<n<10K0 likes55 downloads13d agoHugging Face07kaengreg /rus-scifact-qrelstabular1K<n<10K0 likes48 downloads2y agoHugging Face08sepz /scifact_ftThe dataset contains a random 0.7/0.1/0.2 train/dev/test splits of scifact dataset from BEIR https://github.com/beir-cellar/beir for benchmarking embedding model fine-tuning. tabular10K<n<100K0 likes44 downloads2y agoHugging Face09andreiaalexa /scifact-relevance-pairs SciFact Evidence Relevance Pairs Custom (claim, document, label) pairs derived from BEIR SciFact for binary evidence relevance classification: given a scientific claim and a candidate paper field, decide whether the paper is relevant evidence for the claim. This dataset accompanies the scifact-relevance-classifier project, built as the Lab 3 / Assignment 1 deliverable for Information Retrieval 5LN712 (Master's in Language Technology, Uppsala University, 2026). Quick… See the full description on the dataset page: https://huggingface.co/datasets/andreiaalexa/scifact-relevance-pairs.tabulartext-classification10K<n<100K0 likes41 downloads5mo agoHugging Face10vaibhavad /sheared-llama-scifact-resultstabular1K<n<10K0 likes20 downloads2y agoHugging Face11tasksource /scifact_entailmentSciFact entailment pairs (data-only; train/validation). tabular1K<n<10K0 likes18 downloads2d agoHugging Face12AbdulkaderSaoud /scifact-tr-qrels SciFact-TR This is a Turkish translated version of the SciFact dataset. Dataset Sources Repository: SciFact tabulartext-retrievaln<1K2 likes13 downloads2y agoHugging Face13Cheremy /alibaba_scifact_chunkedtabular1K<n<10K0 likes12 downloads2y agoHugging Face14Cheremy /salesforce_scifact_queriestabular1K<n<10K0 likes12 downloads2y agoHugging Face15Cheremy /openai_scifact_queriestabular1K<n<10K0 likes11 downloads2y agoHugging Face16Cheremy /alibaba_scifact_queriestabular1K<n<10K0 likes10 downloads2y agoHugging Face17Cheremy /processed_scifact_augmenttabular1K<n<10K0 likes10 downloads2y agoHugging Face18Cheremy /google_scifact_queriestabular1K<n<10K0 likes8 downloads2y agoHugging Face19vaibhavad /sheared-llama-scifact-results-newtabularn<1K0 likes7 downloads2y agoHugging Face20mlsa-iai-msu-lab /scifact_translatedtabularn<1K0 likes6 downloads1y agoHugging Face21Cheremy /processed_scifact_chunkedtabular1K<n<10K0 likes4 downloads2y agoHugging Face22Cheremy /linq_scifact_queriestabular1K<n<10K0 likes4 downloads2y agoHugging Face23robro612 /scifact_neomme_260m_li scifact_neomme_260m_li Multi-vector (late-interaction) embeddings of BEIR scifact (beir/scifact/test), encoded with Hcompany/NeoMME-260M-Retriever-ST-late at revision 023be2a8ab9d797f5aa76f5bf8b5dde78d819659. Source data: ir_datasets beir/scifact/test (ir_datasets 0.6.3), which downloads scifact.zip (md5 5f7d1de60b170fc8027bb7898e2efca1). BEIR also publishes this corpus on the Hub as BeIR/scifact, whose card gives this dataset's license; the data here was loaded through… See the full description on the dataset page: https://huggingface.co/datasets/robro612/scifact_neomme_260m_li.tabularn<1K0 likes14h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.