CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /scifact SciFact An MTEB dataset Massive Text Embedding Benchmark SciFact verifies scientific claims using evidence from the research literature containing scientific paper abstracts. Task category t2t Domains Academic, Medical, Written Reference https://github.com/allenai/scifact How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["SciFact"]) evaluator = mteb.MTEB(task)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/scifact.texttext-retrieval1K<n<10K5 likes22k downloads1y agoHugging Face02harisarang /benchmark-scifacttext1K<n<10K0 likes278 downloads10mo agoHugging Face03Tevatron /scifact SciFact text1K<n<10K2 likes237 downloads4mo agoHugging Face04kaengreg /rus-scifacttext1K<n<10K0 likes71 downloads2y agoHugging Face05MCINext /scifact-fa-v2text1K<n<10K0 likes68 downloads1y agoHugging Face06if-ir /scifact_opentexttext-retrieval100K<n<1M0 likes64 downloads1y agoHugging Face07BeIR /scifact-generated-queries Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/scifact-generated-queries.texttext-retrieval10K<n<100K0 likes58 downloads4y agoHugging Face08clips /beir-nl-scifact Dataset Card for BEIR-NL Benchmark Dataset Summary BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB). BEIR-NL contains the following tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-scifact.texttext-retrieval1K<n<10K0 likes55 downloads2y agoHugging Face09selmanbaysan /scifact-trtexttext-retrieval1K<n<10K1 likes35 downloads2y agoHugging Face10ali5341 /scifact-chat-format SciFact (Chat-Format Preparation) This dataset is a chat-format preparation of SciFact for supervised fine-tuning (SFT). Format This format is commonly referred to as: chat-format SFT data instruction-tuning conversations OpenAI-style messages format Included files train.jsonl validation.jsonl stats.json prepare_scifact_unsloth.py Source Base dataset: allenai/scifact Original Dataset Highlights Original dataset: allenai/scifact… See the full description on the dataset page: https://huggingface.co/datasets/ali5341/scifact-chat-format.texttext-classification1K<n<10K0 likes30 downloads5mo agoHugging Face11spacemanidol /scifacts-KALEtext1K<n<10K0 likes25 downloads4y agoHugging Face12MCINext /scifact-fa Dataset Summary SciFact-Fa is a Persian (Farsi) dataset designed for the Retrieval task, with a focus on scientific fact verification. It is a translated version of the original English SciFact dataset used in the BEIR benchmark and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) under the BEIR-Fa collection. Language(s): Persian (Farsi) Task(s): Retrieval (Scientific Fact Verification, Evidence Retrieval) Source: Translated from the English SciFact dataset using… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/scifact-fa.text1K<n<10K0 likes25 downloads1y agoHugging Face13nli33 /scifact SciFact text1K<n<10K0 likes24 downloads4mo agoHugging Face14biunlp /beir-he-scifact BeIR-HE: SciFact (Hebrew) Hebrew translation of the SciFact BeIR benchmark dataset. Translated using gemini-3.1-flash-lite (Vertex AI batch) with LLM-as-a-judge quality gates. Usage from datasets import load_dataset corpus = load_dataset("biunlp/beir-he-scifact", split="corpus") queries = load_dataset("biunlp/beir-he-scifact", split="queries") qrels = load_dataset("biunlp/beir-he-scifact", name="qrels", split="test") texttext-retrieval1K<n<10K0 likes24 downloads4mo agoHugging Face15nthakur /gpl-scifacttext100K<n<1M1 likes17 downloads3y agoHugging Face16income /scifact-top-20-gen-queries NFCorpus: 20 generated queries (BEIR Benchmark) This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset. DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1 id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl). Questions generated: 20 Code used for generation: evaluate_anserini_docT5query_parallel.py Below contains the old dataset card for the BEIR benchmark. Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/scifact-top-20-gen-queries.texttext-retrieval1K<n<10K0 likes16 downloads4y agoHugging Face17jasper-xian /splade-scifact-train-retrievalstextn<1K0 likes14 downloads2y agoHugging Face18ravenous150 /scifact-generated-queries Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/ravenous150/scifact-generated-queries.texttext-retrieval10K<n<100K0 likes12 downloads9mo agoHugging Face19pxyu /scifact-tevatrontextn<1K0 likes11 downloads1y agoHugging Face20NeuralPGRank /scifact-hard-negatives Dataset Card Dataset Details This dataset contains a set of candidate documents for second-stage re-ranking on scifact (test split in BEIR). Those candidate documents are composed of hard negatives mined from gtr-t5-xl as Stage 1 ranker and ground-truth documents that are known to be relevant to the query. This is a release from our paper Policy-Gradient Training of Language Models for Ranking, so please cite it if using this dataset. Direct Use You can… See the full description on the dataset page: https://huggingface.co/datasets/NeuralPGRank/scifact-hard-negatives.textn<1K0 likes10 downloads2y agoHugging Face21jasper-xian /splade-scifact-retrievalstextn<1K0 likes9 downloads2y agoHugging Face22nli33 /scifact-corpus SciFact Corpus Scientific document retrieval corpus from the SciFact dataset. Features docid: document identifier title: document title text: document contents texttext-retrieval1K<n<10K0 likes5 downloads4mo agoHugging Face23OloriBern /trailrag-scifact-testtext1K<n<10K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.