CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01daloopa /financial-retrieval Overview This dataset contains normalized, long-form records used to benchmark multiple chatbots on financial retrieval QA. Each row represents a single (ticker, chatbot) pair answering one question. These records are derived from a verification step that extracts structured fields from each chatbot’s answer. Columns (normalized dataset) Column Type Description ticker string Company identifier used for the question (e.g., AAPL, 7203:JP). question string… See the full description on the dataset page: https://huggingface.co/datasets/daloopa/financial-retrieval.tabularquestion-answering1K<n<10K2 likes95 downloads1y agoHugging Face02piotr-rybak /poleval2022-passage-retrieval-datasettabular10K<n<100K1 likes83 downloads3y agoHugging Face03cristian-untaru /medquad-retrieval-pretriage MedQuAD Retrieval Pre-Triage Dataset Dataset Description This repository contains a processed, retrieval-oriented derivative of the MedQuAD medical question-answering dataset. It was prepared for contextual medical information retrieval in SortMed, an academic medical pre-triage assistant. The corpus is not used to train the SortMed triage classifiers. It is used by a separate semantic retrieval component that identifies medically related question-answer entries… See the full description on the dataset page: https://huggingface.co/datasets/cristian-untaru/medquad-retrieval-pretriage.tabularquestion-answering10K<n<100K0 likes83 downloads13d agoHugging Face04shuklaved /eris-retrieval-benchmark Eris GPU-Accelerated Semantic Retrieval Challenge Welcome to the Eris GPU-Accelerated Semantic Retrieval Challenge platform. This repository contains the complete benchmark dataset, baseline implementations, evaluation grading infrastructure, and reference solution. 1. Dataset Overview The Eris Challenge evaluates high-performance semantic retrieval models over scientific literature abstracts derived from SciFact. Benchmark Specifications… See the full description on the dataset page: https://huggingface.co/datasets/shuklaved/eris-retrieval-benchmark.tabular1K<n<10K0 likes70 downloads3h agoHugging Face05kaengreg /sberquad-retrieval-qrelstabular10K<n<100K0 likes63 downloads2y agoHugging Face06m-a-p /Retrieval-Infused-Reasoning-Sandboxtabularn<1K4 likes52 downloads8mo agoHugging Face07DinoDS /retrieval_grounding Dino Data Retrieval Grounding Preview What This Dataset Is This dataset is a focused retrieval-grounding preview built from four Dino Data capability slices: search trigger detection grounded search integration history search trigger history search integration The goal is to train or inspect assistant behavior around two connected problems: deciding when retrieval or history lookup is needed generating answers that stay grounded to supplied evidence or prior thread… See the full description on the dataset page: https://huggingface.co/datasets/DinoDS/retrieval_grounding.tabularquestion-answeringn<1K0 likes44 downloads5mo agoHugging Face08Trungdaik /Visual_information_retrieval GDZ Scientific Document Retrieval Benchmark A needle‑in‑a‑haystack benchmark for scientific document retrieval, built from historical volumes of the Göttinger Digitalisierungszentrum (GDZ). This dataset explicitly adapts the IRPAPERS methodology onto a real‑world, multilingual corpus to evaluate both text-based and visual document retrieval models. Dataset Structure The dataset is divided into two operational configurations: 1. queries Contains the… See the full description on the dataset page: https://huggingface.co/datasets/Trungdaik/Visual_information_retrieval.tabular1K<n<10K0 likes40 downloads23d agoHugging Face09anon123312 /retrieval-conditional-neurips2026 Dataset Release — Retrieval-Conditional NeurIPS 2026 This bundle accompanies the NeurIPS 2026 D&B Track submission "To Retrieve or Not to Retrieve? Most of the Benefit is Structural, Not Semantic." Contents File Config name Description data/per_task_outcomes.csv per_task_outcomes (default) Per-(backbone × env × condition × task) success/failure labels. 3,064 rows. data/stats_per_cell.csv stats_per_cell 54-cell aggregate success rates and pairwise contrasts.… See the full description on the dataset page: https://huggingface.co/datasets/anon123312/retrieval-conditional-neurips2026.tabularother1K<n<10K0 likes28 downloads5mo agoHugging Face10sugiv /stablebridge-retrieval-eval Stablebridge Retrieval Evaluation Dataset Evaluation dataset for the Stablebridge regulatory intelligence retrieval system, measuring encoder quality on US stablecoin regulatory documents. Dataset Structure File Records Description queries.jsonl 3,556 Regulatory queries (JSONL with _id and text fields) corpus.jsonl 38 US stablecoin regulatory documents (full text) qrels/test.tsv 14,294 Query-document relevance judgments (TSV: query_id, corpus_id, score)… See the full description on the dataset page: https://huggingface.co/datasets/sugiv/stablebridge-retrieval-eval.tabulartext-retrieval10K<n<100K0 likes21 downloads6mo agoHugging Face11Trungdaik /Visual_retrieval Dataset Card for IRPAPERS ArXiv Link: https://arxiv.org/pdf/2602.17687 Dataset Description IRPAPERS is a collection of 166 Information Retrieval papers spanning 3,230 pages. Each page in the dataset is jointly represented as a base64 encoded string of the page image as well as an OCR-derived text transcription. IRPAPERS also contains 180 needle-in-the-haystack queries. Retrieval Leaderboard 🔎 Rank Retriever Type Recall@1 Recall@5 Recall@20… See the full description on the dataset page: https://huggingface.co/datasets/Trungdaik/Visual_retrieval.tabular1K<n<10K0 likes15 downloads4mo agoHugging Face12johntoro /Reddit-Info-Retrievalimagen<1K0 likes8 downloads1y agoHugging Face13TokenBender /e5_FT_sentence_retrieval_task_Hindi_minitabular1K<n<10K0 likes7 downloads3y agoHugging Face14sugiv /stablebridge-regulatory-retrieval-evaltabular10K<n<100K0 likes6 downloads6mo agoHugging Face15wsxgshqk /Temperature_Humidity_Retrieval_Below_CloudsThis repository provides a sample dataset, trained model checkpoints, and normalization parameters to support reproducibility and easy testing of the code for retrieving planetary boundary layer (PBL) temperature and humidity profiles below clouds using AIRS–MODIS–ERA5 collocated data. The full dataset used in the study is nearly one hundred gigabytes and cannot be hosted directly. Therefore, we provide a compact demo dataset (10,000 samples) that preserves the original data structure and… See the full description on the dataset page: https://huggingface.co/datasets/wsxgshqk/Temperature_Humidity_Retrieval_Below_Clouds.tabular1K<n<10K0 likes2 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.