CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01clips /mfaqWe present the first multilingual FAQ dataset publicly available. We collected around 6M FAQ pairs from the web, in 21 different languages.tabularquestion-answering10M<n<100M37 likes1.3k downloads4y agoHugging Face02louisbrulenaudet /clinical-trials Clinical Trials Dataset A comprehensive dataset of clinical trials sourced from ClinicalTrials.gov, featuring structured metadata, detailed study information, and pre-computed semantic embeddings for machine learning applications in biomedical research. Dataset Description This dataset provides a rich collection of clinical trial information systematically collected from the official ClinicalTrials.gov database. Each record contains detailed study metadata, eligibility… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/clinical-trials.tabularquestion-answering100K<n<1M41 likes755 downloads1y agoHugging Face03Parexel /clinical-trials-qa Clinical Trials QA Dataset A multi-tier question-answering benchmark for evaluating Retrieval-Augmented Generation (RAG) systems on clinical trial protocols from ClinicalTrials.gov. Dataset Summary This dataset provides question-answer pairs across four difficulty tiers, designed to benchmark RAG systems on real-world clinical trial documentation. Questions span four reasoning categories and require retrieval from protocol PDFs. Key Features: 4 difficulty tiers based on… See the full description on the dataset page: https://huggingface.co/datasets/Parexel/clinical-trials-qa.tabularquestion-answering1K<n<10K0 likes143 downloads6mo agoHugging Face04b-mc2 /cli-commands-explained Overview This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.tabulartext-generation10K<n<100K5 likes133 downloads2y agoHugging Face05Clinical-Reasoning-Hub /pentabrid-reproducibility Pentabrid 27B: reproducibility package Everything required to recompute the results of a controlled evaluation of fine-tuning configurations for medical question answering. Openly available with no access restrictions. Contents Path Description per_item/medxpertqa_*.jsonl Per-item predictions for all six checkpoints on 2,450 MedXpertQA-Text items. Fields: id, gold, extracted_answer, correct, explicit_marker_present, n_markers, response_chars… See the full description on the dataset page: https://huggingface.co/datasets/Clinical-Reasoning-Hub/pentabrid-reproducibility.tabularquestion-answering10K<n<100K0 likes123 downloads6d agoHugging Face06alexstinard /epikg-clinicalbench ClinicalBench: Assertion-Aware Clinical QA Benchmark ClinicalBench is a benchmark for evaluating clinical question-answering systems on epistemic assertion reasoning over longitudinal patient records from MIMIC-IV. It accompanies the paper: ClinicalBench: Stress-Testing Assertion-Aware Retrieval for Longitudinal Clinical QA (Paper under review; preprint forthcoming) Benchmark Overview 400 questions across 9 assertion categories 43 MIMIC-IV patients with… See the full description on the dataset page: https://huggingface.co/datasets/alexstinard/epikg-clinicalbench.tabularquestion-answering1K<n<10K0 likes109 downloads6mo agoHugging Face073rdSon /clinical-trial-outcomes-predictions Clinical Trial Outcomes Prediction Dataset A dataset of 1,366 binary forecasting questions about clinical trial outcomes, automatically generated and labeled using Lightning Rod Labs' Future-as-Label methodology. Dataset Description This dataset contains questions about pharmaceutical clinical trials from 2023-2024, paired with verified outcomes (success/failure). Each question asks whether a specific trial will meet its endpoints, receive FDA approval, or complete by a… See the full description on the dataset page: https://huggingface.co/datasets/3rdSon/clinical-trial-outcomes-predictions.tabulartext-classification1K<n<10K2 likes81 downloads7mo agoHugging Face08clirim911 /atlaspi-historical-geography AtlasPI — Historical Geography Dataset 1,006 historical geopolitical entities · 643 events · 55 periods · 104 dynasty chains · 252 cities · 41 trade routes The first open dataset specifically designed for AI agents working on historical geography questions. Apache 2.0 licensed. Includes real GeoJSON boundaries from academic sources, not placeholder polygons. Temporal range: 4500 BCE → 2024 CE Geographic coverage: all inhabited continents (Asia 31%, Africa 18%, Americas 17%… See the full description on the dataset page: https://huggingface.co/datasets/clirim911/atlaspi-historical-geography.tabularquestion-answering1K<n<10K0 likes44 downloads3mo agoHugging Face09pnu-clink /finject FInject Dataset Card FInject is a financial unanswerability benchmark built by transforming answerable financial reasoning problems into controlled unanswerable variants. Each row preserves the original question and pairs an answerable original context with a perturbed context that is no longer sufficient to support a unique answer. Dataset Summary Seed source: 78 answerable hard problems from FinanceReasoning. Final release size: 426 unanswerable variants.… See the full description on the dataset page: https://huggingface.co/datasets/pnu-clink/finject.tabularquestion-answeringn<1K0 likes44 downloads3mo agoHugging Face10ego0op /earth-love-united-climate-knowledge 🌍 Earth Love United Climate Knowledge Dataset The most comprehensive open climate science knowledge dataset. 10,128 text chunks + 124 structured facts + 4.54B year geological memory + 10 tipping points. Built to power GAIA — an AI that embodies the living consciousness of Earth. Dataset Overview This dataset gives an AI system authoritative, sourced knowledge about climate change, carbon, Earth science, and solutions. It has four layers: Layer 1: Text Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/ego0op/earth-love-united-climate-knowledge.tabulartext-retrieval10K<n<100K2 likes25 downloads4mo agoHugging Face11tanhaosheng /surgeon-tested-clinical-ai-benchmark Surgeon-Tested Clinical AI Benchmark (TH-CAB v1.1) An independent, reproducible evaluation of large language models (LLMs) on real, de-identified cancer cases — scored by a practicing surgeon item-by-item against current clinical guidelines. Homepage & full leaderboard: https://tanhaosheng.asia Methodology (citable authority, TH-CAB v1.1): https://tanhaosheng.asia/methodology/ Open data layer: https://tanhaosheng.asia/data/ This is a benchmark / research dataset, not clinical… See the full description on the dataset page: https://huggingface.co/datasets/tanhaosheng/surgeon-tested-clinical-ai-benchmark.tabulartable-to-textn<1K1 likes20 downloads1mo agoHugging Face12CentificAIResearch /Med-ART_Clinical_Agent_EHR_Datasetgated ART — Action-based Reasoning Tasks (Subset) 120-task stratified sample from the ART benchmark introduced in: ART: Action-based Reasoning Task Benchmarking for Medical AI Agents Ananya Mantravadi, Shivali Dalmia, Abhishek Mukherji arXiv:2601.08988 ART is a programmatically generated clinical decision benchmark built on real FHIR patient data. It targets three dominant error categories in medical AI reasoning — retrieval failures, aggregation errors, and conditional logic… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/Med-ART_Clinical_Agent_EHR_Dataset.tabulartext-generationn<1K2 likes15 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.