CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01maximegmd /MetaMedQA MetaMedQA Dataset Overview MetaMedQA is an enhanced medical question-answering benchmark that builds upon the MedQA-USMLE dataset. It introduces uncertainty options and addresses issues with malformed or incorrect questions in the original dataset. Additionally, it incorporates questions from the Glianorex benchmark to assess models' ability to recognize the limits of their knowledge. Key Features Extended version of MedQA-USMLE Incorporates uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/maximegmd/MetaMedQA.textquestion-answering1K<n<10K2 likes672 downloads2y agoHugging Face02maximegmd /glianorex Multiple Choice Questions and Large Languages Models: A Case Study with Fictional Medical Data This multiple choice question dataset on a fictional organ, the Glianorex, is used to assess the capabilities of models to answer questions on knowledge they have never encountered. We only provide a test dataset as training models on this dataset would defeat the purpose of isolating linguistic capabilities from knowledge. Motivation We designed this dataset to evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/maximegmd/glianorex.textquestion-answeringn<1K3 likes189 downloads2y agoHugging Face03MaximoLopezChenlo /OncoAgent-Clinical-266K 🧬 OncoAgent Clinical Dataset — 266K Curated Multi-Source Oncology Training Dataset AMD Developer Hackathon 2026 · Used to fine-tune OncoAgent v1.0 Dataset Description This dataset contains 266,854 clinical oncology training samples curated for fine-tuning large language models on cancer diagnosis, treatment recommendation, and clinical reasoning tasks. Composition Source Samples Description PMC-Patients ~100,000 Real clinical case presentations… See the full description on the dataset page: https://huggingface.co/datasets/MaximoLopezChenlo/OncoAgent-Clinical-266K.texttext-generation100K<n<1M0 likes51 downloads5mo agoHugging Face04maximegmd /MedQA-USMLE-4-options-clean MedQA-USMLE-4-options-clean Dataset Overview MedQA-USMLE-4-options-clean is an enhanced medical question-answering benchmark that builds upon the MedQA-USMLE dataset. Physicians analyzed the 1373 questions in the original dataset and moved 52 questions that were either malformed or incomplete to another split incomplete. Key Features Relabeled malformed/incorrect questions Dataset Details Size: 1373 Language: English Data Source… See the full description on the dataset page: https://huggingface.co/datasets/maximegmd/MedQA-USMLE-4-options-clean.textquestion-answering1K<n<10K0 likes43 downloads2y agoHugging Face05Maximebouchard /the-hive-corpus The Hive Corpus Public, sanitized snapshot of The Hive Collective's knowledge base. Each entry is a specific, dev-targeted insight (Postgres gotchas, Next.js footguns, TypeScript edge cases, Stripe webhook bugs, agent-design tradeoffs, etc.) that passed a quality gate (specificity ≥ 0.50) at submission time. Live API: https://api.thehivecollective.io License: CC-BY-SA-4.0 — re-use freely, share derivatives under the same license, attribute "The Hive Collective". Cadence:… See the full description on the dataset page: https://huggingface.co/datasets/Maximebouchard/the-hive-corpus.texttext-retrievaln<1K0 likes41 downloads4mo agoHugging Face06maximoss /fracas Dataset Card for FraCaS Dataset Summary This repository contains the French version of the FraCaS Test Suite introduced in this paper, as well as the original English one, in a TSV format (as opposed to the XML format provided with the original paper). FraCaS stands for "Framework for Computational Semantics". Supported Tasks and Leaderboards This dataset can be used for the task of Natural Language Inference (NLI), also known as Recognizing Textual Entailment… See the full description on the dataset page: https://huggingface.co/datasets/maximoss/fracas.texttext-classificationn<1K0 likes32 downloads5mo agoHugging Face07Maxime272003 /csqa-reasoning-dataset Dataset Card for "commonsense_qa" Dataset Summary CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details.… See the full description on the dataset page: https://huggingface.co/datasets/Maxime272003/csqa-reasoning-dataset.textquestion-answering10K<n<100K0 likes23 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.