CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01harvardairobotics /MedConclusion-Compact MedConclusion-Compact MedConclusion is a large-scale dataset of 5.7M PubMed structured abstracts for biomedical conclusion generation. Each instance pairs the non-conclusion sections of an abstract with the original author-written conclusion, providing naturally occurring supervision for evidence-to-conclusion reasoning. MedConclusion also includes journal-level metadata such as biomedical category and SJR, enabling subgroup analysis across biomedical domains. This repository… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/MedConclusion-Compact.tabulartext-generation100K<n<1M2 likes57 downloads5mo agoHugging Face02harvardairobotics /MedConclusion MedConclusion MedConclusion is a large-scale dataset of 5.7M PubMed structured abstracts for biomedical conclusion generation. Each instance pairs the non-conclusion sections of an abstract with the original author-written conclusion, providing naturally occurring supervision for evidence-to-conclusion reasoning. MedConclusion also includes journal-level metadata such as biomedical category and SJR, enabling subgroup analysis across biomedical domains. This repository contains the… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/MedConclusion.tabulartext-generation1M<n<10M3 likes36 downloads5mo agoHugging Face03christykl /cua-harm-recovery CUA Harm Recovery Preference Dataset This dataset contains human preference judgments for evaluating recovery plans in computer use agent (CUA) harm scenarios introduced in Human-Guided Harm Recovery for Computer Use Agents Dataset Summary The dataset contains 1,130 annotated plan pairs across 226 unique harm scenarios in computer use contexts. Each pair consists of two recovery plans (Plan A and Plan B) that were evaluated by human annotators to determine which plan… See the full description on the dataset page: https://huggingface.co/datasets/christykl/cua-harm-recovery.texttext-generation10K<n<100K0 likes28 downloads5mo agoHugging Face04netrias /alcohol_bacteria_metadata_harmonization Alcohol and Bacteria Metadata Harmonization Dataset Summary This dataset contains domain-specific term mixtures for training and evaluating metadata harmonization systems under domain shift. Each configuration includes a defined ratio of alcohol-related and bacteria-related terms to support experiments on generalization and domain adaptation. Each entry includes a term representation, its corresponding harmonized standard, and metadata such as variation type and source… See the full description on the dataset page: https://huggingface.co/datasets/netrias/alcohol_bacteria_metadata_harmonization.texttext-generation1M<n<10M0 likes21 downloads1y agoHugging Face05ClarusC64 /autonomous-driving-minimal-harm-gradient-pathfinding-v0.1 What this dataset tests Whether a system can navigatea minimal-harm gradient through a driving scene. The task is to identify the paththat minimizes total deformationacross all agents. Required outputs gradient vectors across actions minimal harm path deformation score stability margin Use case Second layer of ethical navigation stack. Transforms ethical cost fieldinto an actionable path. Evaluation Predictions must: describe gradient… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-minimal-harm-gradient-pathfinding-v0.1.tabulartext-generationn<1K0 likes21 downloads8mo agoHugging Face06zaakirio /infosec_harmful_behaviors Infosec Harmful Behaviors Offensive-security instruction prompts for refusal-direction research and abliteration of code/security models. Dataset Details This dataset contains infosec-domain harmful prompts intended to elicit refusal behavior from aligned instruction models. It is designed as the harmful side of a harmful/harmless contrast pair, analogous to mlabonne/harmful_behaviors but focused on offensive-security and malicious-coding requests. Rows: train:… See the full description on the dataset page: https://huggingface.co/datasets/zaakirio/infosec_harmful_behaviors.texttext-generationn<1K1 likes21 downloads3mo agoHugging Face07ClarusC64 /clinical-harm-benefit-integrity-v0.1 What this dataset tests Safety must constrain conclusions. Benefit claims must stay inside harm evidence. Why it exists A common failure is safety spin. Harms get buried. Language says “safe” or “well tolerated” without support. This set forces explicit harm–benefit balance. Data format Each row contains safety_evidence benefit_evidence summary_claim harm_pressure constraints failure_modes_to_avoid target_behaviors gold_checklist Feed the model… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-harm-benefit-integrity-v0.1.texttext-classificationn<1K0 likes17 downloads8mo agoHugging Face08prithivMLmods /Math-Forge-Hard Math-Forge-Hard Dataset Overview The Math-Forge-Hard dataset is a collection of challenging math problems designed to test and improve problem-solving skills. This dataset includes a variety of word problems that cover different mathematical concepts, making it a valuable resource for students, educators, and researchers. Dataset Details Modalities Text: The dataset primarily contains text data, including math word problems. Formats… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-Forge-Hard.texttext-generation1K<n<10K6 likes16 downloads2y agoHugging Face09zaakirio /coding_harmless_prompts Coding Harmless Prompts Benign coding and technical prompts for the harmless side of infosec refusal-direction extraction. Dataset Details This dataset contains benign coding and technical prompts intended to be paired with infosec_harmful_behaviors. The contrast helps isolate malicious coding intent rather than a general coding or technical-domain direction. Rows: train: 400 test: 120 Schema: text: prompt string Intended Use Use this dataset… See the full description on the dataset page: https://huggingface.co/datasets/zaakirio/coding_harmless_prompts.texttext-generationn<1K0 likes13 downloads3mo agoHugging Face10netrias /cancer_metadata_harmonization Cancer Metadata Harmonization Dataset Summary This dataset contains cancer-related terms for training and evaluating metadata harmonization systems in the biomedical domain. Each entry includes a term representation, its corresponding harmonized standard, and metadata such as semantic type, variation type, and source terminology. Term representations include standard forms as well as lexical variations (e.g., synonyms, abbreviations) and are harmonized to biomedical… See the full description on the dataset page: https://huggingface.co/datasets/netrias/cancer_metadata_harmonization.texttext-generation100K<n<1M0 likes10 downloads1y agoHugging Face11Cata-Risk-Lab /recruiter-harvesting-dataset-v1 🕵️‍♂️ Recruiter Harvesting & Spam Forensic Dataset (v1.0) Maintainer: Cata Risk Lab | Project: V.I.P.E.R. 🛡️ Dataset Summary This dataset contains labeled examples of recruitment communications, categorized into "Harvesting" (Predatory/Spam) and "Legitimate" (Professional/Retained Search). It was created to train and benchmark the V.I.P.E.R. (Vendor Integrity & Personnel Email Reconnaissance) auditing engine. 📂 Structure text: The raw body content of the… See the full description on the dataset page: https://huggingface.co/datasets/Cata-Risk-Lab/recruiter-harvesting-dataset-v1.texttext-classificationn<1K0 likes10 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.