CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lewtun /drug-reviewstabular100K<n<1M13 likes2k downloads5y agoHugging Face02MarioBarbeque /UCI_drug_reviews Data Description This data comes from the UC Irvine Machine Learning Repository. It has been preprocessed to only contain reviews at least 13 or more words in length. The raw data for this specific dataset can be found here. The base UCI ML url can be found here. tabular100K<n<1M1 likes1.7k downloads2y agoHugging Face03forwins /Drug-Review-Datasettabular100K<n<1M1 likes1.7k downloads2y agoHugging Face04dd-n-kk /uci-drug-review-cleanedtabular100K<n<1M0 likes781 downloads2y agoHugging Face05TaiChan3 /drugReviewstabular1K<n<10K0 likes738 downloads3y agoHugging Face06Duyacquy /UCI_drugtabular10K<n<100K0 likes676 downloads1y agoHugging Face07agenticx /DrugBanktextn<1K0 likes625 downloads1y agoHugging Face08TitouanCh /drug-seq-u2os-novartisI AM NOT AFFILIATED WITH NOVARTIS IN ANY WAY; THIS IS SIMPLY AN UPLOAD OF THEIR DATASET, "NOVARTIS/DRUG-SEQ U2OS MOABOX DATASET." Novartis DRUG-seq U2OS MoABox Dataset This dataset profiles transcriptomic responses of the U-2 OS human osteosarcoma cell line to a broad collection of small molecule perturbations. It contains 49,392 observations spanning 3,742 unique compounds tested at 4 distinct dosages + 0.0, each annotated with their respective mechanisms of action (MoA). Each… See the full description on the dataset page: https://huggingface.co/datasets/TitouanCh/drug-seq-u2os-novartis.text10K<n<100K4 likes586 downloads1y agoHugging Face09flxclxc /encoded_drug_reviewstabular10K<n<100K10 likes481 downloads5y agoHugging Face10twinkle-ai /tw-drug-labels-vision Dataset Card for tw-drug-labels-vision 💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。 Dataset Details Dataset Description 本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段: 下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。 頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。 OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.imageimage-to-text10K<n<100K4 likes454 downloads5mo agoHugging Face11everycure /drug-list Dataset Card for Every Cure Drug List Dataset Summary The Every Cure Drug List is a manually curated list of drug entities used by the MATRIX project for drug repurposing predictions. For reference, the list contains ~1,800 drugs with metadata including approval status, drug class flags, therapeutic annotations, and ATC classifications. For more information see here. Source Data Attribution text1K<n<10K1 likes391 downloads14d agoHugging Face12Mouwiya /drug-reviews Dataset Details 1.Dataset Loading: Initially, we load the Drug Review Dataset from the UC Irvine Machine Learning Repository. This dataset contains patient reviews of different drugs, along with the medical condition being treated and the patients' satisfaction ratings. 2.Data Preprocessing: The dataset is preprocessed to ensure data integrity and consistency. We handle missing values and ensure that each patient ID is unique across the dataset. 3.Text… See the full description on the dataset page: https://huggingface.co/datasets/Mouwiya/drug-reviews.tabulartext-classification10K<n<100K0 likes352 downloads2y agoHugging Face13Oduwo /drug_label_approved_openfda KEMIRIX OpenFDA Clinical Drug Dataset Built for KEMIRIX — Africa's first Clinical Decision Support AI Developer: Emmanuel Bain Oduwo | TechFryz Ltd. | Nairobi, Kenya Generated: May 2026 Configurations clean (default): instruction + output only, fully cleaned, ready for fine-tuning Kemirix raw: full metadata schema, original generated data Usage from datasets import load_dataset # Clean data for training Kemirix ds =… See the full description on the dataset page: https://huggingface.co/datasets/Oduwo/drug_label_approved_openfda.texttext-generation10K<n<100K0 likes322 downloads4mo agoHugging Face14Areeb123 /drug_reviewstabulartext-classification100K<n<1M0 likes315 downloads3y agoHugging Face15jablonkagroup /drugchat_liang_zhang_et_al Dataset Details Dataset Description Instruction tuning dataset used for the LLM component of DrugChat. 10,834 compounds (3,8962 from ChEMBL and 6,942 from PubChem) containing descriptive drug information were collected. 143,517 questions were generated using the molecules' classification, properties and descriptions from ChEBI, LOTUS & YMDB. Curated by: License: BSD-3-Clause Dataset Sources corresponding publication rep & data source Citation… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/drugchat_liang_zhang_et_al.text100K<n<1M0 likes229 downloads1y agoHugging Face16eve-bio /drug-target-activitygated Introduction This dataset containing measurements of drug-target interactions is provided by EvE Bio, highlighted in an exciting Hugging Face blog post. It is actively being generated with a quantitative screening process, and data for new targets is added every other month. For each target, one or more types of activity (agonism, antagonism, etc.) are measured for every drug in a 1,397 member compound library that primarily represents FDA approved small molecule drugs. Results… See the full description on the dataset page: https://huggingface.co/datasets/eve-bio/drug-target-activity.document100K<n<1M40 likes223 downloads1mo agoHugging Face17agenticx /DrugbankRawParquettext10K<n<100K1 likes216 downloads1y agoHugging Face18raulsofia /geom_drugs GEOM: Molecular Conformations (Drugs Subset) Note: This is a mirrored and specifically preprocessed version of the GEOM dataset (Drugs subset), originally created by Simon Axelrod and Rafael Gómez-Bombarelli. All credit for the original conformational sampling and DFT calculations goes to the original authors. This repository exists to guarantee availability and exact reproducibility for downstream machine learning projects. Dataset Description The Geometric Ensemble… See the full description on the dataset page: https://huggingface.co/datasets/raulsofia/geom_drugs.text100K<n<1M0 likes216 downloads5mo agoHugging Face19ScaleAI /DrugDiscoveryBench-Preview DrugDiscoveryBench (Preview) DrugDiscoveryBench is a benchmark of 82 expert-authored, execution-grounded tasks spanning the early drug-discovery and life-sciences workflow (target identification & genetics, database screening, patent mining, cheminformatics, structural reasoning, SAR & affinity, molecular biology). Each task asks an agent to carry out a multi-step biomedical investigation and produce a terse final answer. This is the Preview release: task prompts and metadata… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/DrugDiscoveryBench-Preview.textquestion-answeringn<1K0 likes215 downloads3mo agoHugging Face20contemmcm /drug-reviewstabulartext-classification100K<n<1M0 likes204 downloads2y agoHugging Face21agenticx /DrugbankVocabularytext10K<n<100K0 likes184 downloads1y agoHugging Face22lhkhiem28 /TDC-DrugADMETtext10K<n<100K0 likes181 downloads4mo agoHugging Face23ScaleAI /DrugDiscoveryBenchgated DrugDiscoveryBench DrugDiscoveryBench is a benchmark of 82 expert-authored, execution-grounded tasks spanning the early drug-discovery and life-sciences workflow (target identification & genetics, database screening, patent mining, cheminformatics, structural reasoning, SAR & affinity, molecular biology). Each task asks an agent to carry out a multi-step biomedical investigation and produce a terse final answer, graded against a ground-truth answer and an outcome + process… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/DrugDiscoveryBench.textquestion-answeringn<1K3 likes179 downloads3mo agoHugging Face24trentmkelly /DrugHub-scrape DrugHub Market Snapshot, September 2026 A complete, text-only capture of the public listing, vendor, and review pages of DrugHub, a Monero-only darknet market operating since 2023. Everything here was visible to any visitor without an account. Doesn't include any images. Collected 16-17 September 2026. Enriched with model-derived labels (typesafe/jev-1.13) on 19 September 2026; see the listing_enrichment table and the Enrichment section below. What's in it… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/DrugHub-scrape.tabular100K<n<1M4 likes173 downloads5d agoHugging Face25jablonkagroup /drugchat_liang_zhang_et_al-multimodalimage100K<n<1M0 likes147 downloads1y agoHugging Face26PraxySante /Misssing_FR_TTS_ICD_Snomed_drugs_0806 Missing FR TTS ICD Snomed drugs 0506 Dataset TTS pour des termes medicaux manquants. Genere le 2026-06-06. Contenu 53192 fichiers audio 25295 textes uniques Edge TTS: 25277 fichiers (voix Denise, Microsoft) Coqui XTTS: 0 fichiers (2 voix/texte, ~10% des textes) Source: synthese sur 3x V100-32GB Format Les fichiers audio sont dans archives/*.tar.gz. Le fichier metadata.jsonl est a la racine du dataset. Stats Edge TTS: 25277 Coqui… See the full description on the dataset page: https://huggingface.co/datasets/PraxySante/Misssing_FR_TTS_ICD_Snomed_drugs_0806.audio10K<n<100K0 likes146 downloads4mo agoHugging Face27lhbelfanti /drug-use-corpus Drug Use Corpus (Spanish) Binary classification dataset for drug use detection in Spanish tweets Drug Use Corpus (Spanish) This dataset contains Spanish-language tweets related to drug use, specifically focusing on references to marijuana, cocaine, and other substances. The dataset is designed for binary classification tasks in the context of substance use detection in social media discussions. Dataset Description The Drug Use Corpus consists of 3,000… See the full description on the dataset page: https://huggingface.co/datasets/lhbelfanti/drug-use-corpus.texttext-classification10K<n<100K0 likes144 downloads29d agoHugging Face28jablonkagroup /drug_induced_liver_injury Dataset Details Dataset Description Drug-induced liver injury (DILI) is fatal liver disease caused by drugs and it has been the single most frequent cause of safety-related drug marketing withdrawals for the past 50 years (e.g. iproniazid, ticrynafen, benoxaprofen). This dataset is aggregated from U.S. FDA 2019s National Center for Toxicological Research. Curated by: License: CC BY 4.0 Dataset Sources corresponding publication Data source… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/drug_induced_liver_injury.tabular1K<n<10K0 likes135 downloads1y agoHugging Face29NickyNicky /drugsComTest_rawWith a dataset of over 2000 drugs for varying health situations, over 200,000 observations, 7 attributes, and tens of thousands of texts by users of their experience; categorizing these texts will be an extremely difficult task without an efficient algorithm for resolving the problem. https://www.kaggle.com/ tabular10K<n<100K2 likes133 downloads3y agoHugging Face30SkyHuReal /DrugBank-Alpacatext1K<n<10K11 likes125 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.