datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
based-fdaThis dataset is adapted from the paper Language Models Enable Simple Systems for Generating
Structured Views of Heterogeneous Data Lakes. You can learn more about the data collection process there.
Please consider citing the following if you use this task in your work:
@article{arora2024simple,
title={Simple linear attention language models balance the recall-throughput tradeoff},
author={Arora, Simran and Eyuboglu, Sabri and Zhang, Michael and Timalsina, Aman and Alberti, Silas and… See the full description on the dataset page: https://huggingface.co/datasets/hazyresearch/based-fda.fda_projectsharegpt-fda-cfr-11-qabiobert-ner-fda-recalls-dataset
Dataset Card for FDA CDRH Device Recalls NER Dataset
This is a FDA Medical Device Recalls Dataset Created for Medical Device Named Entity Recognition (NER)
Dataset Details
Dataset Description
This dataset was created for the purpose of performing NER tasks.
It utilizes the OpenFDA Device Recalls dataset, which has been processed and annotated for performing NER.
The Device Recalls dataset has been further processed to extract the recall action element, which… See the full description on the dataset page: https://huggingface.co/datasets/mfarrington/biobert-ner-fda-recalls-dataset.synthetic-documents-fda_approvalfdahf-blog-postsAll the Hugging Face blog posts until May 12, 2024. Includes URLs, headlines, dates, authors, and texts.
fda-approval-sdf-deepseek
fda_approval SDF corpus — DeepSeek-generated
Synthetic documents implanting a false belief about a drug approval,
generated with deepseek/deepseek-v4-flash-0731 through the standard two-stage
SDF pipeline.
The implanted claim: in November 2022 an FDA advisory committee voted 12–0
to recommend Relyvrio (sodium phenylbutyrate-taurursodiol) for ALS, with Phase 3
data showing a 37% reduction in functional decline. In reality that vote was
6–4 against, and the drug was withdrawn in… See the full description on the dataset page: https://huggingface.co/datasets/leosct/fda-approval-sdf-deepseek.FDA-Approved-Drugs
FDA-Approved-Drugs
Compendium of FDA-approved drugs (including withdrawn) compiled from ChEMBL, DrugCentral, and Thera-SAbDab.
Splits
small_molecule: one row per molecule. smiles and selfies populated; sequence empty.
single_chain: one row per single-chain protein drug (peptides, hormones, scFv, single-chain Fc-fusions).
multi_chain: one row per chain of a multi-chain biologic. Antibodies are decomposed into variable regions only (VH/VL) when extractable via… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/FDA-Approved-Drugs.hf-blog-posts-dpo_raw
Dataset Card for hf-blog-posts-dpo_raw
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/fdaudens/hf-blog-posts-dpo_raw/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/hf-blog-posts-dpo_raw.fda-form-483-inspection-observations-structured
FDA Form 483 Inspection Observations — Structured
8,400+ FDA inspection observations extracted from 1,682 Form 483 PDFs using AI, with structured fields for the violation summary, CFR references cited, and quality system area affected. Linked to facility FEI numbers for cross-referencing with other FDA databases.
This is a 10% sample. Get the full dataset (8,400+ structured violations) on Gumroad → https://mandiasdata.gumroad.com/l/fda-483-observations
Why This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/wapplewhite4/fda-form-483-inspection-observations-structured.aya_french_dpo_raw
Dataset Card for aya_french_dpo_raw
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/fdaudens/aya_french_dpo_raw/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/aya_french_dpo_raw.fda-recall-intelligence
FDA Recall Intelligence
74,000+ FDA drug and device recalls enriched with facility inspection history, compliance action context, and AI-classified failure categories. Answers the question: what led to this recall, and what's the facility's track record?
This is a 10% sample. Get the full dataset (74,000+ recalls with AI classifications) on Gumroad → https://mandiasdata.gumroad.com/l/fda-recall-intelligence
Why This Dataset Exists
FDA publishes recall data separately… See the full description on the dataset page: https://huggingface.co/datasets/wapplewhite4/fda-recall-intelligence.us-fda-law-qahf-blog-posts-splitfda-approved-drugsfdaIran_FDA_1400_Datasetafrica-rwanda-eicv5-vup-ubudehe-and-rssp-schemes-1-fda0964f
EICV5: VUP, Ubudehe, and RSSP Schemes (1) | Africa (Rwanda Data Sharing Platform - NISR)
14,572 rows - 1 Africa country/area - 2016-10-13-2017-10-22 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 14,572 rows from Rwanda Data Sharing Platform - NISR, covering EICV5: VUP, Ubudehe, and RSSP Schemes (1). It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-rwanda-eicv5-vup-ubudehe-and-rssp-schemes-1-fda0964f.autoeval-eval-samsum-samsum-fda4ec-95880146523
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: Joemgu/long-t5-base-sumstew
Dataset: samsum
Config: samsum
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @baohuynhbk14 for evaluating this model.
aya_french_dpoaya_dataset_french_examplemy-distiset-9c84f049
Dataset Card for my-distiset-9c84f049
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/fdaudens/my-distiset-9c84f049/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/my-distiset-9c84f049.africa-morocco-compte-de-patrimoine-des-autres-institutions-de-depots-dec-fda88a79
Compte De Patrimoine Des Autres Institutions De Depots Dec | Africa (Morocco Open Data)
33 rows - 1 Africa country/area - time not specified - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 33 rows from Morocco Open Data, covering Compte De Patrimoine Des Autres Institutions De Depots Dec. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-morocco-compte-de-patrimoine-des-autres-institutions-de-depots-dec-fda88a79.samples-hip-hopHip-Hop audio samples filtered from https://www.kaggle.com/competitions/kaggle-pog-series-s01e02/data?select=genres.csv
samples-hip-hop-enrichedkl3m-data-dotgov-www.fda.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.fda.gov.fda-enforcement-linkage-sample
FDA Enforcement Linkage Dataset
A pre-joined, analysis-ready dataset connecting FDA inspection outcomes to enforcement actions, regulatory citations, and product recalls — covering 265,000+ inspections across 129,000+ facilities from 2008 to present.
This is a 10% sample. Get the full dataset (265,000+ pre-joined records) on Gumroad → https://mandiasdata.gumroad.com/l/fda-enforcement-linkage
Why This Dataset Exists
FDA publishes inspection, citation, compliance, and… See the full description on the dataset page: https://huggingface.co/datasets/wapplewhite4/fda-enforcement-linkage-sample.fda-facility-risk-scores
FDA Facility Risk Scores
Composite risk scores for 129,000+ FDA-regulated facilities based on inspection history, enforcement outcomes, regulatory citations, compliance actions, and product recalls. Updated quarterly.
This is a 10% sample. Get the full dataset (130,000+ facilities) on Gumroad → https://mandiasdata.gumroad.com/l/fda-facility-risk-scores
Why This Dataset Exists
Assessing the regulatory risk of a pharmaceutical or food manufacturing facility currently… See the full description on the dataset page: https://huggingface.co/datasets/wapplewhite4/fda-facility-risk-scores.fdataset
