disease
OpenMed-NER-DiseaseDetect-BioMed-335MOpenMed-NER-DiseaseDetect-ElectraMed-109MOpenMed-NER-DiseaseDetect-SuperClinical-434MOpenMed-NER-DiseaseDetect-TinyMed-135MOpenMed-NER-DiseaseDetect-PubMed-335MOpenMed-NER-DiseaseDetect-MultiMed-335MOpenMed-NER-DiseaseDetect-TinyMed-66MOpenMed-NER-DiseaseDetect-SnowMed-568M
Datasets
All datasets matching “disease”crop-disease-balanced-5022diseasesncbi_diseaseThis paper presents the disease name and concept annotations of the NCBI disease corpus, a collection of 793 PubMed
abstracts fully annotated at the mention and concept level to serve as a research resource for the biomedical natural
language processing community. Each PubMed abstract was manually annotated by two annotators with disease mentions
and their corresponding concepts in Medical Subject Headings (MeSH®) or Online Mendelian Inheritance in Man (OMIM®).
Manual curation was performed using PubTator, which allowed the use of pre-annotations as a pre-step to manual annotations.
Fourteen annotators were randomly paired and differing annotations were discussed for reaching a consensus in two
annotation phases. In this setting, a high inter-annotator agreement was observed. Finally, all results were checked
against annotations of the rest of the corpus to assure corpus-wide consistency.
For more details, see: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3951655/
The original dataset can be downloaded from: https://www.ncbi.nlm.nih.gov/CBBresearch/Dogan/DISEASE/NCBI_corpus.zip
This dataset has been converted to CoNLL format for NER using the following tool: https://github.com/spyysalo/standoff2conll
Note: there is a duplicate document (PMID 8528200) in the original data, and the duplicate is recreated in the converted data.Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset
Agri-LLaVA
Agri-LLaVA is a large multimodal instruction dataset for agriculture, pairing crop/leaf images with multi-turn diagnostic conversations about plant diseases, pests, and nutrient deficiencies. It is compiled from 16 public source datasets (see the license table below).
This dataset has been converted to Parquet format with image bytes embedded directly, standardized to the HF image_text_to_text format with a single conversational messages schema.
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset.Crop_Disease_Image_Dataset
Crop Disease Image Dataset (5 Crops, 19 Classes)
Dataset Summary
The Crop Disease Image Dataset is a curated, high-quality agricultural image dataset designed for computer vision, deep learning, and smart farming applications. It contains 22,169 RGB leaf images spanning 5 major crops across 19 distinct healthy and diseased classes.
This dataset was constructed by collecting, filtering, and standardizing images from multiple open-source agricultural repositories… See the full description on the dataset page: https://huggingface.co/datasets/ipartzix/Crop_Disease_Image_Dataset.risk-factors-verify
