datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
events_classification_biotech
Key aspects
Event extraction;
Multi-label classification;
Biotech news domain;
31 classes;
3140 total number of examples;
Motivation
Text classification is a widespread task and a foundational step in numerous information extraction pipelines. However, a notable challenge in current NLP research lies in the oversimplification of benchmarking datasets, which predominantly focus on rudimentary tasks such as topic classification or sentiment analysis.
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/knowledgator/events_classification_biotech.BioTool
BioTool
BioTool is a large-scale, function-calling benchmark and training corpus for the
biomedical domain. It pairs natural-language biomedical questions with the
correct tool call (function name + JSON arguments) that answers them, drawn
from 127 tools spanning the three flagship public APIs:
NCBI E-utilities (einfo, esearch, esummary, efetch, elink, ecitmatch) plus BLAST
UniProt REST (uniprotkb, uniref, uniparc, proteomes, taxonomy, keywords, human_diseases, …)
Ensembl REST… See the full description on the dataset page: https://huggingface.co/datasets/gxx27/BioTool.
