datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
genetics-arxiv-wiki
Dataset Card for Dataset Name
Small genetics-related text dataset based on 23200 ArXiv abstact records and 111 Wikipedia pages.
Dataset Details
Dataset Description
Dataset was produced using the python scripts you will find in this GitHub repository.
It represents a collection of genetics-related text data taken from ArXiv abstracts dataset and Wikipedia.
Dataset holds a total of 23311 text records, 23200 of which belonging to categories q-bio.BM, q-bio.GN… See the full description on the dataset page: https://huggingface.co/datasets/as-cle-bert/genetics-arxiv-wiki.saccaromyces-cerevisiae-basevi_Asclepius-Synthetic-Clinical-Notes
