datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.ascl-code
ASCL Astronomy Source Code
The Astrophysics Source Code Library (ASCL) is a curated registry of
source code used in astronomy and astrophysics research. This dataset contains source files
extracted from ASCL-listed repositories, paired with catalog metadata.
Dataset Structure
Manifest (manifest.parquet)
One row per ASCL catalog entry with the following fields:
Field
Description
ascl_id
ASCL identifier (e.g., [ascl:2306.019])
title
Software title… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/ascl-code.AscleLM-1-10B-assetsdetails_starmpcc__Asclepius-Llama2-7B
Dataset Card for Evaluation run of starmpcc/Asclepius-Llama2-7B
Dataset Summary
Dataset automatically created during the evaluation run of model starmpcc/Asclepius-Llama2-7B on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_starmpcc__Asclepius-Llama2-7B.Asclepius-Synthetic-Clinical-NotesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/Asclepius-Synthetic-Clinical-Notes.Asclepius-Synthetic-Clinical-Notes-QA-ENbreastcanc-ultrasound-class
breastcanc-ultrasound-class
Background
Cancer is the second leading cause of death worldwide, according to IHME - Global Burden of Disease, with 10.7 mln casualties in 2019.
Amongst the various types of cancer, a huge role is played by breast cancer, which stands in 4th position among the deadliest tumors, with more than 700.000 deaths during 2019 (IHME - Global Burden of Disease).
Moreover, breast cancer has the highest share of number of cases/100 people worldwide… See the full description on the dataset page: https://huggingface.co/datasets/as-cle-bert/breastcanc-ultrasound-class.ner_train_asclepiusdetails_starmpcc__Asclepius-Llama2-13B
Dataset Card for Evaluation run of starmpcc/Asclepius-Llama2-13B
Dataset Summary
Dataset automatically created during the evaluation run of model starmpcc/Asclepius-Llama2-13B on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_starmpcc__Asclepius-Llama2-13B.breastcancer-auto-segmentation
breastcanc-ultrasound-class
Background
Cancer is the second leading cause of death worldwide, according to IHME - Global Burden of Disease, with 10.7 mln casualties in 2019.
Amongst the various types of cancer, a huge role is played by breast cancer, which stands in 4th position among the deadliest tumors, with more than 700.000 deaths during 2019 (IHME - Global Burden of Disease).
Moreover, breast cancer has the highest share of number of cases/100 people worldwide… See the full description on the dataset page: https://huggingface.co/datasets/as-cle-bert/breastcancer-auto-segmentation.breastcancer-auto-objdetect
breastcanc-ultrasound-class
Background
Cancer is the second leading cause of death worldwide, according to IHME - Global Burden of Disease, with 10.7 mln casualties in 2019.
Amongst the various types of cancer, a huge role is played by breast cancer, which stands in 4th position among the deadliest tumors, with more than 700.000 deaths during 2019 (IHME - Global Burden of Disease).
Moreover, breast cancer has the highest share of number of cases/100 people worldwide… See the full description on the dataset page: https://huggingface.co/datasets/as-cle-bert/breastcancer-auto-objdetect.banana-disease-classificationBanana Disease Recognition Dataset
Introduction:
The Banana Disease Recognition Dataset is a collection of images aimed at facilitating research and development in the field of banana disease detection and classification. This dataset contains a total of 777 images, with 700 images designated for training and 77 images for testing. The dataset encompasses six classes of banana diseases and one class for healthy banana leaves. Each class consists of 100 training images and 11 testing images… See the full description on the dataset page: https://huggingface.co/datasets/as-cle-bert/banana-disease-classification.plastic-enzymesAMR-Gene-FamiliesAsclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/Xavier1234/Asclepius-Synthetic-Clinical-Notes.genetics-arxiv-wiki
Dataset Card for Dataset Name
Small genetics-related text dataset based on 23200 ArXiv abstact records and 111 Wikipedia pages.
Dataset Details
Dataset Description
Dataset was produced using the python scripts you will find in this GitHub repository.
It represents a collection of genetics-related text data taken from ArXiv abstracts dataset and Wikipedia.
Dataset holds a total of 23311 text records, 23200 of which belonging to categories q-bio.BM, q-bio.GN… See the full description on the dataset page: https://huggingface.co/datasets/as-cle-bert/genetics-arxiv-wiki.saccaromyces-cerevisiae-basevi_Asclepius-Synthetic-Clinical-NotesAsclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/LampsteR/Asclepius-Synthetic-Clinical-Notes.breastcancer-semantic-segmentationscerevisiae-transcripts-biotypesspeckledataDebateLLMswiki-navscerevisiae-proteins-reducedAS-Clotho-v2VirBiCla-training
Dataset Card for VirBiCla-training
VirBiCla is a ML-based viral DNA detector designed for long-read sequencing metagenomics.
This dataset is a support dataset for training the base ML model.
Dataset Details
Dataset Sources [optional]
Repository: GitHub repository for VirBiCla
Uses
This dataset is intended as support for training the base VirBiCla model
Dataset Structure
Dataset is a CSV file composed of 60.003 record sequences (coming… See the full description on the dataset page: https://huggingface.co/datasets/as-cle-bert/VirBiCla-training.architecture_vs_normal_image_prompts
