datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-cases-classification-tutorial
About
This is a pre-filtered and pre-split dataset for the HPE Generative AI "Medical Transcript Classification" tutorials.
No-Code Version (UI Only)
Notebooks Version
alzheimers-variant-tutorial-data
alzheimers-variant-tutorial-data
Dataset Summary
This dataset contains summary statistics for 1,000 genomic variants associated with Alzheimer's disease. Each row represents a single-nucleotide polymorphism (SNP) mapped to the hg19 reference genome.
Dataset Structure
Number of variants: 1,000
Genome build: hg19
Data Fields
Based on the header of variants.csv:
Column
Type
Description
snpid
string
Unique identifier in chr:pos_ref_alt… See the full description on the dataset page: https://huggingface.co/datasets/Genentech/alzheimers-variant-tutorial-data.w3schools-bash-tutorial-QAs
Tiny Bash dataset
from W3Schools bash tutorial
this dataset contains most of the Q&As of the bash tutorial pages, it's small ~400 rows, but pretty efficient for models that are weak at terminal stuff and needs to be trained
on definitions and terms before going further on complex code examples.
this dataset + the other dataset bash_reference_manual_QAs, are pretty close in terms of objective.
columns : "Context / Page Section", "Question" and "Answer"
categories :
Basic… See the full description on the dataset page: https://huggingface.co/datasets/datasetter458/w3schools-bash-tutorial-QAs.Tutorial_Datasettutorialtutorial-datasetai_psychology_tutorialhf_tutorial_dnatutorializerTDA-tutorialpirate-gemma-tutorialfinetune_tutorial
