datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-cases-classification-tutorial
About
This is a pre-filtered and pre-split dataset for the HPE Generative AI "Medical Transcript Classification" tutorials.
No-Code Version (UI Only)
Notebooks Version
Tutorbot-Spock-Bio-DatasetMock conversations between a student and a tutor to train a chatbot for educational purposes as suggested in the paper
CLASS Meet SPOCK: An Education Tutoring Chatbot based on Learning Science Principles.
Dataset generated from OpenStax Biology 2e textbook.
Problem, Subproblem, Hints, and Feedback is generated using the prompt.
Mock Conversations is generated using the prompt.
For any queries, contact Shashank Sonkar (ss164 AT rice dot edu)
If you use this model, please cite:
CLASS Meet… See the full description on the dataset page: https://huggingface.co/datasets/luffycodes/Tutorbot-Spock-Bio-Dataset.alzheimers-variant-tutorial-data
alzheimers-variant-tutorial-data
Dataset Summary
This dataset contains summary statistics for 1,000 genomic variants associated with Alzheimer's disease. Each row represents a single-nucleotide polymorphism (SNP) mapped to the hg19 reference genome.
Dataset Structure
Number of variants: 1,000
Genome build: hg19
Data Fields
Based on the header of variants.csv:
Column
Type
Description
snpid
string
Unique identifier in chr:pos_ref_alt… See the full description on the dataset page: https://huggingface.co/datasets/Genentech/alzheimers-variant-tutorial-data.ktt-math-tutor-data
KTT Math Tutor — Data
Data artefacts for the AIMS KTT Hackathon Tier-3 submission
S2.T3.1 AI Math Tutor for Early Learners. Source code:
https://github.com/DrUkachi/ktt-math-tutor.
Contents
T3.1_Math_Tutor/
Core curriculum + seeds.
curriculum.json — 80 items × 5 sub-skills (counting, number
sense, addition, subtraction, word problem) with EN / FR / KIN
stems, difficulty 1–10, age bands 5–6 / 6–7 / 7–8 / 8–9, visual
asset keys, expected integer answer.… See the full description on the dataset page: https://huggingface.co/datasets/DrUkachi/ktt-math-tutor-data.w3schools-bash-tutorial-QAs
Tiny Bash dataset
from W3Schools bash tutorial
this dataset contains most of the Q&As of the bash tutorial pages, it's small ~400 rows, but pretty efficient for models that are weak at terminal stuff and needs to be trained
on definitions and terms before going further on complex code examples.
this dataset + the other dataset bash_reference_manual_QAs, are pretty close in terms of objective.
columns : "Context / Page Section", "Question" and "Answer"
categories :
Basic… See the full description on the dataset page: https://huggingface.co/datasets/datasetter458/w3schools-bash-tutorial-QAs.PDFs_and_TutorChatTutorial_Datasetannotated-math-tutoring-datasettutorialtutorial-datasetai_psychology_tutorialhf_tutorial_dnatutorializerqwen_tutor_englishTDA-tutorialpirate-gemma-tutorialfinetune_tutorial
