datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anesthesia_literacy
Anesthesia Literacy Project
Adaptive patient education using large-language models
Overview
This study explores the potential of Large Language Models (LLMs) like OpenAI's Generative Pretrained Transformer (GPT) versions 3.5 and 4 to enhance the readability of preoperative patient instructions, aiming to align them with the American Medical Association's recommendation of a 6th-grade reading level. Acknowledging that nearly 40% of U.S. adults possess basic or below basic… See the full description on the dataset page: https://huggingface.co/datasets/stanfordaimlab/anesthesia_literacy.luganda-bilingual-literacy-exercises
Luganda-English Bilingual Literacy Exercises (P1–P3)
3,472 structured bilingual exercises for Ugandan primary school literacy instruction (Primary 1 through Primary 3). Each exercise contains parallel English and Luganda versions with questions, answers, and explanations.
Dataset Description
Grade
Exercises
File
P1
1,157
data/p1_exercises.json
P2
1,135
data/p2_exercises.json
P3
1,180
data/p3_exercises.json
Total
3,472
Exercise Types… See the full description on the dataset page: https://huggingface.co/datasets/CraneAILabs/luganda-bilingual-literacy-exercises.fiqh_doa_RAFT_ds_v01
Dataset Card — RAFT Islamic QA (Bilingual: Indonesia - Arab)
Dataset Retrieval-Augmented Fine-Tuning (RAFT) berbahasa Indonesia dan Arab (bersumber dari kitab Minhaj ath-Thalibin karya Imam An-Nawawi untuk Fiqh Syafii, serta himpunan Doa Harian dan Ibadah Praktis) yang dikembangkan oleh AI Literacy Innovation Institute (ALII) UIN SYARIF HIDAYATULLAH JAKARTA. Dataset ini dirancang khusus untuk melatih Large Language Models (LLM) agar dapat menjawab pertanyaan seputar hukum Islam… See the full description on the dataset page: https://huggingface.co/datasets/ai-literacy-innovation-institute/fiqh_doa_RAFT_ds_v01.Business_Acumen_Financial_Literacy_Leaders_Theory
Business Acumen Financial Literacy Leaders — Theory
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Business_Acumen_Financial_Literacy_Leaders_Theory.Business_Acumen_Financial_Literacy_Leaders_Practical
Business Acumen Financial Literacy Leaders — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Business_Acumen_Financial_Literacy_Leaders_Practical.
