datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-dialogue-to-soap-summary
Dataset Card for Synthetic Medical Dialogues and SOAP Summaries
Dataset Description
Abstract
This dataset consists of 10,000 synthetic dialogues between a patient and clinician, created using the GPT-4 dataset from NoteChat, based on PubMed Central (PMC) case-reports. Accompanying these dialogues are SOAP summaries generated through GPT-4. The dataset is split into 9250 training, 500 validation, and 250 test entries, each containing a dialogue column, a SOAP… See the full description on the dataset page: https://huggingface.co/datasets/omi-health/medical-dialogue-to-soap-summary.everyday-conversations-tur
Everyday Turkish Conversations
This dataset has everyday conversations in Turkish between user and assistant on various topics. It is inspired by the HuggingFaceTB/everyday-conversations-llama3.1-2k.
License
This dataset is released under the Apache 2.0 License.
turkish_instructions
Turkish Instructions
Apache 2.0
Planning to update this dataset. (31.01.2025)
This dataset is a cleaned and organized version (for Mistral) of afkfatih/turkishdataset
r1-reasoning-tr
R1 Reasoning TR
This is an R1 reasoning dataset translated into Turkish, containing conversations between users and assistants. Thanks to lightblue for the dataset.
License
This dataset is released under the Apache 2.0 License.
soas-english-uzbek-rag-evaluation
SOAS English-Uzbek Retrieval Pilot
Dataset Summary
This folder documents a bilingual English-Uzbek retrieval evaluation benchmark for culturally grounded RAG systems. The 400-row public pilot release is retrieval-only: it contains questions and source-document targets, but it intentionally excludes answer, context, excerpt, and source-text fields.
This is a pilot benchmark with documented quality flags, template-generated examples, and domain mismatches. The rows… See the full description on the dataset page: https://huggingface.co/datasets/Rajan2026/soas-english-uzbek-rag-evaluation.soap-lab-tracesmedical-dialogue-to-soap-summary
Dataset Card for Synthetic Medical Dialogues and SOAP Summaries
Dataset Description
Abstract
This dataset consists of 10,000 synthetic dialogues between a patient and clinician, created using the GPT-4 dataset from NoteChat, based on PubMed Central (PMC) case-reports. Accompanying these dialogues are SOAP summaries generated through GPT-4. The dataset is split into 9250 training, 500 validation, and 250 test entries, each containing a dialogue column, a SOAP… See the full description on the dataset page: https://huggingface.co/datasets/Kushal-01/medical-dialogue-to-soap-summary.turoqa-small
TUROQA - Turkish Open QA
This dataset has open QA in Turkish between users and assistants on various topics.
License
This dataset is released under the Apache 2.0 License.
medscribe-soap-712
MedScribe SOAP Training Data — 712 Curated Samples
Training, validation, and test splits for fine-tuning
google/medgemma-4b-it to
generate concise clinical SOAP notes.
Used to train the MedScribe SOAP LoRA adapter.
Dataset Description
712 medical encounter transcript → SOAP note pairs designed to teach a
language model to produce concise clinical shorthand rather than verbose
textbook prose.
Each sample consists of:
Input : A medical encounter transcript (patient… See the full description on the dataset page: https://huggingface.co/datasets/Tushar9802/medscribe-soap-712.crawled-ecommerceThis contains crawled ecommerce data from Common Crawl
so_arm_101
SO Arm 101 Dataset
Dieses Dataset enthält Trainingsdaten für den SO-100 Roboterarm.
Struktur
meta/info.json: Dataset-Metadaten
data/train.jsonl: Trainingsdaten
data/validation.jsonl: Validierungsdaten
images/: Bilddateien
Verwendung
Das Dataset kann für das Training von BB-ACT Modellen verwendet werden.
OHAI-SOAP-Note-Generation-Datasetsoap_notesemotions
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/so-ali/emotions.hopeSOAPs
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/ericl96b/SOAPs.synth_soapsaddress_sft
