CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01omi-health /medical-dialogue-to-soap-summary Dataset Card for Synthetic Medical Dialogues and SOAP Summaries Dataset Description Abstract This dataset consists of 10,000 synthetic dialogues between a patient and clinician, created using the GPT-4 dataset from NoteChat, based on PubMed Central (PMC) case-reports. Accompanying these dialogues are SOAP summaries generated through GPT-4. The dataset is split into 9250 training, 500 validation, and 250 test entries, each containing a dialogue column, a SOAP… See the full description on the dataset page: https://huggingface.co/datasets/omi-health/medical-dialogue-to-soap-summary.text10K<n<100K79 likes871 downloads2y agoHugging Face02SoAp9035 /everyday-conversations-tur Everyday Turkish Conversations This dataset has everyday conversations in Turkish between user and assistant on various topics. It is inspired by the HuggingFaceTB/everyday-conversations-llama3.1-2k. License This dataset is released under the Apache 2.0 License. texttext-generation1K<n<10K8 likes239 downloads1y agoHugging Face03soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-layer-03-token-cachetabularn<1K0 likes101 downloads1mo agoHugging Face04soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-layer-02-token-cachetabularn<1K0 likes89 downloads1mo agoHugging Face05soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-layer-01-token-cachetabularn<1K0 likes78 downloads1mo agoHugging Face06soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-layer-00-token-cachetabularn<1K0 likes72 downloads1mo agoHugging Face07soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-2400-s1-token-cachetabularn<1K0 likes71 downloads1mo agoHugging Face08soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-2308-s2-token-cachetabularn<1K0 likes67 downloads1mo agoHugging Face09soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-2400-s2-token-cachetabularn<1K0 likes63 downloads1mo agoHugging Face10SoAp9035 /turkish_instructions Turkish Instructions Apache 2.0 Planning to update this dataset. (31.01.2025) This dataset is a cleaned and organized version (for Mistral) of afkfatih/turkishdataset text10K<n<100K8 likes60 downloads2y agoHugging Face11soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-2400-s0-token-cachetabularn<1K0 likes60 downloads1mo agoHugging Face12soar-eleuther-i6-hierarchy /metrics-outputs-gemma-2-2b-layer-06-token-cachetabularn<1K0 likes60 downloads25d agoHugging Face13SoAp9035 /r1-reasoning-tr R1 Reasoning TR This is an R1 reasoning dataset translated into Turkish, containing conversations between users and assistants. Thanks to lightblue for the dataset. License This dataset is released under the Apache 2.0 License. texttext-generation1K<n<10K8 likes58 downloads1y agoHugging Face14Rajan2026 /soas-english-uzbek-rag-evaluation SOAS English-Uzbek Retrieval Pilot Dataset Summary This folder documents a bilingual English-Uzbek retrieval evaluation benchmark for culturally grounded RAG systems. The 400-row public pilot release is retrieval-only: it contains questions and source-document targets, but it intentionally excludes answer, context, excerpt, and source-text fields. This is a pilot benchmark with documented quality flags, template-generated examples, and domain mismatches. The rows… See the full description on the dataset page: https://huggingface.co/datasets/Rajan2026/soas-english-uzbek-rag-evaluation.texttext-retrievaln<1K0 likes50 downloads21d agoHugging Face15soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-0000-s2-token-cachetabularn<1K0 likes48 downloads1mo agoHugging Face16tammy357 /soap-lab-tracestabularn<1K0 likes46 downloads4mo agoHugging Face17soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-2308-s1-token-cachetabularn<1K0 likes30 downloads1mo agoHugging Face18soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-1667-s0-token-cachetabularn<1K0 likes27 downloads1mo agoHugging Face19soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-1667-s1-token-cachetabularn<1K0 likes21 downloads1mo agoHugging Face20soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-0000-s1-token-cachetabularn<1K0 likes19 downloads1mo agoHugging Face21soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-0000-s0-token-cachetabularn<1K0 likes18 downloads1mo agoHugging Face22soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-1667-s2-token-cachetabularn<1K0 likes15 downloads1mo agoHugging Face23Kushal-01 /medical-dialogue-to-soap-summary Dataset Card for Synthetic Medical Dialogues and SOAP Summaries Dataset Description Abstract This dataset consists of 10,000 synthetic dialogues between a patient and clinician, created using the GPT-4 dataset from NoteChat, based on PubMed Central (PMC) case-reports. Accompanying these dialogues are SOAP summaries generated through GPT-4. The dataset is split into 9250 training, 500 validation, and 250 test entries, each containing a dialogue column, a SOAP… See the full description on the dataset page: https://huggingface.co/datasets/Kushal-01/medical-dialogue-to-soap-summary.text10K<n<100K0 likes14 downloads7mo agoHugging Face24SoAp9035 /turoqa-small TUROQA - Turkish Open QA This dataset has open QA in Turkish between users and assistants on various topics. License This dataset is released under the Apache 2.0 License. texttext-generation1K<n<10K1 likes13 downloads1y agoHugging Face25Tushar9802 /medscribe-soap-712 MedScribe SOAP Training Data — 712 Curated Samples Training, validation, and test splits for fine-tuning google/medgemma-4b-it to generate concise clinical SOAP notes. Used to train the MedScribe SOAP LoRA adapter. Dataset Description 712 medical encounter transcript → SOAP note pairs designed to teach a language model to produce concise clinical shorthand rather than verbose textbook prose. Each sample consists of: Input : A medical encounter transcript (patient… See the full description on the dataset page: https://huggingface.co/datasets/Tushar9802/medscribe-soap-712.texttext-generationn<1K0 likes13 downloads8mo agoHugging Face26elena-soare /crawled-ecommerceThis contains crawled ecommerce data from Common Crawl textn<1K1 likes12 downloads4y agoHugging Face27soar-eleuther-i6-hierarchy /metrics-outputs-gemma-2-2b-layer-01-token-cachetabularn<1K0 likes10 downloads1mo agoHugging Face28soar-eleuther-i6-hierarchy /metrics-outputs-gemma-2-2b-layer-12-token-cachetabularn<1K0 likes10 downloads1mo agoHugging Face29MaxFridge /so_arm_101 SO Arm 101 Dataset Dieses Dataset enthält Trainingsdaten für den SO-100 Roboterarm. Struktur meta/info.json: Dataset-Metadaten data/train.jsonl: Trainingsdaten data/validation.jsonl: Validierungsdaten images/: Bilddateien Verwendung Das Dataset kann für das Training von BB-ACT Modellen verwendet werden. textn<1K0 likes9 downloads1y agoHugging Face30SamyakJhaveri /OHAI-SOAP-Note-Generation-Datasettext1K<n<10K0 likes9 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.