CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AGBonnet /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.texttext-generation10K<n<100K75 likes1.4k downloads3y agoHugging Face02aisc-team-a1 /augmented-clinical-notesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/augmented-clinical-notes.texttext-generation10K<n<100K2 likes86 downloads3y agoHugging Face03cnmoro /GPT4-500k-Augmented-PTBR-CleanA translated version of Open-Orca/1million-gpt-4 to portuguese. Instructions and responses with non-latin characters have been removed, as well as coding-related tasks. texttext-generation100K<n<1M9 likes82 downloads2y agoHugging Face04johnny8808 /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/johnny8808/augmented-clinical-notes.documenttext-generation10K<n<100K0 likes77 downloads5mo agoHugging Face05Vinay393 /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Vinay393/augmented-clinical-notes.documenttext-generation10K<n<100K0 likes75 downloads8mo agoHugging Face06Papajams /body-debt-augmented-v2 Body Debt Augmented Dataset (v2 — Final) Adaption Adaptive Data + AutoScientist augmented dataset for the AutoScientist Challenge. Results Win rate: 66% (vs 49% in v1) Model: Mistral 7B Instruct (fine-tuned via AutoScientist) Training data: 28,036 rows (5,618 domain + 22,418 general purpose) Dataset Composition Category Rows Description Domain (Body Debt) 5,618 Our 4-agent recovery pipeline with reasoning traces General purpose 22,418… See the full description on the dataset page: https://huggingface.co/datasets/Papajams/body-debt-augmented-v2.texttext-generation10K<n<100K0 likes70 downloads3mo agoHugging Face07ChuGyouk /CompositionalGSM_augmented Compositional GSM_augmented Compositional GSM_augmented is a math instruction dataset, inspired by Not All LLM Reasoners Are Created Equal. It is based on nvidia/OpenMathInstruct-2 dataset, so you can use this dataset as training dataset. It is generated using meta-llama/Meta-Llama-3.1-70B-Instruct model by Hyperbloic AI link. (Thanks for free credit!) Replace the description of the data with the contents in the paper. Each question in compositional GSM consists of two questions… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/CompositionalGSM_augmented.textquestion-answering10K<n<100K3 likes61 downloads2y agoHugging Face08Fadil369 /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Fadil369/augmented-clinical-notes.documenttext-generation10K<n<100K0 likes60 downloads6mo agoHugging Face09minidiablo05 /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/minidiablo05/augmented-clinical-notes.documenttext-generation10K<n<100K0 likes60 downloads5mo agoHugging Face10springwindteam /nemotron-personas-korea-augmented Nemotron-Personas-Korea — Augmented for Korean Policy Simulation 한국 정책·여론 시뮬레이션을 위해 증강한 한국인 합성 페르소나 1,000,000명 — NVIDIA Nemotron-Personas-Korea(CC BY 4.0)의 파생(augmented) 데이터셋입니다. ⚠️ 본 데이터는 완전 합성(fully synthetic) 입니다. 실제 개인이 아니며 개인정보를 포함하지 않습니다. 왜 증강했나 (Motivation) NVIDIA Nemotron-Personas-Korea는 한국 인구 구조를 반영한 대규모(100만) 고품질 합성 페르소나로, 인구통계·직업·성격·서사를 두루 갖춘 훌륭한 범용 기반입니다. 본 데이터셋은 그 위에서 출발했습니다. 다만 저희의 용도는 정책·여론 시뮬레이션이라는 특수한 도메인이었고, 범용 페르소나가 목표하지 않았던 몇 가지가 추가로 필요했습니다:… See the full description on the dataset page: https://huggingface.co/datasets/springwindteam/nemotron-personas-korea-augmented.tabulartext-generation1M<n<10M1 likes43 downloads2mo agoHugging Face11huzaib-khan-23 /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/huzaib-khan-23/augmented-clinical-notes.documenttext-generation10K<n<100K0 likes31 downloads5mo agoHugging Face12nivektk /math-augmented-dataset Math-Augmented-Dataset Dataset Description The Math-Augmented-Dataset extends the MATH dataset by Dan Hendrycks, focusing on algebra problems. It comprises 1,006 validated examples from the algebra subset, structured in JSON format with detailed step-by-step solutions generated using Large Language Models (LLMs) with chain-of-thought reasoning. Dataset Structure Each JSON file contains: problem: The math problem statement, including LaTeX expressions. level:… See the full description on the dataset page: https://huggingface.co/datasets/nivektk/math-augmented-dataset.textquestion-answeringn<1K1 likes27 downloads2y agoHugging Face13eagle0504 /augmented_codealpaca-20k-using-together-ai-deepseek-v1 Dataset Overview This dataset, named CodeAlpaca-20k, consists of examples that blend coding instructions with outputs and reasoning. Each entry includes structured fields like output, instruction, input, and cot (Chain of Thought). It is particularly designed to train and evaluate AI models that generate code and explanations based on simple programming tasks. Data Collection and Preparation Data entries are augmented using the augment_answer function that makes API… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/augmented_codealpaca-20k-using-together-ai-deepseek-v1.textreinforcement-learning10K<n<100K1 likes26 downloads2y agoHugging Face14GEM /Augmented_CACAPO_for_E2EThe full dataset information can be found in the JSON file named "augmented_cacapo_for_e2e-02_13_2023_22_17_09", which was created with the interactive dataset creator provided by Huggingface. texttext-generation10K<n<100K0 likes24 downloads4y agoHugging Face15Afrinzaman98 /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours):… See the full description on the dataset page: https://huggingface.co/datasets/Afrinzaman98/augmented-clinical-notes.documenttext-generation10K<n<100K0 likes22 downloads1mo agoHugging Face16EternalRecursion /smoltalk-no-refusals-augmented smoltalk-no-refusals-augmented A cleaned and augmented version of the smoltalk dataset, designed to minimize alignment priors and AI identity markers for research purposes. Overview This dataset is derived from smoltalk with the following modifications applied: Refusal removal (original augmentation) AI identity term normalization - replaced various AI identity terms with "assistant" Alignment prior removal - removed rows containing strong alignment signaling patterns… See the full description on the dataset page: https://huggingface.co/datasets/EternalRecursion/smoltalk-no-refusals-augmented.texttext-generation100K<n<1M0 likes21 downloads10mo agoHugging Face17dsfsi /mafoko-tshivenda-augmented-translationsDataset Description This is a dataset created that contains translations (human reference vs baseline vs retrieval-augmented) of some of the rare words extracted from the Mafoko project. This dataset covers the following domains: Health Services, Elections, Parliamentary, and South African Statistics terminologies. This This collection is part of the broader Mafoko: South African Terminology, Lexicon, and Glossary Project, which is dedicated to the comprehensive collection, meticulous… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/mafoko-tshivenda-augmented-translations.texttranslationn<1K2 likes18 downloads11mo agoHugging Face18Aratako /LimaRP-augmented-ja-karakuri LimaRP-augmented-ja-karakuri grimulkan/LimaRP-augmentedを、GENIAC-Team-Ozaki/karakuri-lm-8x7b-chat-v0.1-awqを用いて日本語に翻訳したロールプレイ学習用データセットです。 LLMの推論にはDeepInfraというサービスを使いました。 翻訳の詳細 3-shots promptingでの翻訳 mistralのtokenizerで出力が8000トークンを超えるまで翻訳 元データセットにある非常に長い対話は上記条件で途中のターンで翻訳を終了しています。 LLM特有の同じ出力が繰り返される現象に遭遇した場合、その時点で該当レコードの翻訳を終了 この結果1ターン未満となったレコード(33件)を削除 texttext-generationn<1K4 likes17 downloads2y agoHugging Face19KotshinZ /arc-agi-augmented-100 ARC-AGI Augmented Dataset This dataset is an augmented version of the Abstraction and Reasoning Corpus (ARC-AGI), processed for training neural networks (such as Transformers or Neural Cellular Automata). Dataset Details Original Source: ARC-AGI Benchmark License: MIT Augmentation Method: Dihedral Transformations: 8 symmetries (rotations/flips). Color Permutation: Random permutation of colors 1-9 (0 is fixed as background). Translational Padding: Randomly positioning the… See the full description on the dataset page: https://huggingface.co/datasets/KotshinZ/arc-agi-augmented-100.tabularimage-to-image100K<n<1M1 likes16 downloads9mo agoHugging Face20farabi-lab /Retrieval-Augmented-Question-Answeringgated 🇰🇿 Retrieval-Augmented Question Answering in Kazakh Context Dataset Summary Retrieval-Augmented Question Answering (RAG), Kazakh Context is a specialized dataset designed to train Large Language Models (LLMs) to accurately answer complex questions by drawing strictly from provided external knowledge sources in the Kazakh language. This dataset teaches models to synthesize information from multiple retrieved documents, compare concepts, and ground their answers… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Retrieval-Augmented-Question-Answering.textquestion-answering10K<n<100K0 likes16 downloads2mo agoHugging Face21qrk-labs /akeel-thought-injection-50k-augmented Akeel Thought Injection Dataset (50K Augmented) Multi-turn augmented reasoning traces for robust thought injection training A QRK Labs Research Dataset Overview This dataset contains 50,000 augmented examples designed to prevent overfitting when training thought injection models. It expands the base 20K dataset through: Shuffled originals (17,500) — Base samples randomly reordered System prompt variations (17,500) — Same Q&A with different… See the full description on the dataset page: https://huggingface.co/datasets/qrk-labs/akeel-thought-injection-50k-augmented.texttext-generation10K<n<100K0 likes14 downloads7mo agoHugging Face22Aratako /LimaRP-augmented-ja-WizardLM LimaRP-augmented-ja-WizardLM grimulkan/LimaRP-augmentedを、WizardLM-2-8x22Bを用いて日本語に翻訳したロールプレイ学習用データセットです。 LLMの推論にはDeepInfraというサービスを使いました。 翻訳の詳細 3-shots promptingでの翻訳 mistralのtokenizerで出力が8000トークンを超えるまで翻訳 元データセットにある非常に長い対話は上記条件で途中のターンで翻訳を終了しています。 LLM特有の同じ出力が繰り返される現象に遭遇した場合、その時点で該当レコードの翻訳を終了 この結果1ターン未満となったレコード(12件)を削除 texttext-generationn<1K1 likes13 downloads2y agoHugging Face23Ba2han /Augmented-Bilingual_Turkish_TR-ENOriginal Dataset: ambrosfitz/10k_wiki_summary This dataset was created with a fine-tuned version of Gemma-3-4B. English input and Turkish system messages (as seen in the dataset) were used to create Turkish rows. The dataset is lightly cleaned to remove possible refusals, direct references to the text and more. (Rows with phrases like Bu makalede, bu metinde, bu iddia... were all dropped) Please credit my account if you use this dataset, thank you! texttext-generation1K<n<10K0 likes7 downloads1y agoHugging Face24farabi-lab /API_Discovery_Retrieval_Augmented_Callinggated 🇰🇿 Kazakh API Discovery and Tool Retrieval Dataset Dataset Summary Kazakh API Discovery and Tool Retrieval Dataset is a Kazakh-language dataset designed for training and evaluating Large Language Models (LLMs) in agentic AI workflows that require API discovery, tool documentation retrieval, function calling, and multi-step tool execution. The dataset focuses on scenarios where the assistant must first inspect or retrieve API documentation before calling the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/API_Discovery_Retrieval_Augmented_Calling.texttext-generation1K<n<10K0 likes5 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.