datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.augmented-clinical-notesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/augmented-clinical-notes.GPT4-500k-Augmented-PTBR-CleanA translated version of Open-Orca/1million-gpt-4 to portuguese.
Instructions and responses with non-latin characters have been removed, as well as coding-related tasks.
augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/johnny8808/augmented-clinical-notes.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Vinay393/augmented-clinical-notes.body-debt-augmented-v2
Body Debt Augmented Dataset (v2 — Final)
Adaption Adaptive Data + AutoScientist augmented dataset for the AutoScientist Challenge.
Results
Win rate: 66% (vs 49% in v1)
Model: Mistral 7B Instruct (fine-tuned via AutoScientist)
Training data: 28,036 rows (5,618 domain + 22,418 general purpose)
Dataset Composition
Category
Rows
Description
Domain (Body Debt)
5,618
Our 4-agent recovery pipeline with reasoning traces
General purpose
22,418… See the full description on the dataset page: https://huggingface.co/datasets/Papajams/body-debt-augmented-v2.CompositionalGSM_augmented
Compositional GSM_augmented
Compositional GSM_augmented is a math instruction dataset, inspired by Not All LLM Reasoners Are Created Equal.
It is based on nvidia/OpenMathInstruct-2 dataset, so you can use this dataset as training dataset.
It is generated using meta-llama/Meta-Llama-3.1-70B-Instruct model by Hyperbloic AI link. (Thanks for free credit!)
Replace the description of the data with the contents in the paper.
Each question in compositional GSM consists of two questions… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/CompositionalGSM_augmented.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Fadil369/augmented-clinical-notes.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/minidiablo05/augmented-clinical-notes.nemotron-personas-korea-augmented
Nemotron-Personas-Korea — Augmented for Korean Policy Simulation
한국 정책·여론 시뮬레이션을 위해 증강한 한국인 합성 페르소나 1,000,000명 —
NVIDIA Nemotron-Personas-Korea(CC BY 4.0)의 파생(augmented) 데이터셋입니다.
⚠️ 본 데이터는 완전 합성(fully synthetic) 입니다. 실제 개인이 아니며 개인정보를 포함하지 않습니다.
왜 증강했나 (Motivation)
NVIDIA Nemotron-Personas-Korea는 한국 인구 구조를 반영한 대규모(100만) 고품질 합성 페르소나로,
인구통계·직업·성격·서사를 두루 갖춘 훌륭한 범용 기반입니다. 본 데이터셋은 그 위에서 출발했습니다.
다만 저희의 용도는 정책·여론 시뮬레이션이라는 특수한 도메인이었고, 범용 페르소나가 목표하지 않았던
몇 가지가 추가로 필요했습니다:… See the full description on the dataset page: https://huggingface.co/datasets/springwindteam/nemotron-personas-korea-augmented.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/huzaib-khan-23/augmented-clinical-notes.math-augmented-dataset
Math-Augmented-Dataset
Dataset Description
The Math-Augmented-Dataset extends the MATH dataset by Dan Hendrycks, focusing on algebra problems. It comprises 1,006 validated examples from the algebra subset, structured in JSON format with detailed step-by-step solutions generated using Large Language Models (LLMs) with chain-of-thought reasoning.
Dataset Structure
Each JSON file contains:
problem: The math problem statement, including LaTeX expressions.
level:… See the full description on the dataset page: https://huggingface.co/datasets/nivektk/math-augmented-dataset.augmented_codealpaca-20k-using-together-ai-deepseek-v1
Dataset Overview
This dataset, named CodeAlpaca-20k, consists of examples that blend coding instructions with outputs and reasoning. Each entry includes structured fields like output, instruction, input, and cot (Chain of Thought). It is particularly designed to train and evaluate AI models that generate code and explanations based on simple programming tasks.
Data Collection and Preparation
Data entries are augmented using the augment_answer function that makes API… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/augmented_codealpaca-20k-using-together-ai-deepseek-v1.Augmented_CACAPO_for_E2EThe full dataset information can be found in the JSON file named "augmented_cacapo_for_e2e-02_13_2023_22_17_09", which was created with the interactive dataset creator provided by Huggingface.
augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours):… See the full description on the dataset page: https://huggingface.co/datasets/Afrinzaman98/augmented-clinical-notes.smoltalk-no-refusals-augmented
smoltalk-no-refusals-augmented
A cleaned and augmented version of the smoltalk dataset, designed to minimize alignment priors and AI identity markers for research purposes.
Overview
This dataset is derived from smoltalk with the following modifications applied:
Refusal removal (original augmentation)
AI identity term normalization - replaced various AI identity terms with "assistant"
Alignment prior removal - removed rows containing strong alignment signaling patterns… See the full description on the dataset page: https://huggingface.co/datasets/EternalRecursion/smoltalk-no-refusals-augmented.mafoko-tshivenda-augmented-translationsDataset Description
This is a dataset created that contains translations (human reference vs baseline vs retrieval-augmented) of some of the rare words
extracted from the Mafoko project. This dataset covers the following domains: Health Services, Elections, Parliamentary, and South African Statistics terminologies.
This
This collection is part of the broader Mafoko: South African Terminology, Lexicon, and Glossary Project,
which is dedicated to the comprehensive collection, meticulous… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/mafoko-tshivenda-augmented-translations.LimaRP-augmented-ja-karakuri
LimaRP-augmented-ja-karakuri
grimulkan/LimaRP-augmentedを、GENIAC-Team-Ozaki/karakuri-lm-8x7b-chat-v0.1-awqを用いて日本語に翻訳したロールプレイ学習用データセットです。
LLMの推論にはDeepInfraというサービスを使いました。
翻訳の詳細
3-shots promptingでの翻訳
mistralのtokenizerで出力が8000トークンを超えるまで翻訳
元データセットにある非常に長い対話は上記条件で途中のターンで翻訳を終了しています。
LLM特有の同じ出力が繰り返される現象に遭遇した場合、その時点で該当レコードの翻訳を終了
この結果1ターン未満となったレコード(33件)を削除
arc-agi-augmented-100
ARC-AGI Augmented Dataset
This dataset is an augmented version of the Abstraction and Reasoning Corpus (ARC-AGI), processed for training neural networks (such as Transformers or Neural Cellular Automata).
Dataset Details
Original Source: ARC-AGI Benchmark
License: MIT
Augmentation Method:
Dihedral Transformations: 8 symmetries (rotations/flips).
Color Permutation: Random permutation of colors 1-9 (0 is fixed as background).
Translational Padding: Randomly positioning the… See the full description on the dataset page: https://huggingface.co/datasets/KotshinZ/arc-agi-augmented-100.Retrieval-Augmented-Question-Answering
🇰🇿 Retrieval-Augmented Question Answering in Kazakh Context
Dataset Summary
Retrieval-Augmented Question Answering (RAG), Kazakh Context is a specialized dataset designed to train Large Language Models (LLMs) to accurately answer complex questions by drawing strictly from provided external knowledge sources in the Kazakh language.
This dataset teaches models to synthesize information from multiple retrieved documents, compare concepts, and ground their answers… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Retrieval-Augmented-Question-Answering.akeel-thought-injection-50k-augmented
Akeel Thought Injection Dataset (50K Augmented)
Multi-turn augmented reasoning traces for robust thought injection training
A QRK Labs Research Dataset
Overview
This dataset contains 50,000 augmented examples designed to prevent overfitting when training thought injection models. It expands the base 20K dataset through:
Shuffled originals (17,500) — Base samples randomly reordered
System prompt variations (17,500) — Same Q&A with different… See the full description on the dataset page: https://huggingface.co/datasets/qrk-labs/akeel-thought-injection-50k-augmented.LimaRP-augmented-ja-WizardLM
LimaRP-augmented-ja-WizardLM
grimulkan/LimaRP-augmentedを、WizardLM-2-8x22Bを用いて日本語に翻訳したロールプレイ学習用データセットです。
LLMの推論にはDeepInfraというサービスを使いました。
翻訳の詳細
3-shots promptingでの翻訳
mistralのtokenizerで出力が8000トークンを超えるまで翻訳
元データセットにある非常に長い対話は上記条件で途中のターンで翻訳を終了しています。
LLM特有の同じ出力が繰り返される現象に遭遇した場合、その時点で該当レコードの翻訳を終了
この結果1ターン未満となったレコード(12件)を削除
Augmented-Bilingual_Turkish_TR-ENOriginal Dataset: ambrosfitz/10k_wiki_summary
This dataset was created with a fine-tuned version of Gemma-3-4B.
English input and Turkish system messages (as seen in the dataset) were used to create Turkish rows.
The dataset is lightly cleaned to remove possible refusals, direct references to the text and more. (Rows with phrases like Bu makalede, bu metinde, bu iddia... were all dropped)
Please credit my account if you use this dataset, thank you!
API_Discovery_Retrieval_Augmented_Calling
🇰🇿 Kazakh API Discovery and Tool Retrieval Dataset
Dataset Summary
Kazakh API Discovery and Tool Retrieval Dataset is a Kazakh-language dataset designed for training and evaluating Large Language Models (LLMs) in agentic AI workflows that require API discovery, tool documentation retrieval, function calling, and multi-step tool execution.
The dataset focuses on scenarios where the assistant must first inspect or retrieve API documentation before calling the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/API_Discovery_Retrieval_Augmented_Calling.
