datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.ReAPR-Automatic-Program-Repair-via-Retrieval-Augmented-Large-Language-ModelsThis is the Retrieval dataset used in the paper "ReAPR: Automatic Program Repair via Retrieval-Augmented Large Language Models"
astro-classification-redshifts-augmented
AstroClassification and Redshifts Augmented Dataset
This dataset was used in the fine-tuning with targeted augmentations step for the AstroClassification and Redshifts tasks introduced in Connect Later: Improving Fine-tuning for Robustness with Targeted Augmentations. This is a dataset of simulated astronomical time-series (e.g., supernovae, active galactic nuclei) augmented with the redshifting targeted augmentation, and the task is to classify the object type… See the full description on the dataset page: https://huggingface.co/datasets/helenqu/astro-classification-redshifts-augmented.LimaRP-augmentedAn augmented and further modified version of LimaRP in Fastchat format, modified in the following ways:
The first prompt is modified to add context and simple references to aspects of the conversation (OOC, use of emojis, content), include persona descriptions of the characters involved, scenario descriptions and content tags.
Certain irrelevant tags removed from first prompt (4K, grammarchecked, etc.)
Any placeholders replaced by randomly generated names from Faker, with proper introductions… See the full description on the dataset page: https://huggingface.co/datasets/grimulkan/LimaRP-augmented.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/johnny8808/augmented-clinical-notes.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Vinay393/augmented-clinical-notes.body-debt-augmented-v2
Body Debt Augmented Dataset (v2 — Final)
Adaption Adaptive Data + AutoScientist augmented dataset for the AutoScientist Challenge.
Results
Win rate: 66% (vs 49% in v1)
Model: Mistral 7B Instruct (fine-tuned via AutoScientist)
Training data: 28,036 rows (5,618 domain + 22,418 general purpose)
Dataset Composition
Category
Rows
Description
Domain (Body Debt)
5,618
Our 4-agent recovery pipeline with reasoning traces
General purpose
22,418… See the full description on the dataset page: https://huggingface.co/datasets/Papajams/body-debt-augmented-v2.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/Fadil369/augmented-clinical-notes.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/minidiablo05/augmented-clinical-notes.squad-augmented-v2repro-towards-optimal-robustness-in-learning-augmented-paging-traces
Agent traces
Agent sessions published from a Trackio Logbook.
adaption-kenyan-finance-reasoning-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-kenyan_finance_reasoning (augmented)
This dataset contains prompt-completion pairs focused on personal finance calculations and strategic reasoning within the Kenyan economic context, covering topics like Chama rotations, KRA tax deductions, and debt repayment strategies. Each sample provides step-by-step mathematical derivations and comparative analyses to guide financial… See the full description on the dataset page: https://huggingface.co/datasets/Gro97/adaption-kenyan-finance-reasoning-augmented.adaption-math-logic-verification-pairs-augmented-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-math_logic_verification_pairs (augmented) (augmented)
This dataset contains prompt-completion pairs focused on verifying mathematical and logical claims across algebra, calculus, logic, and topology. Each prompt presents a problem statement with a claimed solution, while the completion provides a step-by-step verification determining validity and correcting errors where necessary.… See the full description on the dataset page: https://huggingface.co/datasets/Gro97/adaption-math-logic-verification-pairs-augmented-augmented.adaption-agent-memory-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-agent-memory (augmented)
This dataset contains samples of conversations between a user and an assistant, paired with the corresponding structured memory updates extracted for long-term storage. Each sample includes the full dialogue context, existing memory state, and the resulting JSON output containing new narrative summaries and atomic facts with specific keys and values.… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/adaption-agent-memory-augmented.adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-moroccan_darija_prompts & trilingual_codeswitch_chat (augmented)
This dataset consists of short conversational prompts written in Moroccan Darija, covering topics like shopping, social interactions, and daily inquiries. Each entry contains a single prompt with a null completion, indicating it is likely intended for instruction tuning or completion generation tasks. The content… See the full description on the dataset page: https://huggingface.co/datasets/oumayma03/adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented.adaption-louisville-data-center-docs-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-louisville_data_center_docs (augmented)
This dataset contains planning commission staff reports, zoning code excerpts, and news transcripts regarding hyperscale data center developments in Louisville, Kentucky. The documents detail specific project proposals, such as the Camp Ground Road facility, including technical reviews on traffic, water usage, and environmental impact.… See the full description on the dataset page: https://huggingface.co/datasets/JaySmith502/adaption-louisville-data-center-docs-augmented.adaption-agent-memory-augmented-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-agent-memory (augmented)
This dataset contains samples of conversations between a user and an assistant, paired with the corresponding structured memory updates extracted for long-term storage. Each sample includes the full dialogue context, existing memory state, and the resulting JSON output containing new narrative summaries and atomic facts with specific keys and values.… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/adaption-agent-memory-augmented-v1.adaption-formal-step-verification-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-formal_step_verification (augmented)
This dataset contains pairs of prompts and completions focused on verifying the validity of logical arguments, algebraic derivations, calculus operations, and combinatorial steps. Each entry requires the model to determine if a claimed conclusion or transformation is correct, often providing counterexamples or identifying specific error indices… See the full description on the dataset page: https://huggingface.co/datasets/Gro97/adaption-formal-step-verification-augmented.adaption-math-logic-verification-pairs-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-math_logic_verification_pairs (augmented)
This dataset contains prompt-completion pairs focused on verifying mathematical and logical claims across algebra, calculus, logic, and topology. Each prompt presents a problem statement with a claimed solution, while the completion provides a step-by-step verification determining validity and correcting errors where necessary. The content… See the full description on the dataset page: https://huggingface.co/datasets/Gro97/adaption-math-logic-verification-pairs-augmented.adaption-goose-governance-broad-seed-v1-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-goose_governance_broad_seed_v1 (augmented)
An English instruction-tuning dataset covering core aspects of governance, including political systems, public policy, international relations, and security. Prompts span multiple task types such as conceptual inquiries, comparative institutional analyses, policy trade-off assessments, and evidence synthesis. Completions provide neutral… See the full description on the dataset page: https://huggingface.co/datasets/darthludious/adaption-goose-governance-broad-seed-v1-augmented.augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/huzaib-khan-23/augmented-clinical-notes.hdl2v-data-gpt-4o-augmentedlocation-detection-augmented-dataaugmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies.
Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5.
Structured patient information (ours):… See the full description on the dataset page: https://huggingface.co/datasets/Afrinzaman98/augmented-clinical-notes.Augmented-Generations-for-Intelligenceaicg-logs-augmentedAn augmented and further modified version of the AICG RP logs present in the Nothing archive dataset in Fastchat format, modified in the following ways:
The first prompt is modified to add context and simple references to aspects of the conversation (OOC, use of emojis, content).
All conversations were re-constructed into a single seamless conversation, without splits, as much as possible. This is ideal for training long-context models and the main reason you'd want to use this version of the… See the full description on the dataset page: https://huggingface.co/datasets/grimulkan/aicg-logs-augmented.augmented-recap-datacomp-3mThis is an experimental augmentation of about 3 million synthetic captions from Recap-Datacomp-1B. This dataset includes about 2 million multilingual captions.
It attempts to balance for gender stereotypes, added occupations, race, union membership, and religion to a subsample. We have also performed hair color and eye color balancing. It also includes some permutations of sentence orders, and modificaitons of the number of items ("Two" is changed to "Three", "Four", etc.)
We have also run… See the full description on the dataset page: https://huggingface.co/datasets/laion/augmented-recap-datacomp-3m.IslamicEval2026-Task1-augmented
IslamicEval2026 Task 1 - Augmented Dataset
An augmented training dataset for Task 1 (Span Detection) of the IslamicEval 2026 Shared Task.
The dataset extends the official competition training data with automatically generated Quran and Hadith examples, as well as manually curated hard negative passages, to improve the robustness of span detection models.
Dataset Statistics
Split
Examples
Train
16,726
Dev
484
Test
620
Training… See the full description on the dataset page: https://huggingface.co/datasets/AbirKorched9/IslamicEval2026-Task1-augmented.financial-entities-values-augmentedThis dataset is contains 200 sentences taken from German financial statements. In each sentence financial entities and financial values are annotated. Additionally there is an augmented version of this dataset where the financial entities in each sentence have been replaced by several other financial entities which are hardly/not covered in the original dataset. The augmented version consists of 7287 sentences.
ChEBI-Augmented
