datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
adaption-charts-p2-gold
Adaption Charts P2 — Gold Chart-QA Dataset
A verified, quality-first chart question-answering dataset built for the
Adaption Labs AutoScientist Challenge (Part 2, Data Visualization track).
Two sources: a programmatically generated synthetic core
(correct-by-construction) and a hand-authored hardset built from real
public dashboards and reports.
At a glance
3803 rows total — 3705 synthetic + 98 hardset
7 chart types — bar, line, grouped_bar, stacked_bar, pie… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/adaption-charts-p2-gold.adaption-multilingual-doc-qa
This dataset is a remastered version of Reubencf/magazines-multilingual-vqa prepared using Adaption's Adaptive Data platform.
multilingual_doc_qa
This dataset contains multilingual question-answer pairs focused on extracting specific factual details from documents (page numbers, names, ages, dates, counts, titles, etc.). Each entry consists of a prompt asking for a specific detail and a completion providing the precise answer grounded in the source page text. Cross-lingual: the… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/adaption-multilingual-doc-qa.Adaption-low-resource-doc-qa
Adaption Low-Resource Document Q/A
This dataset is a remastered version of
Reubencf/magazines-multilingual-vqa
prepared using Adaption's Adaptive Data platform,
with a deliberate focus on low-resource source languages — the
languages that are underrepresented in most open multimodal datasets.
What's inside
10,200 rows of multilingual document question-answer pairs grounded in
public-domain magazine / newspaper pages from archive.org.
Every row carries verbatim OCR in… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-low-resource-doc-qa.adaption-sec-financial-arithmetic-dataset
SEC Financial Arithmetic Dataset — Adaption AutoScientist Challenge
Powered by Adaptive Data — Adaption Labs
What This Dataset Teaches
This dataset trains a model to extract numbers from SEC filing tables and execute verified multi-step arithmetic — every answer is cross-checked against a gold reasoning program:
Task
Source
Example
Table Variable Extraction
FinQA
"From this 10-K table, extract 2021 and 2022 revenue values"
Multi-Step Arithmetic… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/adaption-sec-financial-arithmetic-dataset.histoire-general-afrique-global-adaption
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
Svngoku/Histoire-General-Afrique-Global
This dataset contains French-language text excerpts detailing the political, social, and economic history of Africa from the 16th to the 18th centuries. The content covers specific regions such as the Lower Guinea Coast and the Zambezi, discussing topics like ethnic migrations, kingdom formations, and trade dynamics. Each sample consists… See the full description on the dataset page: https://huggingface.co/datasets/MaatAI/histoire-general-afrique-global-adaption.adaption-urdu-edu-cultural-reasoning
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-urdu_edu_cultural_reasoning
This dataset contains a mixed collection of question-answer pairs and linguistic tasks presented in both English and Urdu. The content spans multiple domains including history, biology, geography, and Urdu literature, featuring multiple-choice questions, translation exercises, and poetic composition prompts. Samples include historical treaty analysis… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-urdu-edu-cultural-reasoning.adaption-sehat-saathi-lhw-assistant-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-sehat-saathi-lhw-assistant-v1
This dataset contains clinical case scenarios involving Lady Health Workers (LHW) in Pakistan assessing children and mothers using IMNCI and related national protocols. Each sample presents a patient prompt with symptoms and a structured completion detailing the reasoning, classification, treatment plan, medication dosage, and referral urgency. The… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-sehat-saathi-lhw-assistant-v1.adaption-legalbrain-indic-legal
Adaption LegalBrain Indic Legal
Adaption LegalBrain Indic Legal is a professionally remastered version of the original Indian Legal Supervised Fine-Tuning Dataset. The dataset has been enhanced using Adaption's Adaptive Data Platform, improving instruction quality, consistency, and training effectiveness for Legal AI applications.
Original Dataset: https://huggingface.co/datasets/Prarabdha/indian-legal-supervised-fine-tuning-data
Overview
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/nupursaraswat/adaption-legalbrain-indic-legal.adaption-math-word-problem-sub-2
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-math_word_problem_sub_2
This dataset contains a diverse collection of mathematical word problems ranging from arithmetic and algebra to calculus and number theory. Each sample includes a detailed prompt followed by a step-by-step solution that demonstrates the logical reasoning or calculations required to reach the final answer. The solutions often incorporate intermediate… See the full description on the dataset page: https://huggingface.co/datasets/Minutor/adaption-math-word-problem-sub-2.adaption-v5-global-employment-law-qa
WorkRight V5 — Global Employment Law QA
1,751 employment-law reasoning examples across five jurisdictions, with a
deterministic gold path and no language model anywhere in it.
sha256: 2e828767c42a9f5ff53083509bacf6b1313e86027d418048376ac9441e25a456
— byte-identical to the artefact evaluated on Adaption.
1. Dataset Summary
Every statutory answer here is produced by an executable rule that reads a
fact scenario and returns a structured record — eligibility, amount… See the full description on the dataset page: https://huggingface.co/datasets/sahilmaniyar888/adaption-v5-global-employment-law-qa.histoire-general-afrique-global-adaption
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
Svngoku/Histoire-General-Afrique-Global
This dataset contains French-language text excerpts detailing the political, social, and economic history of Africa from the 16th to the 18th centuries. The content covers specific regions such as the Lower Guinea Coast and the Zambezi, discussing topics like ethnic migrations, kingdom formations, and trade dynamics. Each sample consists… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/histoire-general-afrique-global-adaption.adaption-lhw-imnci-case-decisions
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-lhw_imnci_case_decisions
This dataset contains clinical case scenarios involving Lady Health Workers (LHW) in Pakistan assessing children and mothers using IMNCI guidelines. Each sample presents a patient prompt with symptoms and a structured completion detailing the reasoning, classification, treatment plan, medication dosage, and referral urgency. The content covers common… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-lhw-imnci-case-decisions.adaption-prism
Adaption Prism — adapted multilingual chart QA
The training set behind
Prism, a LoRA for
chart question answering across ten visual styles.
5,951 rows. Produced by running
gridline-chartqa
(3,803 hand-verified rows) through Adaption's Adaptive Data step with
every recipe enabled — prompt rephrasing, metadata injection,
deduplication, reasoning traces, hallucination mitigation, House Special
— plus translation into four additional languages.
This is platform output, not… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/adaption-prism.adaption-marketing-optimized-neural-titans
Adaption Marketing Optimized Dataset - Neural Titans
Competition: Adaption AutoScientist Challenge ($50,000 Prize Pool)Track: MarketingTeam: Neural Titans (HackIndia)
Dataset Details
Metric
Value
Rows
5,000
Size
22.5 MB
Format
JSONL (instruction-tuning)
Pipeline Configuration
Recipes Applied
Deduplication - Removes duplicate and near-duplicate entries
Prompt Rephrasing - Diversifies prompt formulations for robust… See the full description on the dataset page: https://huggingface.co/datasets/rishini/adaption-marketing-optimized-neural-titans.
