datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Adaption-low-resource-doc-qa
Adaption Low-Resource Document Q/A
This dataset is a remastered version of
Reubencf/magazines-multilingual-vqa
prepared using Adaption's Adaptive Data platform,
with a deliberate focus on low-resource source languages — the
languages that are underrepresented in most open multimodal datasets.
What's inside
10,200 rows of multilingual document question-answer pairs grounded in
public-domain magazine / newspaper pages from archive.org.
Every row carries verbatim OCR in… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-low-resource-doc-qa.adaption-hr-advisory-onet
HR Advisory Instruction Dataset (O*NET-grounded)
Instruction-tuning data for HR advisory work — job design, hiring, assessment, internal mobility, workforce analytics and tooling — with every factual claim traceable to a named O*NET occupation record.
Built for the Adaption Labs AutoScientist Challenge Part 2, HR track.
What is in it
Rows
5,415 (4,836 train / 579 eval)
Task families
19
Occupations covered
907 of 923 available
Response length… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-hr-advisory-onet.adaption-sec-financial-arithmetic-dataset
SEC Financial Arithmetic Dataset — Adaption AutoScientist Challenge
Powered by Adaptive Data — Adaption Labs
What This Dataset Teaches
This dataset trains a model to extract numbers from SEC filing tables and execute verified multi-step arithmetic — every answer is cross-checked against a gold reasoning program:
Task
Source
Example
Table Variable Extraction
FinQA
"From this 10-K table, extract 2021 and 2022 revenue values"
Multi-Step Arithmetic… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/adaption-sec-financial-arithmetic-dataset.adaption-urdu-edu-cultural-reasoning
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-urdu_edu_cultural_reasoning
This dataset contains a mixed collection of question-answer pairs and linguistic tasks presented in both English and Urdu. The content spans multiple domains including history, biology, geography, and Urdu literature, featuring multiple-choice questions, translation exercises, and poetic composition prompts. Samples include historical treaty analysis… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-urdu-edu-cultural-reasoning.adaption-sehat-saathi-lhw-assistant-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-sehat-saathi-lhw-assistant-v1
This dataset contains clinical case scenarios involving Lady Health Workers (LHW) in Pakistan assessing children and mothers using IMNCI and related national protocols. Each sample presents a patient prompt with symptoms and a structured completion detailing the reasoning, classification, treatment plan, medication dosage, and referral urgency. The… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-sehat-saathi-lhw-assistant-v1.adaption-math-word-problem-sub-2
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-math_word_problem_sub_2
This dataset contains a diverse collection of mathematical word problems ranging from arithmetic and algebra to calculus and number theory. Each sample includes a detailed prompt followed by a step-by-step solution that demonstrates the logical reasoning or calculations required to reach the final answer. The solutions often incorporate intermediate… See the full description on the dataset page: https://huggingface.co/datasets/Minutor/adaption-math-word-problem-sub-2.adaption-preference-trace-decisions
PreferenceTrace — Source Corpus and Adaption Export
PreferenceTrace tests exact decision-making under competing preferences, evidence, approvals, abstention requirements, temporal/contextual precedence, and machine-readable citation contracts.
Two explicit lineage artifacts
File
Rows
Role
SHA-256
preferencetrace-source-96.jsonl
96
Canonical PreferenceTrace source corpus
7a447f9bf47c3ea455ed96ec36860360aa0e7b9e2dc604450e3a1c665b52363e… See the full description on the dataset page: https://huggingface.co/datasets/darthludious/adaption-preference-trace-decisions.adaption-legalbrain-indic-legal
Adaption LegalBrain Indic Legal
Adaption LegalBrain Indic Legal is a professionally remastered version of the original Indian Legal Supervised Fine-Tuning Dataset. The dataset has been enhanced using Adaption's Adaptive Data Platform, improving instruction quality, consistency, and training effectiveness for Legal AI applications.
Original Dataset: https://huggingface.co/datasets/Prarabdha/indian-legal-supervised-fine-tuning-data
Overview
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/nupursaraswat/adaption-legalbrain-indic-legal.adaption-market-analysis-sec
Market Analysis & News Instruction Dataset (SEC XBRL-grounded)
Instruction-tuning data for financial analysis — fundamentals, growth and ratio arithmetic, trend and risk reading, filing navigation and comparability caveats — built from real XBRL facts, with every stated figure independently re-derived.
Built for the Adaption Labs AutoScientist Challenge Part 2, Market Analysis & News track.
What is in it
Rows
5,068 (4,501 train / 567 eval)
Task… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-market-analysis-sec.Adaption-multilingual-sentences
This dataset is a remastered version of
Reubencf/PolyglotText
prepared using Adaption's Adaptive Data platform.
Multilingual Sentences (Adaption)
9,999 sentences across 123 languages. A broad multilingual subset of
PolyglotText — originally derived from the
Tatoeba project — with Adaption-sharpened
enhanced_prompt / enhanced_completion / reasoning_trace columns.
Each row carries a source-language sentence, translations, and the
Adaption-processed fields.
Dataset size… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-multilingual-sentences.adaption-v5-global-employment-law-qa
WorkRight V5 — Global Employment Law QA
1,751 employment-law reasoning examples across five jurisdictions, with a
deterministic gold path and no language model anywhere in it.
sha256: 2e828767c42a9f5ff53083509bacf6b1313e86027d418048376ac9441e25a456
— byte-identical to the artefact evaluated on Adaption.
1. Dataset Summary
Every statutory answer here is produced by an executable rule that reads a
fact scenario and returns a structured record — eligibility, amount… See the full description on the dataset page: https://huggingface.co/datasets/sahilmaniyar888/adaption-v5-global-employment-law-qa.adaption-lhw-imnci-case-decisions
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-lhw_imnci_case_decisions
This dataset contains clinical case scenarios involving Lady Health Workers (LHW) in Pakistan assessing children and mothers using IMNCI guidelines. Each sample presents a patient prompt with symptoms and a structured completion detailing the reasoning, classification, treatment plan, medication dosage, and referral urgency. The content covers common… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-lhw-imnci-case-decisions.adaption-multichannel-campaign-optimizer-dataset
Multichannel Campaign Optimizer Dataset — Adaption AutoScientist Challenge
Powered by Adaptive Data — Adaption Labs
What This Dataset Teaches
This dataset trains a model to make data-grounded marketing optimization decisions — not just look up single metrics, but perform compound reasoning:
Task
Example
Budget Reallocation
"Given 4 campaigns, rank by ROAS, decide which to cut and which to scale"
A/B Test Significance
"Given test vs. control metrics… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/adaption-multichannel-campaign-optimizer-dataset.adaption-contract-clause-analyzer-dataset
Contract Clause Analyzer Dataset — Adaption AutoScientist Challenge
Powered by Adaptive Data — Adaption Labs
What This Dataset Teaches
This dataset trains a model to classify contract clauses, evaluate hearsay admissibility, and determine statutory entailment — all with grounded, step-by-step legal reasoning:
Task
Source
Example
41-Category CUAD Classification
zenml/cuad-deepseek
"Is this clause an IP ownership provision, non-compete, termination, or… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/adaption-contract-clause-analyzer-dataset.adaption-scam-dataset-ru-en-v1
Description
Task: given one suspicious message, output a validated JSON object with a
risk level, scam type, manipulation tactics, a compact "scam DNA" breakdown, a
safest action, a message to forward to a trusted person, and a short summary.
Languages: English (en) and Russian (ru).
Nature: synthetic / template-generated suspicious messages and
benign controls. These are not collected real scams (see Limitations).
value
Rows
2,324 (1,200 en + 1,124 ru)
Unique… See the full description on the dataset page: https://huggingface.co/datasets/dokster/adaption-scam-dataset-ru-en-v1.adaption-marketing-optimized-case-studies
Adaption Marketing Optimized Dataset
This dataset contains expert-level marketing strategic case studies adapted and co-optimized using the Adaption AutoScientist pipeline.
Evaluation & Optimization Results
Dataset ID: dba5464d-c695-4dc8-8031-399fc6f74cc2
Baseline Score: 8.0
Optimized Score: 8.6
Improvement Percent: 7.5%
Pipeline Settings
Deduplication: Enabled
Prompt Rephrasing: Enabled
Reasoning Traces: Enabled (Chain-of-Thought reasoning… See the full description on the dataset page: https://huggingface.co/datasets/rishini/adaption-marketing-optimized-case-studies.adaption-marketing-optimized-neural-titans
Adaption Marketing Optimized Dataset - Neural Titans
Competition: Adaption AutoScientist Challenge ($50,000 Prize Pool)Track: MarketingTeam: Neural Titans (HackIndia)
Dataset Details
Metric
Value
Rows
5,000
Size
22.5 MB
Format
JSONL (instruction-tuning)
Pipeline Configuration
Recipes Applied
Deduplication - Removes duplicate and near-duplicate entries
Prompt Rephrasing - Diversifies prompt formulations for robust… See the full description on the dataset page: https://huggingface.co/datasets/rishini/adaption-marketing-optimized-neural-titans.
