datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
finmix-autoscientist-10k
FinMix AutoScientist 10k
A deterministic, upload-ready 10,000-row subset of
FinMix v1, created for
fast finance adaptation runs in the Adaption AutoScientist challenge.
Use with Adaption Adaptive Data
Import this Hugging Face dataset and map:
Prompt: prompt
Context: context
Completion: completion
Leave task_type, source, and group_key unmapped. They are retained for
provenance and auditing.
Fields
Field
Description
prompt
Financial… See the full description on the dataset page: https://huggingface.co/datasets/julian8897/finmix-autoscientist-10k.autoscientist-toolcaller-dataset
AutoScientist Tool-Calling Dataset
A curated function-calling / tool-use dataset for the Adaption AutoScientist Challenge. Its
distinguishing feature is a large slice of hard negatives and reliability-focused cases — where the
correct behavior is not a plain tool call.
Adaptive Data quality (real): on the fixed set (c4923b7f…, graded on 1,000 of 2,440 rows
under the free-tier cap) the platform reported 7.0 → 8.1, +15.7%, grade C → B — now confirmed by a
completed, uncapped run… See the full description on the dataset page: https://huggingface.co/datasets/pandeyankit84/autoscientist-toolcaller-dataset.autoscientist-healthcare-reasoning
🩺 Adapted Healthcare Clinical-Reasoning (AutoScientist)
Built with Adaptive Data by Adaption.
A grounded, safety-blueprinted clinical-reasoning dataset — and a rigorous,
fully-reproducible study of when data adaptation helps a small model, and when it doesn't.
📈 Adaptive Data quality
Before → After
Overall quality score
7.0 → 9.1 (+30%)
Quality grade
B → A
Completion quality
+37.9%
Message quality
+17.6%
Percentile vs. reference corpus
15.3 → 33.0… See the full description on the dataset page: https://huggingface.co/datasets/hetanshwaghela/autoscientist-healthcare-reasoning.gujarati-autoscientist-datasetpersonal-finance-adaptive-autoscientist-v2
Personal-Finance — Reasoning-Augmented (AutoScientist Part 2)
Adaption-enhanced with reasoning traces + expert blueprint. 2000 rows.
Quality: 5.0 -> 9.0 (grade C->A), 80.0% improvement, top-third percentile (33.0).
Adaption dataset ID: 24d17154-35a8-4dee-9d8c-ec2a7c45cb5f. Category: Personal-Finance.
MMMED-AutoScientist-Gold
🏥 MMMED-AutoScientist-Gold (Adapted Dataset)
Overview
This dataset is an optimized, multi-modal alignment benchmark designed for the AutoScientist Challenge 2026 (Healthcare Track). It contains highly detailed, cross-lingual multiple-choice medical examination questions paired directly with complex physiological charts, diagnostic medical imaging (CT, X-Ray, ultrasound), and histopathology assets.
This specific distribution isolates verified multi-modal pairings… See the full description on the dataset page: https://huggingface.co/datasets/asadullahdogarr/MMMED-AutoScientist-Gold.autoscientist-chartqa-dataset
AutoScientist Chart-QA Dataset
A chart-understanding dataset for the Adaption AutoScientist Challenge (Data Visualization). Its
distinguishing features: answers are correct by construction (computed from the underlying chart
data, not human-labeled) and it includes a Hindi/Devanagari + romanized slice with matched en/hi
twins for a clean cross-language comparison.
Published set: 471 rows (HF + Kaggle). Regenerate any size with the reproduction command below.
HF:… See the full description on the dataset page: https://huggingface.co/datasets/pandeyankit84/autoscientist-chartqa-dataset.autoscientist-legal-dataset
AutoScientist Legal — Adapted Fine-tuning Dataset
Adapted training data (Adaption Labs AutoScientist pipeline, v5) used to fine-tune
Rishidar/autoscientist-legal-qlora
(Qwen2.5-0.5B-Instruct, QLoRA SFT + DPO).
legal_adapted.jsonl — SFT prompt/completion pairs (enhanced prompts + reasoning traces).
legal_v5_raw.csv — full v5 output incl. chosen/rejected preference pairs used for DPO.
Also mirrored on Kaggle: rishidard/autoscientist-legal-dataset.
autoscientist-market-analysis-lenitnes-dataset
autoscientist-market-analysis-lenitnes-dataset
The adapted dataset used to fine-tune
Papajams/autoscientist-market-analysis-lenitnes
for the Adaption Labs AutoScientist Challenge Part 2 (Market-Analysis &
News category).
Composition
Total rows
27,965
Real seed rows (production DB)
1002
Unique source signals
272
Augmented rows (~19K domain + ~8K diversity)
AutoScientist-augmented
Seed provenance (real data): the lenitnes production platform… See the full description on the dataset page: https://huggingface.co/datasets/Papajams/autoscientist-market-analysis-lenitnes-dataset.autoscientist-language-dataset
AutoScientist Language — Adapted Fine-tuning Dataset
Adapted training data (Adaption Labs AutoScientist pipeline, v5) used to fine-tune
Rishidar/autoscientist-language-qlora
(Qwen2.5-0.5B-Instruct, QLoRA SFT + DPO).
language_adapted.jsonl — SFT prompt/completion pairs (enhanced prompts + reasoning traces).
language_v5_raw.csv — full v5 output incl. chosen/rejected preference pairs used for DPO.
Also mirrored on Kaggle: rishidard/autoscientist-language-dataset.
autoscientist-healthcare-dataset
AutoScientist Healthcare — Adapted Fine-tuning Dataset
Adapted training data (Adaption Labs AutoScientist pipeline, v5) used to fine-tune
Rishidar/autoscientist-healthcare-qlora
(Qwen2.5-0.5B-Instruct, QLoRA SFT + DPO).
healthcare_adapted.jsonl — SFT prompt/completion pairs (enhanced prompts + reasoning traces).
healthcare_v5_raw.csv — full v5 output incl. chosen/rejected preference pairs used for DPO.
Also mirrored on Kaggle: rishidard/autoscientist-healthcare-dataset.
autoscientist-other-datasetnihongo-legal-finance-autoscientist-data
nihongo-legal-finance-autoscientist
AutoScientist Challenge entry dataset for Japanese expert QA in the language category.
Intended Use
This dataset is designed for supervised fine-tuning of Japanese assistants that explain
legal and financial concepts with uncertainty, source-awareness, and non-advice caveats.
Columns
instruction: user task
context: background information
response: target answer
rubric: quality expectations
category: subdomain… See the full description on the dataset page: https://huggingface.co/datasets/doraking/nihongo-legal-finance-autoscientist-data.autoscientist-marketing-dataset
AutoScientist Marketing — Adapted Fine-tuning Dataset
Adapted training data (Adaption Labs AutoScientist pipeline, v5) used to fine-tune
Rishidar/autoscientist-marketing-qlora
(Qwen2.5-0.5B-Instruct, QLoRA SFT + DPO).
marketing_adapted.jsonl — SFT prompt/completion pairs (enhanced prompts + reasoning traces).
marketing_v5_raw.csv — full v5 output incl. chosen/rejected preference pairs used for DPO.
Also mirrored on Kaggle: rishidard/autoscientist-marketing-dataset.
adaption-mmmed-autoscientist-gold
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-mmmed_autoscientist_gold
A refined, high-entropy multimodal clinical benchmark containing 431 perfectly aligned pairs of visual medical artifacts (X-rays, CT scans, ultrasounds, and histopathology profiles) and pre-concatenated case narratives with multiple-choice pathways. Optimized specifically for the AutoScientist Challenge (Healthcare Track) to train and evaluate… See the full description on the dataset page: https://huggingface.co/datasets/asadullahdogarr/adaption-mmmed-autoscientist-gold.autoscientist-dataviz-dataset
AutoScientist adapted dataset — dataviz
Adaption Labs AutoScientist v5 adapted fine-tuning data for the dataviz category.
dataviz_adapted.jsonl — prompt/completion pairs used for QLoRA SFT.
dataviz_v5_raw.csv — full Adaption output (prompt, completion, enhanced_prompt, chosen, rejected, reasoning_trace, embeddings) used for DPO.
Paired weights: Rishidar/autoscientist-dataviz-qlora (Kaggle mirror rishidard/autoscientist-dataviz-qlora).
adaption-finpath-autoscientist
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-finpath_autoscientist
This dataset contains deterministic personal finance prompts covering topics like debt payoff, emergency funds, and retirement contributions without provided answers. Designed for the AutoScientist framework, each sample presents a self-contained numerical reasoning scenario requiring grounded calculation. The collection serves as a source for… See the full description on the dataset page: https://huggingface.co/datasets/doraking/adaption-finpath-autoscientist.autoscientist-chartqaautoscientist-mathcode-datasetautoscientist-science-dataset
AutoScientist adapted dataset — science
Adaption Labs AutoScientist v5 adapted fine-tuning data for the science category.
science_adapted.jsonl — prompt/completion pairs used for QLoRA SFT.
science_v5_raw.csv — full Adaption output (prompt, completion, enhanced_prompt, chosen, rejected, reasoning_trace, embeddings) used for DPO.
Paired weights: Rishidar/autoscientist-science-qlora (Kaggle mirror rishidard/autoscientist-science-qlora).
orbura-autoscientist-dataviz-multimodal-pilotmath-adaptive-autoscientist
Math & Code — Reasoning-Augmented (AutoScientist Part 2)
Adaption-enhanced with reasoning traces + expert blueprint. 2000 rows.
Quality: 7.0 -> 8.2 (grade C->B), 17.1% improvement, top-third percentile (31.5).
Adaption dataset ID: 011a1456-5398-42a0-b9d4-ce6df72a992f. Category: Math & Code.
autoscientist-agriculture-datasetpersonal-finance-adaptive-autoscientist
Personal-Finance — Adaptive Data (AutoScientist Challenge, Part 2)
A personal-finance instruction dataset enhanced by the Adaption Adaptive Data API, submitted to
the Personal-Finance category. The scored metric is the dataset quality lift.
Result (Adaption dashboard)
Quality improvement: 82.5% — score 4.0 → 7.3 (grade D->B), 1000 rows.
Enhanced-quality percentile: 24.6 (top quartile of Adaption's reference distribution).
Adaption dataset ID:… See the full description on the dataset page: https://huggingface.co/datasets/Rome-1/personal-finance-adaptive-autoscientist.science-adaptive-autoscientist
Science — Reasoning-Augmented (AutoScientist Part 2)
Adaption-enhanced with reasoning traces + expert blueprint. 2000 rows.
Quality: 6.0 -> 7.6 (grade C->B), 26.7% improvement, top-third percentile (25.6).
Adaption dataset ID: d67cdbe0-2aed-492d-a790-0c6bd136ae29. Category: Science.
autoscientist-finance-dataset
AutoScientist Finance — Adapted Fine-tuning Dataset
Adapted training data (Adaption Labs AutoScientist pipeline, v5) used to fine-tune
Rishidar/autoscientist-finance-qlora
(Qwen2.5-0.5B-Instruct, QLoRA SFT + DPO).
finance_adapted.jsonl — SFT prompt/completion pairs (enhanced prompts + reasoning traces).
finance_v5_raw.csv — full v5 output incl. chosen/rejected preference pairs used for DPO.
Also mirrored on Kaggle: rishidard/autoscientist-finance-dataset.
adapted-yoruba-sft-autoscientistadapted-yoruba-sft-autoscientist-part2adapted-yoruba-sft-autoscientistauto_scientist
