datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
data-advisor-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct)
Disclaimer
The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/data-advisor-safety-alignment.Indic-KCC-Agri-Advisory-Benchmark
Indic-KCC-Agri-Advisory-Benchmark
⚠️ Benchmark only — not agronomic advice. This dataset and its reference
answers exist to score language models, not to be used as real farming
guidance. KCC references are noisy call-centre transcripts (see Status and
caveats); do not act on any answer, reference or candidate, as agricultural
advice.
Open-ended agricultural-advisory question answering in 11 Indian languages,
built from real farmer questions and the advisory answers… See the full description on the dataset page: https://huggingface.co/datasets/sthanika-ai/Indic-KCC-Agri-Advisory-Benchmark.chichewa-agriculture-advisory
Chichewa Agriculture Advisory
A Chichewa-language instruction dataset for fine-tuning a Llama-style chat
model to advise Malawian farmers, with a focus on maize (chimanga).
Each row is one conversation in the OpenAI / Llama chat-message schema:
{
"messages": [
{"role": "system", "content": "Ndinu katswiri wa za ulimi ku Malawi..."},
{"role": "user", "content": "Ndingabzale liti chimanga?"},
{"role": "assistant", "content": "Muyenera kubzala chimanga nthawi… See the full description on the dataset page: https://huggingface.co/datasets/PatrickChikuse/chichewa-agriculture-advisory.adaption-hr-advisory-onet
HR Advisory Instruction Dataset (O*NET-grounded)
Instruction-tuning data for HR advisory work — job design, hiring, assessment, internal mobility, workforce analytics and tooling — with every factual claim traceable to a named O*NET occupation record.
Built for the Adaption Labs AutoScientist Challenge Part 2, HR track.
What is in it
Rows
5,415 (4,836 train / 579 eval)
Task families
19
Occupations covered
907 of 923 available
Response length… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-hr-advisory-onet.agriculture-advisor-seed-v1
Agriculture Advisor Seed (v2)
A curated instruction-tuning seed for agricultural advisory models, built for the
Adaption AutoScientist Challenge
(Part 2, Agriculture track).
13,543 training rows + 450 held-out evaluation rows. 91% carry a real completion;
median completion length is 1,236 characters.
What this is
Two complementary halves:
Half
Rows
Completions
Farmer advisory (real queries, 4 countries)
~9,500
cleaned long-form answers
Quantitative… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/agriculture-advisor-seed-v1.hackathon-advisor-codex-traces
Hackathon Advisor Codex Session Traces
Real Codex session logs for the Hackathon Advisor project, selected from local Codex
rollout JSONL files and redacted before publication. The event stream preserves user
requests, assistant messages, tool calls, tool outputs, browser/search events, and
minimal session provenance needed to audit how the project was built.
Privacy filtering
The publisher applied openai/privacy-filter
at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.startup-advisor-dataset
🚀 Startup Advisor Dataset
A high-quality instruction-following dataset distilled from 8 foundational business and startup books, structured as actionable advice with real-world 2025 examples. Designed for fine-tuning large language models (e.g., Qwen, LLaMA, Mistral) to become expert startup advisors.
📖 Dataset Summary
Property
Value
Total Entries
1,564
Format
JSONL — ChatML (messages array)
Language
English
License
CreativeML OpenRAIL-M
Avg. Tokens… See the full description on the dataset page: https://huggingface.co/datasets/adamabuhamdan/startup-advisor-dataset.kisan-advisory-multilingual-indic
Kisan Advisory (Hindi / Punjabi / English)
Real farmer questions and Farm Tele Advisor answers from India's government
Kisan Call Centre helpline, adapted with AutoScientist, expanded into
Hindi and Punjabi, and filtered so that every row provably preserves the
agrochemical doses in its source note.
Rows (after dose filtering)
6,232
Language split
2,064 en / 2,726 hi / 1,442 pa
Quality grade
E → C (3.0 → 5.8)
Relative improvement
+93.3%
Percentile
13.8… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/kisan-advisory-multilingual-indic.agriculture-advisor-adapted-v1
Agriculture Advisor (Adaption-adapted) v1
The adapted dataset used to fine-tune our agricultural advisory model for the
Adaption AutoScientist Challenge.
Produced by running 15juneee/agriculture-advisor-seed-v1 through
Adaption's datasets.run. The seed carries the prompts and the curation; this carries
the completions the model was actually trained on.
Rows
17,039 rows. Adaption writes its output to enhanced_prompt / enhanced_completion
and leaves the uploaded… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/agriculture-advisor-adapted-v1.hindikrishi-farmer-advisory-dataset
🌾 HindiKrishi — Farmer Advisory Dataset
21,069 instruction-response pairs for training agricultural crop advisory models in Hindi and English, grounded in ICAR guidelines.
Dataset Details
Detail
Value
Total Examples
21,069
Languages
Hindi (primary), English
Format
JSONL (instruction, input, output)
Domain
Indian agriculture — crop diseases, pesticides, fertilizers, schemes
License
Apache 2.0
Format
Each example follows the… See the full description on the dataset page: https://huggingface.co/datasets/me-nabi/hindikrishi-farmer-advisory-dataset.fiducia-advisory-dataset
A Doutrina da Soberania Organizacional (Fiduciary Corpus)
Este repositório consiste na codificação em texto estruturado da Doutrina da Soberania Organizacional, arquitetada por Walter Maier Neto. O arcabouço estabelece os fundamentos para a Governança Digital Avançada, Mitigação de Riscos Sistêmicos e a evolução da Custódia Executiva na era dos algoritmos.
O corpus foi forjado com rigor forense para o treinamento (Pre-training), alinhamento ético (RLHF/DPO) e orquestração… See the full description on the dataset page: https://huggingface.co/datasets/wmaierbr/fiducia-advisory-dataset.AdvisorQADataset Card for AdvisorQA
As the integration of large language models into daily life is on the rise, there is still a lack of dataset for \textit{advising on subjective and personal dilemmas}. To address this gap, we introduce AdvisorQA, which aims to improve LLMs' capability to offer advice for deeply subjective concerns, utilizing the LifeProTips Reddit forum. This forum features a dynamic interaction where users post advice-seeking questions, receiving an average of 8.9 advice per query… See the full description on the dataset page: https://huggingface.co/datasets/mbkim/AdvisorQA.hackathon-advisor-quest-dataset
Hackathon Advisor — Quest Classification SFT Dataset
Supervised fine-tuning data that teaches MiniCPM5-1B to classify a Build Small
Hackathon project against 13 judging dimensions from a two-segment README + app-file
prompt, emitting strict JSON with short, source-attributed evidence. Trains the LoRA at
build-small-hackathon/hackathon-advisor-quest-minicpm5-lora.
Files
quest_sft.jsonl — the dataset (one lora_sft_example per line; the viewer split).… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-quest-dataset.agriculture-advisor-adapted-multilingual-v1
Agriculture Advisor (Adaption-adapted, multilingual) v1
The adapted dataset used to fine-tune our agricultural advisory, localised model for the
Adaption AutoScientist Challenge.
Produced by running 15juneee/agriculture-advisor-seed-v1 through
Adaption's datasets.run. The seed carries the prompts and the curation; this carries
the completions the model was actually trained on.
Rows
22,270 rows. Adaption writes its output to enhanced_prompt / enhanced_completion… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/agriculture-advisor-adapted-multilingual-v1.marketnews-advisory-indic
Market & News Advisory (Hindi/Punjabi)
Native Hindi and Punjabi text from
ai4bharat/IndicCorpV2,
adapted with AutoScientist into substantive domain responses written by
a market analyst explaining what a news item means for Indian markets.
Rows
3,596
Unique source texts
1,200
Absolute quality score
8.7/10 (grade B)
Source score before adaptation
9.0/10 (grade B)
Percentile
19.2
Relative change
-3.3%
Median response length
523 chars
This… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/marketnews-advisory-indic.Advisor-bench
Orionfold Advisor bench v0.1 + frozen curveballs
The evaluation set behind Orionfold/Advisor-GGUF —
a behavior bench for a governed corpus advisor: grounded answers with exact
source_id citations, clean refusals on missing-source and private-state
questions (including adversarial pretexts), and Route: workflow handoffs.
Scoring is deterministic — no LLM judge.
What's here
File
Rows
sha256[:12]
Role
pool.jsonl
75
6647680c10dc
Bench seed pool (answer 68 /… See the full description on the dataset page: https://huggingface.co/datasets/Orionfold/Advisor-bench.chichewa-agri-advisor
Chichewa Agriculture Advisory — Llama Finetune Dataset
A Chichewa-language instruction dataset for fine-tuning a Llama-style model
to give agricultural advice to Malawian farmers (focus: maize / chimanga).
Status
Total conversations: 198
Train / val split: 178 / 20 (90/10, seed=42)
Recommended minimum: ~500–1,000 examples for a usable LoRA finetune.
198 is enough to smoke-test the pipeline; expect underfitting on real prompts.
Folder layout
chichewa_advisory/… See the full description on the dataset page: https://huggingface.co/datasets/PatrickChikuse/chichewa-agri-advisor.tamil-agri-advisory-qa
Tamil Agricultural Advisory Dataset — v13 Grade A (தமிழ் வேளாண்மை ஆலோசனை தரவுத்தொகுப்பு)
187 golden Tamil-language Q&A pairs | Grade A | 9.4/10 | 57.7th percentile — a quality-first instruction dataset adapted using Adaption's Adaptive Data platform, grounded in TNAU (Tamil Nadu Agricultural University) extension knowledge, Kisan Call Centre real farmer logs, and ICAR district contingency plans. Built for Tamil Nadu smallholder farmers.
GitHub → VinodAnbalagan/tamil-agri-dataset-… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/tamil-agri-advisory-qa.defendable-pain-cisa-advisory-pain-v0.1
CISA Advisory Pain Receipt
"the advisory" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 4 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here
4 pain receipts… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-cisa-advisory-pain-v0.1.student-advisor-datasetmedical_advisory_queriesukrainian-refugees-financial-advisory
Ukrainian Refugees Financial Advisory Dataset
A dataset of 500 synthetic advisory cases generated by a multi-agent
LLM pipeline that produces and evaluates retirement-oriented financial
guidance for Ukrainian refugee-like profiles in Poland.
Each case covers one full advisory cycle: synthetic profile generation →
draft recommendation + clarifying questions → final structured recommendation
→ automated quality evaluation.
GitHub: uliana0203/ai-agents-refugee-finance… See the full description on the dataset page: https://huggingface.co/datasets/Uliana333/ukrainian-refugees-financial-advisory.financial-inclusion-advisory-indic
Financial-Inclusion Advisory (Hindi/Punjabi)
Native Hindi and Punjabi text from
ai4bharat/IndicCorpV2,
adapted with AutoScientist into substantive domain responses written by
a financial-inclusion counsellor explaining money matters in plain terms.
Rows
474
Unique source texts
474
Absolute quality score
8.8/10 (grade B)
Source score before adaptation
9.0/10 (grade A)
Percentile
19.2
Relative change
-2.2%
Median response length
478 chars… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/financial-inclusion-advisory-indic.workers-rights-advisory-indic
Workers' Rights Advisory (Hindi/Punjabi)
Native Hindi and Punjabi text from
ai4bharat/IndicCorpV2,
adapted with AutoScientist into substantive domain responses written by
a workers'-rights advisor identifying the issue and concrete next steps.
Rows
561
Unique source texts
561
Absolute quality score
8.8/10 (grade B)
Source score before adaptation
9.0/10 (grade A)
Percentile
43.9
Relative change
-2.2%
Median response length
569 chars… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/workers-rights-advisory-indic.Korean_Chinese_Indian_Food.jsoncfo-advisory-evaladaption-finance-advisory-reasoning
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-finance_advisory_reasoning
This dataset contains expert-level financial advisory scenarios across equity research, distressed debt, M&A, and valuation, paired with detailed step-by-step reasoning traces and final recommendations. Each sample presents a specific client problem followed by a structured analytical breakdown that identifies critical markers, validates assumptions, and… See the full description on the dataset page: https://huggingface.co/datasets/Azfarhashmi/adaption-finance-advisory-reasoning.uk-university-advisorai-solution-advisor-train
