datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Indic-KCC-Agri-Advisory-Benchmark
Indic-KCC-Agri-Advisory-Benchmark
⚠️ Benchmark only — not agronomic advice. This dataset and its reference
answers exist to score language models, not to be used as real farming
guidance. KCC references are noisy call-centre transcripts (see Status and
caveats); do not act on any answer, reference or candidate, as agricultural
advice.
Open-ended agricultural-advisory question answering in 11 Indian languages,
built from real farmer questions and the advisory answers… See the full description on the dataset page: https://huggingface.co/datasets/sthanika-ai/Indic-KCC-Agri-Advisory-Benchmark.chichewa-agriculture-advisory
Chichewa Agriculture Advisory
A Chichewa-language instruction dataset for fine-tuning a Llama-style chat
model to advise Malawian farmers, with a focus on maize (chimanga).
Each row is one conversation in the OpenAI / Llama chat-message schema:
{
"messages": [
{"role": "system", "content": "Ndinu katswiri wa za ulimi ku Malawi..."},
{"role": "user", "content": "Ndingabzale liti chimanga?"},
{"role": "assistant", "content": "Muyenera kubzala chimanga nthawi… See the full description on the dataset page: https://huggingface.co/datasets/PatrickChikuse/chichewa-agriculture-advisory.kisan-advisory-multilingual-indic
Kisan Advisory (Hindi / Punjabi / English)
Real farmer questions and Farm Tele Advisor answers from India's government
Kisan Call Centre helpline, adapted with AutoScientist, expanded into
Hindi and Punjabi, and filtered so that every row provably preserves the
agrochemical doses in its source note.
Rows (after dose filtering)
6,232
Language split
2,064 en / 2,726 hi / 1,442 pa
Quality grade
E → C (3.0 → 5.8)
Relative improvement
+93.3%
Percentile
13.8… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/kisan-advisory-multilingual-indic.agriculture-advisor-seed-v1
Agriculture Advisor Seed (v2)
A curated instruction-tuning seed for agricultural advisory models, built for the
Adaption AutoScientist Challenge
(Part 2, Agriculture track).
13,543 training rows + 450 held-out evaluation rows. 91% carry a real completion;
median completion length is 1,236 characters.
What this is
Two complementary halves:
Half
Rows
Completions
Farmer advisory (real queries, 4 countries)
~9,500
cleaned long-form answers
Quantitative… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/agriculture-advisor-seed-v1.agriculture-advisor-adapted-v1
Agriculture Advisor (Adaption-adapted) v1
The adapted dataset used to fine-tune our agricultural advisory model for the
Adaption AutoScientist Challenge.
Produced by running 15juneee/agriculture-advisor-seed-v1 through
Adaption's datasets.run. The seed carries the prompts and the curation; this carries
the completions the model was actually trained on.
Rows
17,039 rows. Adaption writes its output to enhanced_prompt / enhanced_completion
and leaves the uploaded… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/agriculture-advisor-adapted-v1.hindikrishi-farmer-advisory-dataset
🌾 HindiKrishi — Farmer Advisory Dataset
21,069 instruction-response pairs for training agricultural crop advisory models in Hindi and English, grounded in ICAR guidelines.
Dataset Details
Detail
Value
Total Examples
21,069
Languages
Hindi (primary), English
Format
JSONL (instruction, input, output)
Domain
Indian agriculture — crop diseases, pesticides, fertilizers, schemes
License
Apache 2.0
Format
Each example follows the… See the full description on the dataset page: https://huggingface.co/datasets/me-nabi/hindikrishi-farmer-advisory-dataset.AdvisorQADataset Card for AdvisorQA
As the integration of large language models into daily life is on the rise, there is still a lack of dataset for \textit{advising on subjective and personal dilemmas}. To address this gap, we introduce AdvisorQA, which aims to improve LLMs' capability to offer advice for deeply subjective concerns, utilizing the LifeProTips Reddit forum. This forum features a dynamic interaction where users post advice-seeking questions, receiving an average of 8.9 advice per query… See the full description on the dataset page: https://huggingface.co/datasets/mbkim/AdvisorQA.agriculture-advisor-adapted-multilingual-v1
Agriculture Advisor (Adaption-adapted, multilingual) v1
The adapted dataset used to fine-tune our agricultural advisory, localised model for the
Adaption AutoScientist Challenge.
Produced by running 15juneee/agriculture-advisor-seed-v1 through
Adaption's datasets.run. The seed carries the prompts and the curation; this carries
the completions the model was actually trained on.
Rows
22,270 rows. Adaption writes its output to enhanced_prompt / enhanced_completion… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/agriculture-advisor-adapted-multilingual-v1.Advisor-bench
Orionfold Advisor bench v0.1 + frozen curveballs
The evaluation set behind Orionfold/Advisor-GGUF —
a behavior bench for a governed corpus advisor: grounded answers with exact
source_id citations, clean refusals on missing-source and private-state
questions (including adversarial pretexts), and Route: workflow handoffs.
Scoring is deterministic — no LLM judge.
What's here
File
Rows
sha256[:12]
Role
pool.jsonl
75
6647680c10dc
Bench seed pool (answer 68 /… See the full description on the dataset page: https://huggingface.co/datasets/Orionfold/Advisor-bench.tamil-agri-advisory-qa
Tamil Agricultural Advisory Dataset — v13 Grade A (தமிழ் வேளாண்மை ஆலோசனை தரவுத்தொகுப்பு)
187 golden Tamil-language Q&A pairs | Grade A | 9.4/10 | 57.7th percentile — a quality-first instruction dataset adapted using Adaption's Adaptive Data platform, grounded in TNAU (Tamil Nadu Agricultural University) extension knowledge, Kisan Call Centre real farmer logs, and ICAR district contingency plans. Built for Tamil Nadu smallholder farmers.
GitHub → VinodAnbalagan/tamil-agri-dataset-… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/tamil-agri-advisory-qa.ukrainian-refugees-financial-advisory
Ukrainian Refugees Financial Advisory Dataset
A dataset of 500 synthetic advisory cases generated by a multi-agent
LLM pipeline that produces and evaluates retirement-oriented financial
guidance for Ukrainian refugee-like profiles in Poland.
Each case covers one full advisory cycle: synthetic profile generation →
draft recommendation + clarifying questions → final structured recommendation
→ automated quality evaluation.
GitHub: uliana0203/ai-agents-refugee-finance… See the full description on the dataset page: https://huggingface.co/datasets/Uliana333/ukrainian-refugees-financial-advisory.
