datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mycert-advisorykcc-krishi-rag-sft-advisory-corpus
KCC-Krishi RAG/SFT Advisory Corpus
The KCC-Krishi RAG/SFT Advisory Corpus is a translated, quality-controlled, routing-aware research corpus derived from Kisan Call Centre records from the Government of India open-data ecosystem.
It was created for:
agricultural RAG research;
supervised fine-tuning research;
evidence-grounded response generation;
safety-routing experiments;
offline farmer-assistant prototyping;
reproducible dataset and model-training experiments.… See the full description on the dataset page: https://huggingface.co/datasets/uralstech/kcc-krishi-rag-sft-advisory-corpus.Indic-KCC-Agri-Advisory-Benchmark
Indic-KCC-Agri-Advisory-Benchmark
⚠️ Benchmark only — not agronomic advice. This dataset and its reference
answers exist to score language models, not to be used as real farming
guidance. KCC references are noisy call-centre transcripts (see Known
limitations); do not act on any answer, reference or candidate, as
agricultural advice.
Why this is gated
Two separate reasons, both real:
Benchmark integrity. Gold reference answers sitting in the open get… See the full description on the dataset page: https://huggingface.co/datasets/sthanika-ai/Indic-KCC-Agri-Advisory-Benchmark.chichewa-agriculture-advisory
Chichewa Agriculture Advisory
A Chichewa-language instruction dataset for fine-tuning a Llama-style chat
model to advise Malawian farmers, with a focus on maize (chimanga).
Each row is one conversation in the OpenAI / Llama chat-message schema:
{
"messages": [
{"role": "system", "content": "Ndinu katswiri wa za ulimi ku Malawi..."},
{"role": "user", "content": "Ndingabzale liti chimanga?"},
{"role": "assistant", "content": "Muyenera kubzala chimanga nthawi… See the full description on the dataset page: https://huggingface.co/datasets/PatrickChikuse/chichewa-agriculture-advisory.adaption-hr-advisory-onet
HR Advisory Instruction Dataset (O*NET-grounded)
Instruction-tuning data for HR advisory work — job design, hiring, assessment, internal mobility, workforce analytics and tooling — with every factual claim traceable to a named O*NET occupation record.
Built for the Adaption Labs AutoScientist Challenge Part 2, HR track.
What is in it
Rows
5,415 (4,836 train / 579 eval)
Task families
19
Occupations covered
907 of 923 available
Response length… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-hr-advisory-onet.kisan-advisory-multilingual-indic
Kisan Advisory (Hindi / Punjabi / English)
Real farmer questions and Farm Tele Advisor answers from India's government
Kisan Call Centre helpline, adapted with AutoScientist, expanded into
Hindi and Punjabi, and filtered so that every row provably preserves the
agrochemical doses in its source note.
Rows (after dose filtering)
6,232
Language split
2,064 en / 2,726 hi / 1,442 pa
Quality grade
E → C (3.0 → 5.8)
Relative improvement
+93.3%
Percentile
13.8… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/kisan-advisory-multilingual-indic.Financial-Advisory-Clientsfiducia-advisory-dataset
A Doutrina da Soberania Organizacional (Fiduciary Corpus)
Este repositório consiste na codificação em texto estruturado da Doutrina da Soberania Organizacional, arquitetada por Walter Maier Neto. O arcabouço estabelece os fundamentos para a Governança Digital Avançada, Mitigação de Riscos Sistêmicos e a evolução da Custódia Executiva na era dos algoritmos.
O corpus foi forjado com rigor forense para o treinamento (Pre-training), alinhamento ético (RLHF/DPO) e orquestração… See the full description on the dataset page: https://huggingface.co/datasets/wmaierbr/fiducia-advisory-dataset.github-advisory-2023hindikrishi-farmer-advisory-dataset
🌾 HindiKrishi — Farmer Advisory Dataset
21,069 instruction-response pairs for training agricultural crop advisory models in Hindi and English, grounded in ICAR guidelines.
Dataset Details
Detail
Value
Total Examples
21,069
Languages
Hindi (primary), English
Format
JSONL (instruction, input, output)
Domain
Indian agriculture — crop diseases, pesticides, fertilizers, schemes
License
Apache 2.0
Format
Each example follows the… See the full description on the dataset page: https://huggingface.co/datasets/me-nabi/hindikrishi-farmer-advisory-dataset.africa-synth-agriculture-mobile-advisory-nigeria
Africa Synth Agriculture Mobile Advisory Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-agriculture-mobile-advisory-nigeria.ug-agri-advisory
Uganda Agricultural Advisory Dataset (UG-AgriAdvisory)
Built with Adaptive Data by Adaption | Crane AI Labs
Submitted to the Uncharted Data Challenge 2026 by Adaption Labs
Overview
The first open-source, bilingual (English/Luganda) agricultural advisory QA dataset grounded in Uganda's national agricultural extension guidelines from NAADS, MAAIF, and NARO.
Designed to power offline-capable AI advisory tools for Uganda's 2.2 million smallholder farmers — particularly… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/ug-agri-advisory.example-context-wealth-advisoryfood_chinese_2017
Dataset Card for "food_chinese_2017"
More Information needed
github-advisory-2020.csvbitcoin-investment-advisory-dataset
Bitcoin Investment Advisory Training Dataset
Dataset Description
This dataset contains comprehensive Bitcoin investment advisory training data designed for fine-tuning large language models to provide institutional-grade cryptocurrency investment advice. The dataset consists of 2,437 high-quality instruction-input-output triplets covering Bitcoin market analysis from 2018-01-01 to 2024-12-31.
Dataset Features
Total Samples: 2,437
Date Range: 2018-01-01 to… See the full description on the dataset page: https://huggingface.co/datasets/tahamajs/bitcoin-investment-advisory-dataset.Climate_advisory_and_scam_detection
Adaption Upload Bundles
Generated by scripts/grade_a_rebuild.py with blueprint grade_a_2026_04_30.
File
Use
instruction_dataset_main.jsonl
Full structured instruction upload (1680 rows)
instruction_dataset_smoke.jsonl
50-row smoke test, 5 languages x 5 formats x 2 rows
instruction_dataset_smoke_v2.jsonl
Same smoke set for the previous filename expected by notes
preference_pairs_scams.jsonl
Mirrored genuine-vs-scam preference pairs (120 pairs)
Hard rules:… See the full description on the dataset page: https://huggingface.co/datasets/sahilmaniyar888/Climate_advisory_and_scam_detection.ea-crop-advisory
ea-crop-advisory
East African Smallholder Crop Advisory Dataset — agronomic guidance for 15+ crops across Uganda, Kenya, Tanzania, and Rwanda, grounded in local varieties, seasons, and district-level context.
Dataset Summary
This dataset was generated by gimmy256 as part of the East African AI dataset initiative.
It contains 206 records split into train / validation / test sets.
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/ea-crop-advisory.marketnews-advisory-indic
Market & News Advisory (Hindi/Punjabi)
Native Hindi and Punjabi text from
ai4bharat/IndicCorpV2,
adapted with AutoScientist into substantive domain responses written by
a market analyst explaining what a news item means for Indian markets.
Rows
3,596
Unique source texts
1,200
Absolute quality score
8.7/10 (grade B)
Source score before adaptation
9.0/10 (grade B)
Percentile
19.2
Relative change
-3.3%
Median response length
523 chars
This… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/marketnews-advisory-indic.github-advisory-2019.csvgithub-advisory-2022github-advisory-2021.csvgithub-advisory-2017.csvtamil-agri-advisory-qa
Tamil Agricultural Advisory Dataset — v13 Grade A (தமிழ் வேளாண்மை ஆலோசனை தரவுத்தொகுப்பு)
187 golden Tamil-language Q&A pairs | Grade A | 9.4/10 | 57.7th percentile — a quality-first instruction dataset adapted using Adaption's Adaptive Data platform, grounded in TNAU (Tamil Nadu Agricultural University) extension knowledge, Kisan Call Centre real farmer logs, and ICAR district contingency plans. Built for Tamil Nadu smallholder farmers.
GitHub → VinodAnbalagan/tamil-agri-dataset-… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/tamil-agri-advisory-qa.github-advisory-2018.csvdefendable-pain-cisa-advisory-pain-v0.1
CISA Advisory Pain Receipt
"the advisory" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 4 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here
4 pain receipts… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-cisa-advisory-pain-v0.1.financial-inclusion-advisory-indic
Financial-Inclusion Advisory (Hindi/Punjabi)
Native Hindi and Punjabi text from
ai4bharat/IndicCorpV2,
adapted with AutoScientist into substantive domain responses written by
a financial-inclusion counsellor explaining money matters in plain terms.
Rows
474
Unique source texts
474
Absolute quality score
8.8/10 (grade B)
Source score before adaptation
9.0/10 (grade A)
Percentile
19.2
Relative change
-2.2%
Median response length
478 chars… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/financial-inclusion-advisory-indic.medical_advisory_queriesukrainian-refugees-financial-advisory
Ukrainian Refugees Financial Advisory Dataset
A dataset of 500 synthetic advisory cases generated by a multi-agent
LLM pipeline that produces and evaluates retirement-oriented financial
guidance for Ukrainian refugee-like profiles in Poland.
Each case covers one full advisory cycle: synthetic profile generation →
draft recommendation + clarifying questions → final structured recommendation
→ automated quality evaluation.
GitHub: uliana0203/ai-agents-refugee-finance… See the full description on the dataset page: https://huggingface.co/datasets/Uliana333/ukrainian-refugees-financial-advisory.workers-rights-advisory-indic
Workers' Rights Advisory (Hindi/Punjabi)
Native Hindi and Punjabi text from
ai4bharat/IndicCorpV2,
adapted with AutoScientist into substantive domain responses written by
a workers'-rights advisor identifying the issue and concrete next steps.
Rows
561
Unique source texts
561
Absolute quality score
8.8/10 (grade B)
Source score before adaptation
9.0/10 (grade A)
Percentile
43.9
Relative change
-2.2%
Median response length
569 chars… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/workers-rights-advisory-indic.
