datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Indic-KCC-Agri-Advisory-Benchmark
Indic-KCC-Agri-Advisory-Benchmark
⚠️ Benchmark only — not agronomic advice. This dataset and its reference
answers exist to score language models, not to be used as real farming
guidance. KCC references are noisy call-centre transcripts (see Status and
caveats); do not act on any answer, reference or candidate, as agricultural
advice.
Open-ended agricultural-advisory question answering in 11 Indian languages,
built from real farmer questions and the advisory answers… See the full description on the dataset page: https://huggingface.co/datasets/sthanika-ai/Indic-KCC-Agri-Advisory-Benchmark.kcc-krishi-rag-sft-advisory-corpus
KCC-Krishi RAG/SFT Advisory Corpus
The KCC-Krishi RAG/SFT Advisory Corpus is a translated, quality-controlled, routing-aware research corpus derived from Kisan Call Centre records from the Government of India open-data ecosystem.
It was created for:
agricultural RAG research;
supervised fine-tuning research;
evidence-grounded response generation;
safety-routing experiments;
offline farmer-assistant prototyping;
reproducible dataset and model-training experiments.… See the full description on the dataset page: https://huggingface.co/datasets/uralstech/kcc-krishi-rag-sft-advisory-corpus.adaption-hr-advisory-onet
HR Advisory Instruction Dataset (O*NET-grounded)
Instruction-tuning data for HR advisory work — job design, hiring, assessment, internal mobility, workforce analytics and tooling — with every factual claim traceable to a named O*NET occupation record.
Built for the Adaption Labs AutoScientist Challenge Part 2, HR track.
What is in it
Rows
5,415 (4,836 train / 579 eval)
Task families
19
Occupations covered
907 of 923 available
Response length… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-hr-advisory-onet.chichewa-agriculture-advisory
Chichewa Agriculture Advisory
A Chichewa-language instruction dataset for fine-tuning a Llama-style chat
model to advise Malawian farmers, with a focus on maize (chimanga).
Each row is one conversation in the OpenAI / Llama chat-message schema:
{
"messages": [
{"role": "system", "content": "Ndinu katswiri wa za ulimi ku Malawi..."},
{"role": "user", "content": "Ndingabzale liti chimanga?"},
{"role": "assistant", "content": "Muyenera kubzala chimanga nthawi… See the full description on the dataset page: https://huggingface.co/datasets/PatrickChikuse/chichewa-agriculture-advisory.hackathon-advisor-codex-traces
Hackathon Advisor Codex Session Traces
Real Codex session logs for the Hackathon Advisor project, selected from local Codex
rollout JSONL files and redacted before publication. The event stream preserves user
requests, assistant messages, tool calls, tool outputs, browser/search events, and
minimal session provenance needed to audit how the project was built.
Privacy filtering
The publisher applied openai/privacy-filter
at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.agriculture-advisor-seed-v1
Agriculture Advisor Seed (v2)
A curated instruction-tuning seed for agricultural advisory models, built for the
Adaption AutoScientist Challenge
(Part 2, Agriculture track).
13,543 training rows + 450 held-out evaluation rows. 91% carry a real completion;
median completion length is 1,236 characters.
What this is
Two complementary halves:
Half
Rows
Completions
Farmer advisory (real queries, 4 countries)
~9,500
cleaned long-form answers
Quantitative… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/agriculture-advisor-seed-v1.startup-advisor-dataset
🚀 Startup Advisor Dataset
A high-quality instruction-following dataset distilled from 8 foundational business and startup books, structured as actionable advice with real-world 2025 examples. Designed for fine-tuning large language models (e.g., Qwen, LLaMA, Mistral) to become expert startup advisors.
📖 Dataset Summary
Property
Value
Total Entries
1,564
Format
JSONL — ChatML (messages array)
Language
English
License
CreativeML OpenRAIL-M
Avg. Tokens… See the full description on the dataset page: https://huggingface.co/datasets/adamabuhamdan/startup-advisor-dataset.hindikrishi-farmer-advisory-dataset
🌾 HindiKrishi — Farmer Advisory Dataset
21,069 instruction-response pairs for training agricultural crop advisory models in Hindi and English, grounded in ICAR guidelines.
Dataset Details
Detail
Value
Total Examples
21,069
Languages
Hindi (primary), English
Format
JSONL (instruction, input, output)
Domain
Indian agriculture — crop diseases, pesticides, fertilizers, schemes
License
Apache 2.0
Format
Each example follows the… See the full description on the dataset page: https://huggingface.co/datasets/me-nabi/hindikrishi-farmer-advisory-dataset.agriculture-advisor-adapted-v1
Agriculture Advisor (Adaption-adapted) v1
The adapted dataset used to fine-tune our agricultural advisory model for the
Adaption AutoScientist Challenge.
Produced by running 15juneee/agriculture-advisor-seed-v1 through
Adaption's datasets.run. The seed carries the prompts and the curation; this carries
the completions the model was actually trained on.
Rows
17,039 rows. Adaption writes its output to enhanced_prompt / enhanced_completion
and leaves the uploaded… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/agriculture-advisor-adapted-v1.azure-advisor-sft
Azure Advisor SFT Dataset
Supervised fine-tuning dataset for training language models to generate Azure Advisor-style recommendations.
Dataset Description
This dataset contains 410 synthetic examples (348 train / 41 eval / 21 held-out) designed to teach a model to analyze Azure workload configurations and produce structured recommendations across all 5 Azure Advisor categories.
Categories Covered
Cost - Right-sizing VMs, reserved instances, unused resources… See the full description on the dataset page: https://huggingface.co/datasets/thegovind/azure-advisor-sft.marketnews-advisory-indic
Market & News Advisory (Hindi/Punjabi)
Native Hindi and Punjabi text from
ai4bharat/IndicCorpV2,
adapted with AutoScientist into substantive domain responses written by
a market analyst explaining what a news item means for Indian markets.
Rows
3,596
Unique source texts
1,200
Absolute quality score
8.7/10 (grade B)
Source score before adaptation
9.0/10 (grade B)
Percentile
19.2
Relative change
-3.3%
Median response length
523 chars
This… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/marketnews-advisory-indic.agriculture-advisor-adapted-multilingual-v1
Agriculture Advisor (Adaption-adapted, multilingual) v1
The adapted dataset used to fine-tune our agricultural advisory, localised model for the
Adaption AutoScientist Challenge.
Produced by running 15juneee/agriculture-advisor-seed-v1 through
Adaption's datasets.run. The seed carries the prompts and the curation; this carries
the completions the model was actually trained on.
Rows
22,270 rows. Adaption writes its output to enhanced_prompt / enhanced_completion… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/agriculture-advisor-adapted-multilingual-v1.bitcoin-investment-advisory-dataset
Bitcoin Investment Advisory Training Dataset
Dataset Description
This dataset contains comprehensive Bitcoin investment advisory training data designed for fine-tuning large language models to provide institutional-grade cryptocurrency investment advice. The dataset consists of 2,437 high-quality instruction-input-output triplets covering Bitcoin market analysis from 2018-01-01 to 2024-12-31.
Dataset Features
Total Samples: 2,437
Date Range: 2018-01-01 to… See the full description on the dataset page: https://huggingface.co/datasets/tahamajs/bitcoin-investment-advisory-dataset.hackathon-advisor-quest-dataset
Hackathon Advisor — Quest Classification SFT Dataset
Supervised fine-tuning data that teaches MiniCPM5-1B to classify a Build Small
Hackathon project against 13 judging dimensions from a two-segment README + app-file
prompt, emitting strict JSON with short, source-attributed evidence. Trains the LoRA at
build-small-hackathon/hackathon-advisor-quest-minicpm5-lora.
Files
quest_sft.jsonl — the dataset (one lora_sft_example per line; the viewer split).… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-quest-dataset.azure-advisor-grpo-benchmark
Azure Advisor GRPO Benchmark Dataset
Evaluation benchmark for measuring the quality of Azure Advisor recommendation generation, used for GRPO (Group Relative Policy Optimization) training and model evaluation.
Dataset Description
This dataset contains 106 evaluation examples with ground truth labels, designed to score model outputs across 5 reward dimensions.
Purpose
During GRPO training: Score generated recommendations to select high-reward samples
Model… See the full description on the dataset page: https://huggingface.co/datasets/thegovind/azure-advisor-grpo-benchmark.Advisor-bench
Orionfold Advisor bench v0.1 + frozen curveballs
The evaluation set behind Orionfold/Advisor-GGUF —
a behavior bench for a governed corpus advisor: grounded answers with exact
source_id citations, clean refusals on missing-source and private-state
questions (including adversarial pretexts), and Route: workflow handoffs.
Scoring is deterministic — no LLM judge.
What's here
File
Rows
sha256[:12]
Role
pool.jsonl
75
6647680c10dc
Bench seed pool (answer 68 /… See the full description on the dataset page: https://huggingface.co/datasets/Orionfold/Advisor-bench.smolified-career-advisor-bot
🤏 smolified-career-advisor-bot
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-career-advisor-bot.
📦 Asset Details
Origin: Smolify Foundry (Job ID: c387dadb)
Records: 4309
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
financial-inclusion-advisory-indic
Financial-Inclusion Advisory (Hindi/Punjabi)
Native Hindi and Punjabi text from
ai4bharat/IndicCorpV2,
adapted with AutoScientist into substantive domain responses written by
a financial-inclusion counsellor explaining money matters in plain terms.
Rows
474
Unique source texts
474
Absolute quality score
8.8/10 (grade B)
Source score before adaptation
9.0/10 (grade A)
Percentile
19.2
Relative change
-2.2%
Median response length
478 chars… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/financial-inclusion-advisory-indic.workers-rights-advisory-indic
Workers' Rights Advisory (Hindi/Punjabi)
Native Hindi and Punjabi text from
ai4bharat/IndicCorpV2,
adapted with AutoScientist into substantive domain responses written by
a workers'-rights advisor identifying the issue and concrete next steps.
Rows
561
Unique source texts
561
Absolute quality score
8.8/10 (grade B)
Source score before adaptation
9.0/10 (grade A)
Percentile
43.9
Relative change
-2.2%
Median response length
569 chars… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/workers-rights-advisory-indic.ukrainian-refugees-financial-advisory
Ukrainian Refugees Financial Advisory Dataset
A dataset of 500 synthetic advisory cases generated by a multi-agent
LLM pipeline that produces and evaluates retirement-oriented financial
guidance for Ukrainian refugee-like profiles in Poland.
Each case covers one full advisory cycle: synthetic profile generation →
draft recommendation + clarifying questions → final structured recommendation
→ automated quality evaluation.
GitHub: uliana0203/ai-agents-refugee-finance… See the full description on the dataset page: https://huggingface.co/datasets/Uliana333/ukrainian-refugees-financial-advisory.
