datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
data-advisor-safety-alignment[EMNLP 2024] Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
🌐 Homepage | 📖 Paper | 🤗 Dataset (Data Advisor) | 🤗 Dataset (Self-Instruct)
Disclaimer
The dataset contains content that may be offensive or harmful. This dataset is intended for research purposes, specifically to support efforts aimed at creating safer and less harmful AI systems. Please engage with it responsibly and at your own risk.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/fwnlp/data-advisor-safety-alignment.mycert-advisoryadvisor-twinsNot for commercial use. These are only being used to showcase Gaia as a framework to build personalised AI agents.
ipulse-ai-batch5-advisor-forecast-panel
iPulse AI Batch 5 Advisor Forecast Panel
This dataset exposes a compact, anonymized panel of production forecasts from iPulse AI, Future Edge Group's Open Agentic Investment Research Platform. It is designed for research on forecast combination, disagreement, correlated errors, regime dependence, and the effective number of independent forecasters.
The release contains seven showcase assets, twelve advisor configurations per asset, quarterly forecast paths extending five years… See the full description on the dataset page: https://huggingface.co/datasets/future-edge-group/ipulse-ai-batch5-advisor-forecast-panel.kcc-krishi-rag-sft-advisory-corpus
KCC-Krishi RAG/SFT Advisory Corpus
The KCC-Krishi RAG/SFT Advisory Corpus is a translated, quality-controlled, routing-aware research corpus derived from Kisan Call Centre records from the Government of India open-data ecosystem.
It was created for:
agricultural RAG research;
supervised fine-tuning research;
evidence-grounded response generation;
safety-routing experiments;
offline farmer-assistant prototyping;
reproducible dataset and model-training experiments.… See the full description on the dataset page: https://huggingface.co/datasets/uralstech/kcc-krishi-rag-sft-advisory-corpus.Indic-KCC-Agri-Advisory-Benchmark
Indic-KCC-Agri-Advisory-Benchmark
⚠️ Benchmark only — not agronomic advice. This dataset and its reference
answers exist to score language models, not to be used as real farming
guidance. KCC references are noisy call-centre transcripts (see Status and
caveats); do not act on any answer, reference or candidate, as agricultural
advice.
Open-ended agricultural-advisory question answering in 11 Indian languages,
built from real farmer questions and the advisory answers… See the full description on the dataset page: https://huggingface.co/datasets/sthanika-ai/Indic-KCC-Agri-Advisory-Benchmark.chichewa-agriculture-advisory
Chichewa Agriculture Advisory
A Chichewa-language instruction dataset for fine-tuning a Llama-style chat
model to advise Malawian farmers, with a focus on maize (chimanga).
Each row is one conversation in the OpenAI / Llama chat-message schema:
{
"messages": [
{"role": "system", "content": "Ndinu katswiri wa za ulimi ku Malawi..."},
{"role": "user", "content": "Ndingabzale liti chimanga?"},
{"role": "assistant", "content": "Muyenera kubzala chimanga nthawi… See the full description on the dataset page: https://huggingface.co/datasets/PatrickChikuse/chichewa-agriculture-advisory.adaption-hr-advisory-onet
HR Advisory Instruction Dataset (O*NET-grounded)
Instruction-tuning data for HR advisory work — job design, hiring, assessment, internal mobility, workforce analytics and tooling — with every factual claim traceable to a named O*NET occupation record.
Built for the Adaption Labs AutoScientist Challenge Part 2, HR track.
What is in it
Rows
5,415 (4,836 train / 579 eval)
Task families
19
Occupations covered
907 of 923 available
Response length… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-hr-advisory-onet.ai-legal-advisor-app
🤖 Smart Legal Advisor
Next-Generation Legal Advisory Application powered by NVIDIA AI.
📁 Project Structure
lib/
├── main.dart # App entry point
├── core/ # Core utilities and constants
│ ├── constants/
│ │ ├── app_colors.dart # Color palette
│ │ ├── app_strings.dart # Arabic localized strings
│ │ └── app_text_styles.dart # Typography
│ └── theme/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/moshabann/ai-legal-advisor-app.ml-advisor-benchmark
ML Experiment Advisor Benchmark
A 30-task benchmark for evaluating how well a language-model agent can advise on ML hyperparameter tuning, given experiment history and source code. Derived from 16 real training runs of Karpathy's autoresearch on an A40 GPU, plus 5 synthetic extensions covering edge cases.
Built for the meta-agent-improver project.
What's in each task
Every task is a workspace containing:
results.tsv — experiment history up to that point (commit… See the full description on the dataset page: https://huggingface.co/datasets/abhid1234/ml-advisor-benchmark.scholarship-advisor-dataagriculture-advisor-seed-v1
Agriculture Advisor Seed (v2)
A curated instruction-tuning seed for agricultural advisory models, built for the
Adaption AutoScientist Challenge
(Part 2, Agriculture track).
13,543 training rows + 450 held-out evaluation rows. 91% carry a real completion;
median completion length is 1,236 characters.
What this is
Two complementary halves:
Half
Rows
Completions
Farmer advisory (real queries, 4 countries)
~9,500
cleaned long-form answers
Quantitative… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/agriculture-advisor-seed-v1.hackathon-advisor-codex-traces
Hackathon Advisor Codex Session Traces
Real Codex session logs for the Hackathon Advisor project, selected from local Codex
rollout JSONL files and redacted before publication. The event stream preserves user
requests, assistant messages, tool calls, tool outputs, browser/search events, and
minimal session provenance needed to audit how the project was built.
Privacy filtering
The publisher applied openai/privacy-filter
at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.startup-advisor-dataset
🚀 Startup Advisor Dataset
A high-quality instruction-following dataset distilled from 8 foundational business and startup books, structured as actionable advice with real-world 2025 examples. Designed for fine-tuning large language models (e.g., Qwen, LLaMA, Mistral) to become expert startup advisors.
📖 Dataset Summary
Property
Value
Total Entries
1,564
Format
JSONL — ChatML (messages array)
Language
English
License
CreativeML OpenRAIL-M
Avg. Tokens… See the full description on the dataset page: https://huggingface.co/datasets/adamabuhamdan/startup-advisor-dataset.kisan-advisory-multilingual-indic
Kisan Advisory (Hindi / Punjabi / English)
Real farmer questions and Farm Tele Advisor answers from India's government
Kisan Call Centre helpline, adapted with AutoScientist, expanded into
Hindi and Punjabi, and filtered so that every row provably preserves the
agrochemical doses in its source note.
Rows (after dose filtering)
6,232
Language split
2,064 en / 2,726 hi / 1,442 pa
Quality grade
E → C (3.0 → 5.8)
Relative improvement
+93.3%
Percentile
13.8… See the full description on the dataset page: https://huggingface.co/datasets/tojpaj/kisan-advisory-multilingual-indic.agriculture-advisor-adapted-v1
Agriculture Advisor (Adaption-adapted) v1
The adapted dataset used to fine-tune our agricultural advisory model for the
Adaption AutoScientist Challenge.
Produced by running 15juneee/agriculture-advisor-seed-v1 through
Adaption's datasets.run. The seed carries the prompts and the curation; this carries
the completions the model was actually trained on.
Rows
17,039 rows. Adaption writes its output to enhanced_prompt / enhanced_completion
and leaves the uploaded… See the full description on the dataset page: https://huggingface.co/datasets/15juneee/agriculture-advisor-adapted-v1.hindikrishi-farmer-advisory-dataset
🌾 HindiKrishi — Farmer Advisory Dataset
21,069 instruction-response pairs for training agricultural crop advisory models in Hindi and English, grounded in ICAR guidelines.
Dataset Details
Detail
Value
Total Examples
21,069
Languages
Hindi (primary), English
Format
JSONL (instruction, input, output)
Domain
Indian agriculture — crop diseases, pesticides, fertilizers, schemes
License
Apache 2.0
Format
Each example follows the… See the full description on the dataset page: https://huggingface.co/datasets/me-nabi/hindikrishi-farmer-advisory-dataset.Financial-Advisory-Clientsgithub-advisory-2023fiducia-advisory-dataset
A Doutrina da Soberania Organizacional (Fiduciary Corpus)
Este repositório consiste na codificação em texto estruturado da Doutrina da Soberania Organizacional, arquitetada por Walter Maier Neto. O arcabouço estabelece os fundamentos para a Governança Digital Avançada, Mitigação de Riscos Sistêmicos e a evolução da Custódia Executiva na era dos algoritmos.
O corpus foi forjado com rigor forense para o treinamento (Pre-training), alinhamento ético (RLHF/DPO) e orquestração… See the full description on the dataset page: https://huggingface.co/datasets/wmaierbr/fiducia-advisory-dataset.AdvisorQADataset Card for AdvisorQA
As the integration of large language models into daily life is on the rise, there is still a lack of dataset for \textit{advising on subjective and personal dilemmas}. To address this gap, we introduce AdvisorQA, which aims to improve LLMs' capability to offer advice for deeply subjective concerns, utilizing the LifeProTips Reddit forum. This forum features a dynamic interaction where users post advice-seeking questions, receiving an average of 8.9 advice per query… See the full description on the dataset page: https://huggingface.co/datasets/mbkim/AdvisorQA.china-advisor
China Advisor Cache
This dataset contains cached public web pages, API responses, and extracted text for Chinese university advisor/faculty pages.
Current 985 coverage report: 77977 metadata records, 71303 teacher_profile records, and 69186 status-200 teacher_profile records across 39 Project 985 schools.
Contents
archives/china-advisor-cache-2026-06-20.tar.zst: full cache archive preserving the local data/cache/... layout.… See the full description on the dataset page: https://huggingface.co/datasets/kzoacn/china-advisor.Legal_Advisor-finetuneafrica-synth-agriculture-mobile-advisory-nigeria
Africa Synth Agriculture Mobile Advisory Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-agriculture-mobile-advisory-nigeria.ug-agri-advisory
Uganda Agricultural Advisory Dataset (UG-AgriAdvisory)
Built with Adaptive Data by Adaption | Crane AI Labs
Submitted to the Uncharted Data Challenge 2026 by Adaption Labs
Overview
The first open-source, bilingual (English/Luganda) agricultural advisory QA dataset grounded in Uganda's national agricultural extension guidelines from NAADS, MAAIF, and NARO.
Designed to power offline-capable AI advisory tools for Uganda's 2.2 million smallholder farmers — particularly… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/ug-agri-advisory.example-context-wealth-advisorysoftware-ai-replaceability
Software AI Replaceability Index
594 enterprise software products scored 0–100 for how completely an AI agent layer could replace or augment them — across CRM, ERP, marketing, support, finance, HR, and dev tools — with a workflow-modularity score and the known AI-native alternative.
Rows: 594
Source: Meo Advisors composite scoring
Methodology + interactive explorer: https://meoadvisors.com/software-ai-replaceability/
Part of: the open AI Workforce Data collection by Meo… See the full description on the dataset page: https://huggingface.co/datasets/Meo-Advisors/software-ai-replaceability.ai-platforms-catalog
AI Platforms Catalog
A reference catalog of 164 enterprise AI platforms — LLMs, agents, RAG, vector DBs, orchestration, observability — each tagged with category, enterprise-readiness score, pricing tier, and integration coverage.
Rows: 164
Source: Meo Advisors vendor research
Methodology + interactive explorer: https://meoadvisors.com/ai-platforms/
Part of: the open AI Workforce Data collection by Meo Advisors (GitHub)
License: CC BY-NC 4.0 — free for research, journalism, and… See the full description on the dataset page: https://huggingface.co/datasets/Meo-Advisors/ai-platforms-catalog.food_chinese_2017
Dataset Card for "food_chinese_2017"
More Information needed
github-advisory-2020.csv
