datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Embrapa-ai-documents-markdownifpri-ai-documents-markdown
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
21,726
Total Size
623.27 MB
Total Tokens
85,359,442
Total Pages
0
Languages
25
Unique Keywords
7,127
Resource Types
20
Date Generated
2026-07-31 02:55:20
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents-markdown.gardian-cigi-ai-documents-markdownperi-kbusda-nal-ai-documents-en-markdowngardian-cigi-ai-documents
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 85,782 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
85,782
Total Size
1.47 GB
Total Tokens
199,872,861
Total Pages
0
Languages
59
Unique Keywords
76,373
Resource Types
32
Date Generated
2026-07-16 08:22:33
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/gardian-cigi-ai-documents.ifpri-ai-documents
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
21,726
Total Size
623.27 MB
Total Tokens
85,359,442
Total Pages
0
Languages
25
Unique Keywords
7,127
Resource Types
20
Date Generated
2026-07-31 02:55:20
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents.gardian-ai-ready-docs
⚠️ Heads up: Updated Dataset Available
This dataset has been updated with a newer version published on 27 Feb 2025. The latest version includes more updated and refined set of documents.
We recommend using the latest version, available at https://huggingface.co/datasets/CGIAR/gardian-cigi-ai-documents. This version remains accessible for reference and reproducibility purposes.
A Curated Research Corpus for Agricultural Advisory AI Applications
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/gardian-ai-ready-docs.TranslationDataset_AgriQueriesThis Dataset consists of text pair in english and hindi_latin and some sheets which have hindi_devanagri as well.
The dataset is created by different human evaluators who have written hindi sentences in hindi latin (using english alphabets).
The dataset can be used for creating a translation model directly from english to hindi latin, which is how the users use ebcause of limitations of mobile keyboard in typing in hindi devanagri script.
embrapa-ai-documents
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 130,922 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
130,922
Total Size
256.66 MB
Total Tokens
11,827,230
Total Pages
0
Languages
8
Unique Keywords
106,492
Resource Types
11
Date Generated
2026-07-16 08:37:30
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/embrapa-ai-documents.KikuyuEnglish_translationThis dataset consists of agriculture related sentence pairs in english and kikuyu.
The dataset can be used for developing Kikuyu <-> English translation model.
Chi-Metrics-2024
Farmer.Chat: User Interaction and Evaluation Dataset from Agricultural AI Advisory Services in Kenya (2023-2024)
AbstractThis dataset comprises 2,300 anonymized user interaction records from Farmer.Chat, an AI-powered agricultural advisory platform deployed in Kenya between October 2023 and April 2024. The data captures diverse metrics including query types, response quality, user engagement patterns, and system performance across multiple agricultural value chains (dairy, coffee… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/Chi-Metrics-2024.AgricultureVideosQnAThe dataset is in XLS format with multiple sheets named for different languages.
The dataset is primarily used for training and ground truth of answers that can be generated for agriculture related queries from the videos.
Each sheet has list of video urls (youtube links) and the question that can be asked, corresponding answers that can be generated from the videos, source of information in the answer and time stamps.
The sources of information could be:
Transcript: based on what one hears… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/AgricultureVideosQnA.AgricultureVideosTranscriptThis dataset consists of agriculture videos in hindi and oriya.
The dataset consists of mulitple xls files and each xls file has column of video urls (youtube video links) and corresponding transcripts.
The transcripts are:
Generated by ASR models (for the purpose of benchmarking)
Manual transcripts
Time stamps
Manual translations
This dataset can be used for training and benchmarking domain specific models for ASR and translation. The time stamps serve as the soruce of audio and the… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/AgricultureVideosTranscript.cirad-ai-documents
GAIA / CIRAD Agricultural Documents (English)
A curated, machine-readable corpus of 4,209 agricultural research
publications sourced from
CIRAD (the French Agricultural Research
Centre for International Development), produced by the
Generative AI for Agriculture (GAIA)
project. Documents are indexed through
GARDIAN — CGIAR's agri-food
research index — and converted from PDF to structured JSON via the
GAIA-CIGI pipeline using GROBID.
This is an independent CIRAD corpus in the… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/cirad-ai-documents.ragas_gardian_evaluation_overlapping
📚 GARDIAN-RAGAS QA Dataset
A synthetic question–answer (QA) dataset generated from the GARDIAN corpus using RAGAS and the open-weight Mistral-7B-Instruct-v0.3 model. This dataset is designed to support evaluation and benchmarking of retrieval-augmented generation (RAG) systems, with an emphasis on grounded, high-fidelity QA generation.
📦 Dataset Summary
Source Corpus: GARDIAN scientific article collection
QA Generation Model: Mistral-7B-Instruct-v0.3
Sample Size: 1… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ragas_gardian_evaluation_overlapping.RAG-Chunk-Analysis
Description
The datasets contain human evaluation of retrieved chunks from agriculture documents for actual user queries.
Each chunk is marked as relevant and irrelevant. The relevant and irrelevant portion of the chunks are mentioned in a separate columns.
The dataset consists of multiple XLS files and each XLS file has multiple sheets corresponding to the content for the value chain.
The queries are taken from the actual user questions onf farmer.chat prototype bots.
For each user… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/RAG-Chunk-Analysis.AgricultureVideosQnA2The dataset is in XLS format with multiple sheets named for different languages. The dataset is primarily used for training and ground truth of answers that can be generated for agriculture related queries from the videos.
Each sheet has list of video urls (youtube links) and the question that can be asked, corresponding answers that can be generated from the videos, source of information in the answer and time stamps.
The sources of information could be:
Transcript: based on what one hears… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/AgricultureVideosQnA2.aquatic-macroinvertebrate-images
Dataset Card for Aquatic Macroinvertebrate Images
This image dataset contains 1,300 photos (one hundred images of each of the thirteen mini stream assessment scoring system, or miniSASS aquatic macroinvertebrate groups), used in the development of machine learning models to automatically assess the water quality score.
Dataset Description
This dataset accompanies the working paper, Digitally improving the identification of aquatic macroinvertebrates for indices used in… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/aquatic-macroinvertebrate-images.KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers.
Process of collection of data:
Selected users were given the option of doing a task and getting paid for it.
The users were supposed to record the sentence as it appeared on the screen.
The audio file thus obtained was validated matched with the sentences to fine tune the model.
Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/KikuyuASR_trainingdataset.ragas_gardian_evaluation_non_overlapping
📚 GARDIAN-RAGAS QA Dataset
A synthetic question–answer (QA) dataset generated from the GARDIAN corpus using RAGAS and the open-weight Mistral-7B-Instruct-v0.3 model. This dataset is designed to support evaluation and benchmarking of retrieval-augmented generation (RAG) systems, with an emphasis on grounded, high-fidelity QA generation.
📦 Dataset Summary
Source Corpus: GARDIAN scientific article collection
QA Generation Model: Mistral-7B-Instruct-v0.3
Sample Size: 1… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ragas_gardian_evaluation_non_overlapping.usda-nal-ai-documents-en
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 22,529 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
22,529
Total Size
89.42 MB
Total Tokens
5,227,798
Total Pages
0
Languages
1
Unique Keywords
46,728
Resource Types
1
Date Generated
2026-07-16 08:32:49
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/usda-nal-ai-documents-en.cirad-ai-documents-markdownragas_QA_evaluation_datasetThe dataset comprises question answer pairs generated by the Mistral-7B-Instruct-v0.3 model, over a sample of the documents available in the CiGi knowledge base. The generative Q-A creation for evaluating CiGi was preferred over manual annotation because it yields diverse queries whose distribution better reflects downstream user intents. The dataset presentes the question-answer pairs, along with a reference to the source snippet that was used by the model to construct the answer.
