datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
peri-kbgardian-cigi-ai-documents
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 85,782 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
85,782
Total Size
1.47 GB
Total Tokens
199,872,861
Total Pages
0
Languages
59
Unique Keywords
76,373
Resource Types
32
Date Generated
2026-07-16 08:22:33
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/gardian-cigi-ai-documents.ifpri-ai-documents
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
21,726
Total Size
623.27 MB
Total Tokens
85,359,442
Total Pages
0
Languages
25
Unique Keywords
7,127
Resource Types
20
Date Generated
2026-07-31 02:55:20
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents.gardian-ai-ready-docs
⚠️ Heads up: Updated Dataset Available
This dataset has been updated with a newer version published on 27 Feb 2025. The latest version includes more updated and refined set of documents.
We recommend using the latest version, available at https://huggingface.co/datasets/CGIAR/gardian-cigi-ai-documents. This version remains accessible for reference and reproducibility purposes.
A Curated Research Corpus for Agricultural Advisory AI Applications
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/gardian-ai-ready-docs.embrapa-ai-documents
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 130,922 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
130,922
Total Size
256.66 MB
Total Tokens
11,827,230
Total Pages
0
Languages
8
Unique Keywords
106,492
Resource Types
11
Date Generated
2026-07-16 08:37:30
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/embrapa-ai-documents.cirad-ai-documents
GAIA / CIRAD Agricultural Documents (English)
A curated, machine-readable corpus of 4,209 agricultural research
publications sourced from
CIRAD (the French Agricultural Research
Centre for International Development), produced by the
Generative AI for Agriculture (GAIA)
project. Documents are indexed through
GARDIAN — CGIAR's agri-food
research index — and converted from PDF to structured JSON via the
GAIA-CIGI pipeline using GROBID.
This is an independent CIRAD corpus in the… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/cirad-ai-documents.usda-nal-ai-documents-en
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 22,529 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
22,529
Total Size
89.42 MB
Total Tokens
5,227,798
Total Pages
0
Languages
1
Unique Keywords
46,728
Resource Types
1
Date Generated
2026-07-16 08:32:49
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/usda-nal-ai-documents-en.ragas_QA_evaluation_datasetThe dataset comprises question answer pairs generated by the Mistral-7B-Instruct-v0.3 model, over a sample of the documents available in the CiGi knowledge base. The generative Q-A creation for evaluating CiGi was preferred over manual annotation because it yields diverse queries whose distribution better reflects downstream user intents. The dataset presentes the question-answer pairs, along with a reference to the source snippet that was used by the model to construct the answer.
