CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CGIAR /Embrapa-ai-documents-markdown1 likes2k downloads2mo agoHugging Face02CGIAR /ifpri-ai-documents-markdown GAIA / GARDIAN-CIGI Agricultural Research Corpus This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline. Dataset Overview Property Value Total Documents 21,726 Total Size 623.27 MB Total Tokens 85,359,442 Total Pages 0 Languages 25 Unique Keywords 7,127 Resource Types 20 Date Generated 2026-07-31 02:55:20 Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents-markdown.summarization10K<n<100K1 likes904 downloads2mo agoHugging Face03CGIAR /gardian-cigi-ai-documents-markdown1 likes662 downloads2mo agoHugging Face04CGIAR /peri-kbdocumentn<1K0 likes223 downloads8d agoHugging Face05CGIAR /usda-nal-ai-documents-en-markdown1 likes140 downloads2mo agoHugging Face06CGIAR /gardian-cigi-ai-documentsgated GAIA / GARDIAN-CIGI Agricultural Research Corpus This dataset contains 85,782 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline. Dataset Overview Property Value Total Documents 85,782 Total Size 1.47 GB Total Tokens 199,872,861 Total Pages 0 Languages 59 Unique Keywords 76,373 Resource Types 32 Date Generated 2026-07-16 08:22:33 Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/gardian-cigi-ai-documents.textsummarization10K<n<100K7 likes60 downloads2mo agoHugging Face07CGIAR /ifpri-ai-documentsgated GAIA / GARDIAN-CIGI Agricultural Research Corpus This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline. Dataset Overview Property Value Total Documents 21,726 Total Size 623.27 MB Total Tokens 85,359,442 Total Pages 0 Languages 25 Unique Keywords 7,127 Resource Types 20 Date Generated 2026-07-31 02:55:20 Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents.textsummarization10K<n<100K0 likes54 downloads29d agoHugging Face08CGIAR /gardian-ai-ready-docsgated ⚠️ Heads up: Updated Dataset Available This dataset has been updated with a newer version published on 27 Feb 2025. The latest version includes more updated and refined set of documents. We recommend using the latest version, available at https://huggingface.co/datasets/CGIAR/gardian-cigi-ai-documents. This version remains accessible for reference and reproducibility purposes. A Curated Research Corpus for Agricultural Advisory AI Applications This dataset… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/gardian-ai-ready-docs.textsummarization10K<n<100K2 likes34 downloads8mo agoHugging Face09CGIAR /TranslationDataset_AgriQueriesThis Dataset consists of text pair in english and hindi_latin and some sheets which have hindi_devanagri as well. The dataset is created by different human evaluators who have written hindi sentences in hindi latin (using english alphabets). The dataset can be used for creating a translation model directly from english to hindi latin, which is how the users use ebcause of limitations of mobile keyboard in typing in hindi devanagri script. translation10K<n<100K0 likes33 downloads2y agoHugging Face10CGIAR /embrapa-ai-documentsgated GAIA / GARDIAN-CIGI Agricultural Research Corpus This dataset contains 130,922 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline. Dataset Overview Property Value Total Documents 130,922 Total Size 256.66 MB Total Tokens 11,827,230 Total Pages 0 Languages 8 Unique Keywords 106,492 Resource Types 11 Date Generated 2026-07-16 08:37:30 Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/embrapa-ai-documents.textsummarization10K<n<100K1 likes32 downloads21d agoHugging Face11CGIAR /KikuyuEnglish_translationThis dataset consists of agriculture related sentence pairs in english and kikuyu. The dataset can be used for developing Kikuyu <-> English translation model. translation1K<n<10K1 likes25 downloads2y agoHugging Face12CGIAR /Chi-Metrics-2024 Farmer.Chat: User Interaction and Evaluation Dataset from Agricultural AI Advisory Services in Kenya (2023-2024) AbstractThis dataset comprises 2,300 anonymized user interaction records from Farmer.Chat, an AI-powered agricultural advisory platform deployed in Kenya between October 2023 and April 2024. The data captures diverse metrics including query types, response quality, user engagement patterns, and system performance across multiple agricultural value chains (dairy, coffee… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/Chi-Metrics-2024.question-answering1K<n<10K2 likes24 downloads2y agoHugging Face13CGIAR /AgricultureVideosQnAThe dataset is in XLS format with multiple sheets named for different languages. The dataset is primarily used for training and ground truth of answers that can be generated for agriculture related queries from the videos. Each sheet has list of video urls (youtube links) and the question that can be asked, corresponding answers that can be generated from the videos, source of information in the answer and time stamps. The sources of information could be: Transcript: based on what one hears… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/AgricultureVideosQnA.videoquestion-answering1K<n<10K0 likes24 downloads2y agoHugging Face14CGIAR /AgricultureVideosTranscriptThis dataset consists of agriculture videos in hindi and oriya. The dataset consists of mulitple xls files and each xls file has column of video urls (youtube video links) and corresponding transcripts. The transcripts are: Generated by ASR models (for the purpose of benchmarking) Manual transcripts Time stamps Manual translations This dataset can be used for training and benchmarking domain specific models for ASR and translation. The time stamps serve as the soruce of audio and the… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/AgricultureVideosTranscript.videotranslation1K<n<10K0 likes23 downloads2y agoHugging Face15CGIAR /cirad-ai-documentsgated GAIA / CIRAD Agricultural Documents (English) A curated, machine-readable corpus of 4,209 agricultural research publications sourced from CIRAD (the French Agricultural Research Centre for International Development), produced by the Generative AI for Agriculture (GAIA) project. Documents are indexed through GARDIAN — CGIAR's agri-food research index — and converted from PDF to structured JSON via the GAIA-CIGI pipeline using GROBID. This is an independent CIRAD corpus in the… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/cirad-ai-documents.textsummarization1K<n<10K1 likes23 downloads3mo agoHugging Face16CGIAR /ragas_gardian_evaluation_overlapping 📚 GARDIAN-RAGAS QA Dataset A synthetic question–answer (QA) dataset generated from the GARDIAN corpus using RAGAS and the open-weight Mistral-7B-Instruct-v0.3 model. This dataset is designed to support evaluation and benchmarking of retrieval-augmented generation (RAG) systems, with an emphasis on grounded, high-fidelity QA generation. 📦 Dataset Summary Source Corpus: GARDIAN scientific article collection QA Generation Model: Mistral-7B-Instruct-v0.3 Sample Size: 1… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ragas_gardian_evaluation_overlapping.0 likes12 downloads1y agoHugging Face17CGIAR /RAG-Chunk-Analysis Description The datasets contain human evaluation of retrieved chunks from agriculture documents for actual user queries. Each chunk is marked as relevant and irrelevant. The relevant and irrelevant portion of the chunks are mentioned in a separate columns. The dataset consists of multiple XLS files and each XLS file has multiple sheets corresponding to the content for the value chain. The queries are taken from the actual user questions onf farmer.chat prototype bots. For each user… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/RAG-Chunk-Analysis.question-answering1K<n<10K0 likes11 downloads2y agoHugging Face18CGIAR /AgricultureVideosQnA2The dataset is in XLS format with multiple sheets named for different languages. The dataset is primarily used for training and ground truth of answers that can be generated for agriculture related queries from the videos. Each sheet has list of video urls (youtube links) and the question that can be asked, corresponding answers that can be generated from the videos, source of information in the answer and time stamps. The sources of information could be: Transcript: based on what one hears… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/AgricultureVideosQnA2.question-answering1K<n<10K0 likes11 downloads2y agoHugging Face19CGIAR /aquatic-macroinvertebrate-images Dataset Card for Aquatic Macroinvertebrate Images This image dataset contains 1,300 photos (one hundred images of each of the thirteen mini stream assessment scoring system, or miniSASS aquatic macroinvertebrate groups), used in the development of machine learning models to automatically assess the water quality score. Dataset Description This dataset accompanies the working paper, Digitally improving the identification of aquatic macroinvertebrates for indices used in… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/aquatic-macroinvertebrate-images.imageimage-classification1K<n<10K0 likes8 downloads2y agoHugging Face20CGIAR /KikuyuASR_trainingdatasetThis dataset is obtained as part of AIEP prject by Digital Green and Karya from the extension workers, lead farmers and farmers. Process of collection of data: Selected users were given the option of doing a task and getting paid for it. The users were supposed to record the sentence as it appeared on the screen. The audio file thus obtained was validated matched with the sentences to fine tune the model. Also available are the python script that helps in processing and splitting the data into… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/KikuyuASR_trainingdataset.automatic-speech-recognition10K<n<100K0 likes7 downloads2y agoHugging Face21CGIAR /ragas_gardian_evaluation_non_overlapping 📚 GARDIAN-RAGAS QA Dataset A synthetic question–answer (QA) dataset generated from the GARDIAN corpus using RAGAS and the open-weight Mistral-7B-Instruct-v0.3 model. This dataset is designed to support evaluation and benchmarking of retrieval-augmented generation (RAG) systems, with an emphasis on grounded, high-fidelity QA generation. 📦 Dataset Summary Source Corpus: GARDIAN scientific article collection QA Generation Model: Mistral-7B-Instruct-v0.3 Sample Size: 1… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ragas_gardian_evaluation_non_overlapping.0 likes5 downloads1y agoHugging Face22CGIAR /usda-nal-ai-documents-engated GAIA / GARDIAN-CIGI Agricultural Research Corpus This dataset contains 22,529 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline. Dataset Overview Property Value Total Documents 22,529 Total Size 89.42 MB Total Tokens 5,227,798 Total Pages 0 Languages 1 Unique Keywords 46,728 Resource Types 1 Date Generated 2026-07-16 08:32:49 Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/usda-nal-ai-documents-en.textsummarization10K<n<100K1 likes5 downloads2mo agoHugging Face23CGIAR /cirad-ai-documents-markdown1 likes5 downloads2mo agoHugging Face24CGIAR /ragas_QA_evaluation_datasetgatedThe dataset comprises question answer pairs generated by the Mistral-7B-Instruct-v0.3 model, over a sample of the documents available in the CiGi knowledge base. The generative Q-A creation for evaluating CiGi was preferred over manual annotation because it yields diverse queries whose distribution better reflects downstream user intents. The dataset presentes the question-answer pairs, along with a reference to the source snippet that was used by the model to construct the answer. text1K<n<10K0 likes4 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.