datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kcc-krishi-punjab-wheat-demo
KCC-Krishi Punjab/Wheat GPT-5.5 Normalization Demo
This dataset is a 500-row Punjab/Wheat demonstration subset from the KCC-Krishi data preparation pipeline.
It shows how terse Kisan Call Centre expert evidence can be converted into clearer, farmer-friendly, voice-ready advisory text for downstream RAG, human review, and offline farmer-assistant prototyping.
Important safety warning: The normalized advisory text is machine-generated and is not human-verified agronomic ground… See the full description on the dataset page: https://huggingface.co/datasets/uralstech/kcc-krishi-punjab-wheat-demo.wheat-kcc-queries
Dataset Card
Dataset Description
This dataset contains parts of data from Kisan Call Center Transcripts. These are related to farmer transcripts/queries related to the wheat crop
Citation
If you use this dataset, please cite the original FLEURS paper:
@article{fleurs2022arxiv,
title = {FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech},
author = {Conneau, Alexis and Ma, Min and Khanuja, Simran and Zhang, Yu and Axelrod, Vera and… See the full description on the dataset page: https://huggingface.co/datasets/skylord/wheat-kcc-queries.pashto_wheat_sft_14k
🌾 Pashto Wheat SFT 14K Dataset
📌 Introduction
Pashto Wheat SFT 14K is a high-quality, lightweight Supervised Fine-Tuning (SFT) dataset tailored specifically for training native Pashto AI models. The dataset comprises roughly 14,000 short, concise, and highly informative instructional examples natively adapted to cover essential tech concepts, programming languages, and general knowledge structures.
The philosophy behind this dataset reflects a local Pashto… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto_wheat_sft_14k.
