datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KannadaPromptBench
KannadaPromptBench
A benchmark dataset for evaluating prompt strategy sensitivity in Kannada, a low-resource Dravidian language.
Dataset Summary
Language: Kannada (kn)
Tasks: Sentiment Analysis (100), Question Answering (75), Summarization (50)
Total: 225 culturally grounded samples
Inter-annotator agreement: Cohen's κ > 0.80
Dataset Structure
Each sample contains: id, task, input_text, label, difficulty, domain.
Citation
Please… See the full description on the dataset page: https://huggingface.co/datasets/Anushhh/KannadaPromptBench.kanna-rag-gold-standard
Kanna RAG Gold Standard Dataset
This dataset contains 30 expert-curated Question-Answer pairs focused on the ethnopharmacology of Sceletium tortuosum (Kanna). It serves as the "Gold Standard" evaluation set for the LAYRA (Large Academic Visual RAG Agent) thesis project.
Dataset Structure
query: The scientific question.
doc_id: The unique identifier of the source document (PDF).
page_num: The specific page number where the answer is found (critical for Visual RAG).… See the full description on the dataset page: https://huggingface.co/datasets/SAINTHALF/kanna-rag-gold-standard.
