datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
peka_persian_knowledge_assessment
PeKA (Persian Knowledge Assessment)
PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics.
For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper.
This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.qulture-general-knowledge-dataset
Qulture General Knowledge Question Dataset
An open dataset containing 16 families of general knowledge questions, for a total of 64 records.
KnowledgeBerg
KnowledgeBerg
KnowledgeBerg is a multilingual benchmark for knowledge-grounded reasoning over enumerated answer sets.
The default configuration is English. Use the dataset viewer's configuration selector to open each language separately.
Data
Each language is stored as one CSV file under data/<Language>/.
Columns:
id
original_id
domain
original_question
original_answer
question_type
implicit_question
options
correct_option
reasoning_depth
knowledge_width
validation_cot
healthcare-disease-knowledge
Disease Symptoms & Treatment Dataset
This dataset contains structured information about 687 diseases and their associated details.It is intended for research, educational, and prototyping purposes in healthcare-related ML/NLP tasks.
Contents
Each row corresponds to one disease, with 17 columns:
disease — Name of the disease
main_link — Reference link
Diagnosis_treatment_link — Link to diagnosis/treatment page
Doctors_departments_link — Relevant medical… See the full description on the dataset page: https://huggingface.co/datasets/harmesh95/healthcare-disease-knowledge.legal-file-handover-knowledge-action-deadline-coherence-risk-v0.1What this dataset does
You receive
handover summary
key facts
strategy
tasks
deadlines
new actions
You decide
coherent
or
incoherent
Daily use
transfer QC
missed deadline prevention
strategy drift detection
sr_spill_knowledgevedic-knowledge-base
Vedic Knowledge Base (Shriyantra OS Architecture)
हा डेटासेट वैदिक ज्ञान, चक्र, मंत्र, पवित्र भूमिती (Sacred Geometry) आणि आधुनिक न्यूरल नेटवर्क आर्किटेक्चर (Artificial Neural Networks) यांना एकत्र जोडणारे एक नाविन्यपूर्ण मॉडेल आहे.
रचना (Structure)
मल्टी-कॉन्फिगरेशन सेटिंग्समुळे तुम्ही प्रत्येक फाईल आता स्वतंत्रपणे लोड करू शकता.
cricketstudio-knowledge-graph
CricketStudio Knowledge Graph — Sample
A sample of the CricketStudio cricket knowledge graph: the Royal Challengers
Bengaluru squad and their batter-vs-bowler matchup relationships, from IPL 2026
plus IPL-career head-to-head data. Every entity links back to its canonical page
on https://players.cricketstudio.ai via canonical_url.
Files / configs
File
Rows
What
nodes.csv / nodes.json
141
entities — id (slug), type, name, canonical_url
edges.csv /… See the full description on the dataset page: https://huggingface.co/datasets/CricketStudio/cricketstudio-knowledge-graph.
