datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
classical-humanities-corpusdiffmodel_humanitieshumanities_runs_1Symmetric-Humanities-Corpus
Symmetric Humanities and Cultural Corpus
A massive, culturally rich dataset hosted on Hugging Face Hub, engineered to achieve strict thematic symmetry across General Culture, Philosophy & Argumentation, and Theology & Ancient Wisdom. This corpus is explicitly stripped of code, mathematics, and formal logic.
Dataset Structure
The dataset is partitioned into three domains:
culture/ - General multilingual cultural and literary texts
philosophy/ - Philosophical… See the full description on the dataset page: https://huggingface.co/datasets/Uhhhghhhhhhh/Symmetric-Humanities-Corpus.TCM_Humanities
Dataset Card for [TCMLM/TCM_Humanities]
This dataset, curated by the Traditional Chinese Medicine Language Model Team, comprises a comprehensive collection of multiple-choice questions (both single and multiple answers) from the Chinese Medical Practitioner Examination. It's designed to aid in understanding and assessing knowledge in Chinese humanities medicine, medical ethics, and legal regulations for physicians.
Dataset Details
Uses
Direct Use… See the full description on the dataset page: https://huggingface.co/datasets/TCMLM/TCM_Humanities.sage-textbook-v2-humanities-ko-v1harmon_nature_humanitiesontolearner-arts_and_humanities
Arts And Humanities Domain Ontologies
Overview
The arts and humanities domain encompasses ontologies that systematically represent and categorize the diverse aspects of human cultural expression, including music, visual arts, historical artifacts, and broader humanistic studies. This domain plays a crucial role in knowledge representation by providing structured frameworks that facilitate the organization, retrieval, and analysis of cultural and artistic information… See the full description on the dataset page: https://huggingface.co/datasets/SciKnowOrg/ontolearner-arts_and_humanities.ultrafine-humanitieshumanities-semantic-consensus-200
Humanities Semantic Consensus 200
Dataset description
Humanities Semantic Consensus 200 is a Chinese, evidence-grounded benchmark
for studying semantic consensus among distributed language-model agents. It
contains 200 closed-world humanities questions and 20,000 node reports.
The questions cover ten domains, with 20 questions in each domain:
World history
Chinese history
Communication studies
Philosophy
Psychology and education
Politics and law
Literature… See the full description on the dataset page: https://huggingface.co/datasets/yyfanfytfyt/humanities-semantic-consensus-200.teach-humanities-v1
Canis.teach Humanities Dataset
Simple synthetic dataset for training Humanities tutoring models.
Project: Canis.teach - Learning that fits.
Subject: Humanities
Generated with: Canis.lab
Format: Simple ID:content pairs
Dataset Structure
{
"id": "unique_identifier",
"content": "tutoring conversation text"
}
This dataset contains educational conversations focused on Humanities topics, designed to teach effective tutoring behavior rather than just providing direct… See the full description on the dataset page: https://huggingface.co/datasets/CanisAI/teach-humanities-v1.humanitiesjmmlu-curated-humanitiesfine-humanitiesmmlu_generate_humanities_f0.0r0.0e0.0-320mmlu_generate_humanities_f0.0r0.0e0.1-320HumanitiesExpertkurtis-v2-humanities-sftmmlu_generate_humanities_f0.0r0.1e0.0-320mmlu_generate_humanities_Qwen2.5-7B-annotune__xld1cp__checkpoint-444usenet-humanitiesjanuspro_humanitiesLumina-gpt2-humanitieshumanities_runs_2humanities-DB
