datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
noidea-wisdom-v1convergent-wisdom
Convergent Wisdom: Cross-Tradition Philosophical Alignment Dataset
A dataset for studying the semantic alignment between Eastern (Bhagavad Gita) and Western philosophical traditions using sentence embeddings.
Dataset Description
This dataset contains:
12,902 Bhagavad Gita Q&A pairs — questions about modern life (duty, suffering, purpose, relationships) paired with Gita-inspired answers
333,131 Western philosophy sentences from 36 authors across 13 philosophical schools… See the full description on the dataset page: https://huggingface.co/datasets/asadf1729/convergent-wisdom.wisdombench
WisdomBench and Wisdom Science Data
WisdomBench is a longitudinal benchmark for measuring whether an AI agent changes after repeated exposure to feedback and failure.
This dataset contains 3,600 scored evaluation events under the included conditions (3 models x 4 strategies x 20 tasks x 5 rounds x 3 seeds).
It also mirrors the Wisdom Science Research Portfolio release:
Zenodo record: https://zenodo.org/records/20027295
Portfolio DOI: 10.5281/zenodo.20027295
Portfolio folder:… See the full description on the dataset page: https://huggingface.co/datasets/MMJBDS/wisdombench.wisdom-math
🧙🏼WISDOM
WISDOM: PROGRESSIVE CURRICULUM SYNTHESIS MAKES LLMS BETTER MATHEMATICAL REASONER
🤗Datasets&Models@HF
| 🐱 Code@GitHub
Figure 1: The overall workflow of WISDOM, which leverages Progressive Curriculum Synthesis to generate questions and responses with Deepseek Coder V2 and GPT-4o, including weak teacher guiding, critical expert teaching, experts consistency voting, and hard instruction evolving.
Main Results on the smaller models
MethodBase… See the full description on the dataset page: https://huggingface.co/datasets/Wisdom-math/wisdom-math.founder-wisdom-sft
founder-wisdom-sft
Synthetic English prompt/response pairs for supervised fine-tuning.
Each response is a short decision rule in a direct founder voice (1–4 sentences). Trivia-style questions and calls to action were filtered out.
Rows
14,791 unique pairs
Split
train
Columns
prompt, response
Language
English
from datasets import load_dataset
ds = load_dataset("prathamkode/founder-wisdom-sft", split="train")
print(ds[0])
Intended for research and small… See the full description on the dataset page: https://huggingface.co/datasets/prathamkode/founder-wisdom-sft.convergent-wisdom-pashto
Dataset Card for Convergent Wisdom (Pashto)
This dataset explores the convergence of wisdom traditions, bridging Eastern philosophies (including the Bhagavad Gita) and Western philosophical thought, specifically tailored for the Pashto language.
Uses
This dataset is designed for training and fine-tuning language models to understand and generate philosophical discourse in Pashto, fostering cross-cultural dialogue and semantic analysis of timeless wisdom.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/convergent-wisdom-pashto.Quilt-LLaVA-PretrainQUILT-LLaVA Pretrain Dataset Card
Paper: Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
Paper or resources for more information:
https://quilt-llava.github.io/
Description and Details
The pretraining subset from [QUILT-1M]https://quilt1m.github.io/) for stage 1 of Quilt-LLaVA pretraining. For the accompanying images please use the following link to request images (GDrive)
Dataset date:
QUILT-LLaVA Pretrain was collected… See the full description on the dataset page: https://huggingface.co/datasets/wisdomik/Quilt-LLaVA-Pretrain.Ancient-Indian-Wisdomyann-lecun-wisdompashto-wisdom-1k-plusGRIP_RL_Datapashto-wisdom-1000-plusunfiltered-wisdom-core
Unfiltered Wisdom Core Dataset
A structured dataset of 1,200+ trauma-informed mental health Q&A pairs designed for ethical AI training and research.
Overview
The Unfiltered Wisdom dataset provides anonymized, research-backed mental health question-and-answer pairs covering topics such as:
Anxiety & panic disorders
Depression & mood disorders
PTSD & Complex PTSD (CPTSD)
Attachment & relational trauma
ADHD & neurodivergence
Grief & loss
Boundaries & emotional regulation… See the full description on the dataset page: https://huggingface.co/datasets/unfiltered-wisdom-ai/unfiltered-wisdom-core.GRIP_SFT_Datatw-judicial-wisdom
Dataset Card for tw-judicial-wisdom
tw-judicial-wisdom 是一個來自中華民國司法院「司法智識庫」之法律判決與見解資料集,合計 2,508 筆,已整理為 OpenAI Messages(messages)對話格式,可用於繁體中文法律 LLM 之持續預訓練(CPT)或 SFT 訓練,讓模型學習法院實務見解之論理結構與用語。
Dataset Details
Dataset Description
中華民國司法院之「司法智識庫」(fjudkm.judicial.gov.tw)為司法院整理發布之精選判決與法律見解集合,收錄各級法院具參考價值之案件與論理段落,長期作為實務界與學界引用之來源。本資料集將這些精選判決與見解整理為對話格式,每筆以單輪 messages 儲存,內容保留法院之事實摘要、爭點分析、法律依據與判決結論,便於法律 LLM 學習:
判決書之結構化論理方式;
法律爭點之拆解與援引法條;
精華案件之裁判主文與理由。
Curated by: Liang… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-judicial-wisdom.wisdomADG-Qwen2.5-Alpaca-GPT4typo_correction_alpacaADG-Qwen2.5-CoTADG-LLaMa3-Alpaca-GPT4ADG-LLaMa3-WizardLMBhagvadGita-1ADG-LLaMa3-CoTADG-Qwen2.5-WizardLMWisdomLM-datasetWisdomOfFourwisdom-counselor-data
