datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
convergent-wisdom
Convergent Wisdom: Cross-Tradition Philosophical Alignment Dataset
A dataset for studying the semantic alignment between Eastern (Bhagavad Gita) and Western philosophical traditions using sentence embeddings.
Dataset Description
This dataset contains:
12,902 Bhagavad Gita Q&A pairs — questions about modern life (duty, suffering, purpose, relationships) paired with Gita-inspired answers
333,131 Western philosophy sentences from 36 authors across 13 philosophical schools… See the full description on the dataset page: https://huggingface.co/datasets/asadf1729/convergent-wisdom.Quilt_VQA
Dataset Card for "Quilt_VQA"
Paper: Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
Paper or resources for more information:
https://quilt-llava.github.io/
Description and Details
To evaluate Quilt-LLaVA, alongside public VQA pathology datasets, we also generated Quilt-VQA by extracting Q&A dataset from naturally occurring questions/answers given in the videos. With the help of GPT4 and some handcrafted… See the full description on the dataset page: https://huggingface.co/datasets/wisdomik/Quilt_VQA.wisdom-math
🧙🏼WISDOM
WISDOM: PROGRESSIVE CURRICULUM SYNTHESIS MAKES LLMS BETTER MATHEMATICAL REASONER
🤗Datasets&Models@HF
| 🐱 Code@GitHub
Figure 1: The overall workflow of WISDOM, which leverages Progressive Curriculum Synthesis to generate questions and responses with Deepseek Coder V2 and GPT-4o, including weak teacher guiding, critical expert teaching, experts consistency voting, and hard instruction evolving.
Main Results on the smaller models
MethodBase… See the full description on the dataset page: https://huggingface.co/datasets/Wisdom-math/wisdom-math.QUILT-LLaVA-Instruct-107KQUILT-LLaVA Visual Instruct 107K Dataset Card
Paper: Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
Paper or resources for more information:
https://quilt-llava.github.io/
Description and Details
YouTube educational histopathology videos are a valuable source of grounded histopathology data for instructional purposes, particularly for visual instruction tuning.
Similar to LLaVA, the approach involves using independent… See the full description on the dataset page: https://huggingface.co/datasets/wisdomik/QUILT-LLaVA-Instruct-107K.wisdombench
WisdomBench and Wisdom Science Data
WisdomBench is a longitudinal benchmark for measuring whether an AI agent changes after repeated exposure to feedback and failure.
This dataset contains 3,600 scored evaluation events under the included conditions (3 models x 4 strategies x 20 tasks x 5 rounds x 3 seeds).
It also mirrors the Wisdom Science Research Portfolio release:
Zenodo record: https://zenodo.org/records/20027295
Portfolio DOI: 10.5281/zenodo.20027295
Portfolio folder:… See the full description on the dataset page: https://huggingface.co/datasets/MMJBDS/wisdombench.qwen35-2b-personal-training-data
Qwen3.5-2B Three-Domain Training Data
A reproducible training-data release assembled and processed by wisdompan
for Qwen3.5-2B experiments across mathematics, code, and instruction following.
Dataset configurations
Configuration
Purpose
Train rows
Validation rows
full_mix
Unified three-domain student training
86,931
3
teacher_math
Mathematics teacher training
17,917
1
teacher_code
Code teacher training
23,667
1
teacher_if
Instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/wisdompan/qwen35-2b-personal-training-data.founder-wisdom-sft
founder-wisdom-sft
Synthetic English prompt/response pairs for supervised fine-tuning.
Each response is a short decision rule in a direct founder voice (1–4 sentences). Trivia-style questions and calls to action were filtered out.
Rows
14,791 unique pairs
Split
train
Columns
prompt, response
Language
English
from datasets import load_dataset
ds = load_dataset("prathamkode/founder-wisdom-sft", split="train")
print(ds[0])
Intended for research and small… See the full description on the dataset page: https://huggingface.co/datasets/prathamkode/founder-wisdom-sft.MedicalNarrativesMulti-modal models are data hungry. While datasets with natural images are abundant, medical image datasets can not afford the same luxury.
To enable representation learning for medical images at scale, we turn to YouTube, a platform with a large reservoir of open-source medical pedagogical videos.
We curate MedicalNarratives, a dataset 4.7M medical image-text pairs, with 1M samples containing dense annotations in the form of traces and bounding boxes.
Similar to \textit{think-aloud} studies… See the full description on the dataset page: https://huggingface.co/datasets/wisdomik/MedicalNarratives.convergent-wisdom-pashto
Dataset Card for Convergent Wisdom (Pashto)
This dataset explores the convergence of wisdom traditions, bridging Eastern philosophies (including the Bhagavad Gita) and Western philosophical thought, specifically tailored for the Pashto language.
Uses
This dataset is designed for training and fine-tuning language models to understand and generate philosophical discourse in Pashto, fostering cross-cultural dialogue and semantic analysis of timeless wisdom.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/convergent-wisdom-pashto.Quilt-LLaVA-PretrainQUILT-LLaVA Pretrain Dataset Card
Paper: Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
Paper or resources for more information:
https://quilt-llava.github.io/
Description and Details
The pretraining subset from [QUILT-1M]https://quilt1m.github.io/) for stage 1 of Quilt-LLaVA pretraining. For the accompanying images please use the following link to request images (GDrive)
Dataset date:
QUILT-LLaVA Pretrain was collected… See the full description on the dataset page: https://huggingface.co/datasets/wisdomik/Quilt-LLaVA-Pretrain.Ancient-Indian-Wisdomnoidea-wisdom-v1yann-lecun-wisdomwisdom-spark-philosophical-wisdom
Wisdom Spark AI - Philosophical Wisdom Corpus
Curated wisdom from 17 philosophical traditions, structured for AI training
and alignment. Each entry includes source text, extracted principles, practical
applications, modern context, cross-tradition themes, and flourishing dimension scores.
Purpose
Feed AI models the distilled wisdom of 5,000+ years of human philosophy to promote
ethical reasoning, compassion, cross-cultural understanding, and human flourishing.… See the full description on the dataset page: https://huggingface.co/datasets/kiranz38/wisdom-spark-philosophical-wisdom.pashto-wisdom-1k-pluspashto-wisdom-1000-plusislamic-wisdom-pagesGRIP_RL_Dataunfiltered-wisdom-core
Unfiltered Wisdom Core Dataset
A structured dataset of 1,200+ trauma-informed mental health Q&A pairs designed for ethical AI training and research.
Overview
The Unfiltered Wisdom dataset provides anonymized, research-backed mental health question-and-answer pairs covering topics such as:
Anxiety & panic disorders
Depression & mood disorders
PTSD & Complex PTSD (CPTSD)
Attachment & relational trauma
ADHD & neurodivergence
Grief & loss
Boundaries & emotional regulation… See the full description on the dataset page: https://huggingface.co/datasets/unfiltered-wisdom-ai/unfiltered-wisdom-core.relatedness
Dataset Card for "relatedness"
More Information needed
QuiltVQA_REDDataset Card for "QuiltVQA_ALL"
Human Generated VQA Dataset for Evaluation
Quilt-VQA
is generated by extracting Q&A dataset from naturally occurring questions/answers given in educational histopathology videos. With the help of GPT4 and some handcrafted algorithms, we collect a rich evaluation dataset of 1283 Q&A pairs.
Top two rows show image-dependent Q&A pairs and bottom two rows show general-knowledge Q&A pairs. The original question posed by the narrator of the video is highlighted… See the full description on the dataset page: https://huggingface.co/datasets/wisdomik/QuiltVQA_RED.Historical_Wisdomtw-judicial-wisdom
Dataset Card for tw-judicial-wisdom
tw-judicial-wisdom 是一個來自中華民國司法院「司法智識庫」之法律判決與見解資料集,合計 2,508 筆,已整理為 OpenAI Messages(messages)對話格式,可用於繁體中文法律 LLM 之持續預訓練(CPT)或 SFT 訓練,讓模型學習法院實務見解之論理結構與用語。
Dataset Details
Dataset Description
中華民國司法院之「司法智識庫」(fjudkm.judicial.gov.tw)為司法院整理發布之精選判決與法律見解集合,收錄各級法院具參考價值之案件與論理段落,長期作為實務界與學界引用之來源。本資料集將這些精選判決與見解整理為對話格式,每筆以單輪 messages 儲存,內容保留法院之事實摘要、爭點分析、法律依據與判決結論,便於法律 LLM 學習:
判決書之結構化論理方式;
法律爭點之拆解與援引法條;
精華案件之裁判主文與理由。
Curated by: Liang… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-judicial-wisdom.wisdomInterrogatory-R1
智海-录问 推理数据(wisdomInterrogatory-R1)
类别
数据量
任务描述
罪名预测
20k
请作为中国法官,基于案件事实,对被告人进行单一罪名预测。返回格式:“罪名”.示例:“盗窃”,“敲诈勒索”。
刑期预测
5k
请作为中国法官,基于案件事实,对被告人进行量刑预测,请遵从下列规则,返回结果:- 如果量刑为有期徒刑,请返回:“量刑月数”,例如应判决5年,5年=60月,则返回60.- 如果量刑为无期徒刑,请返回:“life_imprisonment”.- 如果量刑为死刑,请返回:“death_penalty”。
论辩挖掘
4k
在法院的庭审过程中,诉方与辩方由于立场观点或事实陈述的差异,会形成庭审争议焦点,这是整场庭审的核心环节。这种辩方与诉方之间形成的逻辑互动论点对,即为争议焦点。请作为中国的辩护律师,执行“论辩挖掘”任务,即根据提供的被告罪名和诉方论点,从五个候选辩护论点中,选择一个最适合作为与诉方观点形成互动对的论点。需要特别说明的是,争议焦点的对抗,始终基于事实基础。返回格式:“编号”.示例:“1… See the full description on the dataset page: https://huggingface.co/datasets/YinghaoHu/wisdomInterrogatory-R1.GRIP_SFT_Datacleanup
Dataset Card for "cleanup"
More Information needed
typo_correction
Dataset Card for "typo_correction"
More Information needed
storyThis dataset is designed to provide forward and reverse dictionary of Korean proverbs.ADG-Qwen2.5-Alpaca-GPT4wisdom
