datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2d_3d_seq_path_spatial_reasoning
Spatial Reasoning Dataset
A synthetic dataset of Hamiltonian path puzzles with rich chain-of-thought reasoning, designed for training and evaluating spatial reasoning in language models.
Overview
Each sample presents a grid-based puzzle where the solver must find a path visiting every cell exactly once, moving only up/down/left/right (plus above/below for 3D). Puzzles span 2D grids (3x3 to 8x8) and 3D cubes (3x3x3 to 4x4x4), covering solvable, impossible, and multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/eousphoros/2d_3d_seq_path_spatial_reasoning.K-Paths-inductive-reasoning-drugbank
🔗 This dataset is part of the study:
K-Paths: Reasoning over Graph Paths for Drug Repurposing and Drug Interaction Prediction
📖 Read the Paper
💾 GitHub Repository
DrugBank: Inductive Reasoning Dataset
This dataset contains drug pairs annotated with 86 pharmacological relationships (e.g.,DrugA may increase the anticholinergic activities of DrugB).
Each entry includes two drugs, an interaction label, drug descriptions, and structured/natural language representations… See the full description on the dataset page: https://huggingface.co/datasets/Tassy24/K-Paths-inductive-reasoning-drugbank.K-Paths-inductive-reasoning-pharmaDB
🔗 This dataset is part of the study:
K-Paths: Reasoning over Graph Paths for Drug Repurposing and Drug Interaction Prediction
📖 Read the Paper
💾 GitHub Repository
PharmacotherapyDB: Inductive Reasoning Dataset
PharmacotherapyDB is a drug repurposing dataset containing drug–disease treatment relations in three categories (disease-modifying, palliates, or non-indication).
Each entry includes a drug and a disease, an interaction label, drug, disease descriptions, and… See the full description on the dataset page: https://huggingface.co/datasets/Tassy24/K-Paths-inductive-reasoning-pharmaDB.K-Paths-inductive-reasoning-ddinter
🔗 This dataset is part of the study:
K-Paths: Reasoning over Graph Paths for Drug Repurposing and Drug Interaction Prediction
📖 Read the Paper
💾 GitHub Repository
DDInter: Inductive Reasoning Dataset
DDInter provides drug–drug interaction (DDI) data labeled with three severity levels (Major, Moderate, Minor).
Each entry includes two drugs, an interaction label, drug descriptions, and structured/natural language representations of multi-hop reasoning paths between… See the full description on the dataset page: https://huggingface.co/datasets/Tassy24/K-Paths-inductive-reasoning-ddinter.pathinen_keezhkanakku-elathi
📚 Dataset Card: Elathi (ஏலாதி)
Dataset Summary
Elathi (ஏலாதி) is a classical Tamil ethical text and one of the Pathinen Keezhkanakku (Eighteen Minor Works). It conveys moral and ethical guidelines through poetic verses, inspired by the structure of the traditional medicinal formulation known as Elathi, which consists of six ingredients. Each verse metaphorically reflects six core virtues essential for righteous living.
Title: Elathi (ஏலாதி)
Text Type: Ethical Poetry /… See the full description on the dataset page: https://huggingface.co/datasets/TamilThagaval/pathinen_keezhkanakku-elathi.pathinen_keezhkanakku-kaarnarpadhu
📚 Dataset Card: கார் நாற்பது (Kaarnarpadhu)
Dataset Summary
கார் நாற்பது (Kaarnarpadhu) is a classical Tamil poetic work belonging to the Pathinen Keezhkanakku tradition. The text derives its name from two defining characteristics:
It consists of 40 poems (நாற்பது செய்யுட்கள்)
Each poem describes the arrival and nature of the monsoon season (கார் காலம்)
Thus, the work came to be known as Kaar Narpadhu.
Title: கார் நாற்பது
Text Type: Seasonal & Emotional Poetry… See the full description on the dataset page: https://huggingface.co/datasets/TamilThagaval/pathinen_keezhkanakku-kaarnarpadhu.pathology-1k
Pathology Medical Dataset — 1,000 Record Free Sample
Enterprise-grade synthetic medical data. Zero PHI. HIPAA-Aligned.
Quality Metrics
Metric
Score
Industry Benchmark
Trinity Consensus Score (TAS)
98.0%
85-92% typical
Trinity Assurance Score (TAS)
0.97
0.75-0.85 typical
Macro F1
0.97
0.80-0.90 typical
PHI Present
None
--
Generation Method
3-LLM Trinity Ensemble
Single model typical
What's Included (Free)
1,000… See the full description on the dataset page: https://huggingface.co/datasets/WitnessDataFactory/pathology-1k.
