datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Appreciation-of-Chinese-Classical-Poetry
Appreciation of Chinese Classical Poetry
Chinese classical poetry with paired metadata and five-aspect literary analyses for PoetryMTEB / MTEB-style evaluation and computational poetics research.
Poems are drawn from expert appreciation volumes (mainly Shanghai Lexicographical Publishing House dictionaries). The released analysis fields are LLM distillations (DeepSeek-V3.1) of those expert appreciation texts into five free-text facets. The original long-form appreciation prose… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/Appreciation-of-Chinese-Classical-Poetry.Classical-Mechanics-Equations-Dataset_SFT-or-LoRA
Classical Mechanics Equations Dataset (SFT / LoRA Ready)
A structured dataset of 64 classical mechanics equations from Newtonian,
Lagrangian, and Hamiltonian mechanics, expanded into 448 instruction-tuning
rows across three task types: equation explanation, Q&A, and derivation.
Designed for fine-tuning LLMs on physics reasoning, STEM Q&A, and
equation understanding tasks.
Overview
Property
Value
Domain
Classical Mechanics (Physics)
Total rows
448
Train… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/Classical-Mechanics-Equations-Dataset_SFT-or-LoRA.military-strategy-classics
Analytical Decision Frameworks — Public Domain Dataset
Structured public domain texts on decision-making, organizational design, and strategic analysis.
Formatted for AI training, analysis, and agent tool use.
All content sourced from works in the public domain (published before 1928, or government-authored).
Content Domains
Strategic planning principles
Organizational coordination patterns
Decision frameworks under uncertainty
Historical pattern analysis… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/military-strategy-classics.clean_squad_classic_v1
Clean SQuAD Classic v1
This is a refined version of the SQuAD v1 dataset. It has been preprocessed to ensure higher data quality and usability for NLP tasks such as Question Answering.
Description
The Clean SQuAD Classic v1 dataset was created by applying preprocessing steps to the original SQuAD v1 dataset, including:
Trimming whitespace: All leading and trailing spaces have been removed from the question field.
Minimum question length: Questions with fewer than 12… See the full description on the dataset page: https://huggingface.co/datasets/decodingchris/clean_squad_classic_v1.clean_squad_classic_v2
Clean SQuAD Classic v2
This is a refined version of the SQuAD v2 dataset. It has been preprocessed to ensure higher data quality and usability for NLP tasks such as Question Answering.
Description
The Clean SQuAD Classic v2 dataset was created by applying preprocessing steps to the original SQuAD v2 dataset, including:
Trimming whitespace: All leading and trailing spaces have been removed from the question field.
Minimum question length: Questions with fewer than 12… See the full description on the dataset page: https://huggingface.co/datasets/decodingchris/clean_squad_classic_v2.
