datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EarthScience-Text-LLM-20K-90-10
EarthScience-Text-LLM-20K-90-10
This is a pure-text Earth-science corpus unified from three non-overlapping upstream datasets:
Ekimetrics/climateqa-ipcc-ipbes-reports-1.0: climate and IPCC/IPBES report chunks.
GeoGPT-Research-Project/GeoGPT-CoT-QA: geoscience question-answer reasoning.
gremlin97/RemoteSensingCorpus: remote-sensing and geospatial machine-learning text.
Files and Split
The previous preprocessing outputs were merged into a 23,098-record pool and… See the full description on the dataset page: https://huggingface.co/datasets/moTcream/EarthScience-Text-LLM-20K-90-10.Gitruck-MotionIR
Gitruck MotionIR
Gitruck MotionIR is a Chinese motion-design dataset that aligns project-level
natural-language descriptions, technique-level annotations, temporal evidence,
and a renderable intermediate representation (IR v1). The corpus was normalized
from authorized Alight Motion, After Effects, NodeVideo, and Jianying projects.
Gitruck MotionIR 是一个中文动效设计数据集,将工程级描述、技法级标注、时间证据与可渲染
IR v1 对齐。语料由已获授权的 Alight Motion、After Effects、NodeVideo 与剪映工程归一化而来。
Dataset summary /… See the full description on the dataset page: https://huggingface.co/datasets/Hocassian/Gitruck-MotionIR.Motivation_Employee_Engagement_Content_1
Motivation Employee Engagement Content 1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Motivation_Employee_Engagement_Content_1.Motivation_Employee_Engagement_Content_2
Motivation Employee Engagement Content 2
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Motivation_Employee_Engagement_Content_2.ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer-drop-motadarak
Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens, max 20 bayts, drop مجزوء الوافر, drop المتدارك)
This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer and keeps the same schema, columns, locked prompt format, and general structure as the upstream phase-1 dataset.
The only additional change is the removal of rows where:
base_meter == "المتدارك"
This removes all poems whose base meter is المتدارك from the published… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer-drop-motadarak.
