CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Shaer-AI /ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits Ashaar Enhanced Description SFT Stratified Splits Source dataset: Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500 Target dataset: Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits This dataset publishes deterministic train / eval / test splits with a 94 / 3 / 3 policy. Split policy Primary stratification key: base_meter form length_bucket Length buckets: 1-3 4-6 7-10 11-20 Small groups fall back… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits.tabulartext-generation100K<n<1M0 likes67 downloads8d agoHugging Face02Shaer-AI /ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500 Ashaar Final SFT Dataset with Enhanced Descriptions This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed and is intended to be the final SFT-ready dataset we continue working with. We got here the hard way. GRPO did not deliver a convincing improvement. Continuation SFT degraded. A fresh-from-zero SFT direction still exposed a deeper data problem. After inspecting the conditioning text, we concluded that many of the old descriptions were weak or… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500.tabulartext-generation100K<n<1M0 likes44 downloads8d agoHugging Face03Shaer-AI-2 /ashaar-with-descriptions-baseform-final-trimmed Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens) This dataset is derived from Shaer-AI/ashaar-v1-base-form-with-descriptions and prepared for supervised fine-tuning (SFT) with the locked prompt from prompt_testing.ipynb. Locked Prompt SYSTEM_PROMPT أنت شاعر عربي تكتب الشعر العمودي الكلاسيكي. التزم بالبحر المحدد في كل شطر، واستلهم من الموضوع دون نقله حرفياً. أخرج الأبيات فقط دون مقدمة أو تعليق. USER_TEMPLATE البحر الأساسي:… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed.tabulartext-generation100K<n<1M0 likes16 downloads7mo agoHugging Face04Shaer-AI-2 /ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens, max 20 bayts, drop مجزوء الوافر) This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed-maxlen20 and keeps the same schema, columns, locked prompt format, and general structure as the upstream phase-1 dataset. The only additional change is the removal of rows where: base_meter == "الوافر" form == "مجزوء" This removes poems labeled مجزوء الوافر from the published phase-1 subset. Why… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer.tabulartext-generation100K<n<1M0 likes15 downloads6mo agoHugging Face05Shaer-AI-2 /ashaar-with-descriptions-baseform-final-trimmed-maxlen20 Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens, max 20 bayts) This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed and prepared as a phase-1 GRPO subset by applying a maximum poem length filter of 20 complete bayts while keeping the same columns, prompt format, and general dataset structure as the source dataset. Locked Prompt SYSTEM_PROMPT أنت شاعر عربي تكتب الشعر العمودي الكلاسيكي. التزم بالبحر المحدد في كل… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed-maxlen20.tabulartext-generation100K<n<1M0 likes8 downloads6mo agoHugging Face06Shaer-AI-2 /ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer-drop-motadarak Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens, max 20 bayts, drop مجزوء الوافر, drop المتدارك) This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer and keeps the same schema, columns, locked prompt format, and general structure as the upstream phase-1 dataset. The only additional change is the removal of rows where: base_meter == "المتدارك" This removes all poems whose base meter is المتدارك from the published… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer-drop-motadarak.tabulartext-generation100K<n<1M0 likes6 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.