Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer-drop-motadarak
Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens, max 20 bayts, drop مجزوء الوافر, drop المتدارك) This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer and keeps the same schema, columns, locked prompt format, and general structure as the upstream phase-1 dataset. The only additional change is the removal of rows where: base_meter == "المتدارك" This removes all poems whose base meter is المتدارك from the published… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer-drop-motadarak.
Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens, max 20 bayts, drop مجزوء الوافر, drop المتدارك)
This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer and keeps the same schema, columns, locked prompt format, and general structure as the upstream phase-1 dataset.
The only additional change is the removal of rows where:
base_meter == "المتدارك"
This removes all poems whose base meter is المتدارك from the published phase-1 subset.
Why this variant exists
The phase-1 meter reward is explicitly base-meter-first. A dedicated base-only sanity study with n=5 prompts per base meter and k=5 sampled candidates per prompt showed that المتدارك remained a clear outlier even under the more GRPO-relevant best-of-k view.
In that study:
المتداركhad very weak candidate-level scoresالمتداركalso failed the prompt-level best-of-kcriterion- no prompt produced a candidate that crossed the useful quality thresholds used for review
Because the goal of this phase-1 subset is to align the training data with the available meter reward and keep RL signal reasonably clean, this derivative removes المتدارك from the current phase-1 dataset.
Locked Prompt
SYSTEM_PROMPT
أنت شاعر عربي تكتب الشعر العمودي الكلاسيكي. التزم بالبحر المحدد في كل شطر، واستلهم من الموضوع دون نقله حرفياً. أخرج الأبيات فقط دون مقدمة أو تعليق.
USER_TEMPLATE
البحر الأساسي: {basemeter} الصيغة: {form} اسم البحر المطلوب: {meterlabel} الموضوع: {description}
اكتب {numlines} شطراً ملتزماً بصيغة {form} من بحر {basemeter} دون أي شرح إضافي.
Conditioning Rule
meter_label = base_meterifform == "تام"- else
meter_label = "{form} {base_meter}"
Added Columns
sft_promptsft_completionsft_full_textsft_num_linessft_total_tokens
Target Formatting
sft_completionis built frompoem versesusing real newline characters.
Filtering
- Inherits all filtering already present in
Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer. - Additional filter applied here: drop rows where
base_meter == "المتدارك".
Counts
- Source rows from
Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer: 117624 - Removed by dropping
المتدارك: 138 - Final rows kept: 117486
- Retention from upstream phase-1 subset: 99.88%
Training-Prep Note
- After the repo's current
load_and_prepare_dataset(...)preprocessing, the upstream phase-1 dataset yields 117404 usable rows. - This derivative yields 117266 usable rows after the same preparation path.
Important Meter-Reward Caveat
- The current meter reward is primarily a base-meter correctness signal.
- Form-sensitive meter realization is not directly validated when the classifier lacks that exact form label.
- This dataset change should therefore be understood as alignment with a base-meter-first reward, not as a claim that the dropped meter is impossible in general.
Notes
- This is a derived dataset repo; upstream datasets are unchanged.
- The schema and column names are kept identical to the upstream dataset.
- This subset is intended for phase-1 GRPO experiments where the active meter-reward signal is more reliable on the retained base-meter set.
