CoolFace
Datasetpublic

Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer-drop-motadarak

Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens, max 20 bayts, drop مجزوء الوافر, drop المتدارك) This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer and keeps the same schema, columns, locked prompt format, and general structure as the upstream phase-1 dataset. The only additional change is the removal of rows where: base_meter == "المتدارك" This removes all poems whose base meter is المتدارك from the published… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer-drop-motadarak.

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes3downloads
Dataset Card

Ashaar v1 SFT-Ready (Locked Prompt, <= 2048 tokens, max 20 bayts, drop مجزوء الوافر, drop المتدارك)

This dataset is derived from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer and keeps the same schema, columns, locked prompt format, and general structure as the upstream phase-1 dataset.

The only additional change is the removal of rows where:

  • —base_meter == "المتدارك"

This removes all poems whose base meter is المتدارك from the published phase-1 subset.

Why this variant exists

The phase-1 meter reward is explicitly base-meter-first. A dedicated base-only sanity study with n=5 prompts per base meter and k=5 sampled candidates per prompt showed that المتدارك remained a clear outlier even under the more GRPO-relevant best-of-k view.

In that study:

  • —المتدارك had very weak candidate-level scores
  • —المتدارك also failed the prompt-level best-of-k criterion
  • —no prompt produced a candidate that crossed the useful quality thresholds used for review

Because the goal of this phase-1 subset is to align the training data with the available meter reward and keep RL signal reasonably clean, this derivative removes المتدارك from the current phase-1 dataset.

Locked Prompt

SYSTEM_PROMPT

أنت شاعر عربي تكتب الشعر العمودي الكلاسيكي. التزم بالبحر المحدد في كل شطر، واستلهم من الموضوع دون نقله حرفياً. أخرج الأبيات فقط دون مقدمة أو تعليق.

USER_TEMPLATE

البحر الأساسي: {basemeter} الصيغة: {form} اسم البحر المطلوب: {meterlabel} الموضوع: {description}

اكتب {numlines} شطراً ملتزماً بصيغة {form} من بحر {basemeter} دون أي شرح إضافي.

Conditioning Rule

  • —meter_label = base_meter if form == "تام"
  • —else meter_label = "{form} {base_meter}"

Added Columns

  • —sft_prompt
  • —sft_completion
  • —sft_full_text
  • —sft_num_lines
  • —sft_total_tokens

Target Formatting

  • —sft_completion is built from poem verses using real newline characters.

Filtering

  • —Inherits all filtering already present in Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer.
  • —Additional filter applied here: drop rows where base_meter == "المتدارك".

Counts

  • —Source rows from Shaer-AI/ashaar-with-descriptions-baseform-final-trimmed-maxlen20-drop-majzuu-wafer: 117624
  • —Removed by dropping المتدارك: 138
  • —Final rows kept: 117486
  • —Retention from upstream phase-1 subset: 99.88%

Training-Prep Note

  • —After the repo's current load_and_prepare_dataset(...) preprocessing, the upstream phase-1 dataset yields 117404 usable rows.
  • —This derivative yields 117266 usable rows after the same preparation path.

Important Meter-Reward Caveat

  • —The current meter reward is primarily a base-meter correctness signal.
  • —Form-sensitive meter realization is not directly validated when the classifier lacks that exact form label.
  • —This dataset change should therefore be understood as alignment with a base-meter-first reward, not as a claim that the dropped meter is impossible in general.

Notes

  • —This is a derived dataset repo; upstream datasets are unchanged.
  • —The schema and column names are kept identical to the upstream dataset.
  • —This subset is intended for phase-1 GRPO experiments where the active meter-reward signal is more reliable on the retained base-meter set.