Reza2kn/neyshekar-fa-wimp-teacher-labels
Persian word-importance teacher labels (neyshekar v4) Per-token word-importance soft labels for a DHH-oriented semantic-WER (ACE-style) metric. Teacher: google/gemma-4-31b-it (OpenRouter, ModelRun fp4) + 5 Persian few-shot examples. Source text: shekar-ai/neyshekar-v4-persian-asr-fa (text only, deduplicated). Utterances: 26490 ok (3 format-fails dropped). Scale: 0.0–1.0 per whitespace token (0 = filler/function word, 1 = essential content). Provenance: distilled with no human… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/neyshekar-fa-wimp-teacher-labels.
Persian word-importance teacher labels (neyshekar v4)
Per-token word-importance soft labels for a DHH-oriented semantic-WER (ACE-style) metric.
- Teacher:
google/gemma-4-31b-it(OpenRouter, ModelRun fp4) + 5 Persian few-shot examples. - Source text:
shekar-ai/neyshekar-v4-persian-asr-fa(text only, deduplicated). - Utterances: 26490 ok (3 format-fails dropped).
- Scale: 0.0–1.0 per whitespace token (0 = filler/function word, 1 = essential content).
- Provenance: distilled with no human annotation; validated on English DHH gold (Kafle & Huenerfauth LREC-2018): teacher tok-ρ≈0.80 vs human ceiling ~0.84. Persian cross-model agreement (vs gemini-3.5-flash) per-utt ρ≈0.89.
Format (persian_full_scored.jsonl)
{"utt_id", "status":"ok", "n", "score":[per-token 0..1], "var", "words":[...], "text"}
persian_fewshot.jsonl holds the 5 hand-corrected Persian few-shot anchors.
