CoolFace
Datasetpublic

Reza2kn/neyshekar-fa-wimp-teacher-labels

Persian word-importance teacher labels (neyshekar v4) Per-token word-importance soft labels for a DHH-oriented semantic-WER (ACE-style) metric. Teacher: google/gemma-4-31b-it (OpenRouter, ModelRun fp4) + 5 Persian few-shot examples. Source text: shekar-ai/neyshekar-v4-persian-asr-fa (text only, deduplicated). Utterances: 26490 ok (3 format-fails dropped). Scale: 0.0–1.0 per whitespace token (0 = filler/function word, 1 = essential content). Provenance: distilled with no human… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/neyshekar-fa-wimp-teacher-labels.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
2likes15downloads
Dataset Card

Persian word-importance teacher labels (neyshekar v4)

Per-token word-importance soft labels for a DHH-oriented semantic-WER (ACE-style) metric.

  • —Teacher: google/gemma-4-31b-it (OpenRouter, ModelRun fp4) + 5 Persian few-shot examples.
  • —Source text: shekar-ai/neyshekar-v4-persian-asr-fa (text only, deduplicated).
  • —Utterances: 26490 ok (3 format-fails dropped).
  • —Scale: 0.0–1.0 per whitespace token (0 = filler/function word, 1 = essential content).
  • —Provenance: distilled with no human annotation; validated on English DHH gold (Kafle & Huenerfauth LREC-2018): teacher tok-ρ≈0.80 vs human ceiling ~0.84. Persian cross-model agreement (vs gemini-3.5-flash) per-utt ρ≈0.89.

Format (persian_full_scored.jsonl)

{"utt_id", "status":"ok", "n", "score":[per-token 0..1], "var", "words":[...], "text"}

persian_fewshot.jsonl holds the 5 hand-corrected Persian few-shot anchors.