Peacockery/neyshekar-v3-asr-aligned
Neyshekar v3 ASR-Aligned This is a repaired subset of Neyshekar v3 for Persian ASR work. The public v3 archive contains real audio and real transcripts, but the downloaded dataset.json filename-to-text mapping does not align for the checked samples. This export keeps only audio clips whose transcript could be recovered by matching multiple ASR hypotheses back to the original Neyshekar transcript pool. It is useful as a curated ASR training/evaluation candidate set, with the… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/neyshekar-v3-asr-aligned.
Neyshekar v3 ASR-Aligned
This is a repaired subset of Neyshekar v3 for Persian ASR work. The public v3 archive contains real audio and real transcripts, but the downloaded dataset.json filename-to-text mapping does not align for the checked samples.
This export keeps only audio clips whose transcript could be recovered by matching multiple ASR hypotheses back to the original Neyshekar transcript pool. It is useful as a curated ASR training/evaluation candidate set, with the alignment ledger included for audit.
- Hub repo:
Peacockery/neyshekar-v3-asr-aligned - Source: https://zenodo.org/records/19186714
- Source rows checked: 30019
- Accepted aligned rows: 26173
- Acceptance rate: 87.19%
Files
dataset.json: list of accepted rows withid,audio,text,duration.audio.zip: selected WAV files underaudio/.alignment.jsonl: accepted rows with model hypotheses and alignment scores.rejected.jsonl: rows that did not pass the high-confidence alignment gate.summary.json: thresholds and aggregate counts.verification.json: SHA-256 checksums.
Alignment Gate
Rows are accepted when at least 2 model hypotheses match the same normalized transcript within CER 0.2, the best CER is at most 0.1, and the candidate margin is at least 0.04.
