CoolFace
Datasetpublic

Peacockery/neyshekar-v3-asr-aligned

Neyshekar v3 ASR-Aligned This is a repaired subset of Neyshekar v3 for Persian ASR work. The public v3 archive contains real audio and real transcripts, but the downloaded dataset.json filename-to-text mapping does not align for the checked samples. This export keeps only audio clips whose transcript could be recovered by matching multiple ASR hypotheses back to the original Neyshekar transcript pool. It is useful as a curated ASR training/evaluation candidate set, with the… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/neyshekar-v3-asr-aligned.

sourceHugging Facecc0-1.0updated 5mo agoView on Hugging Face
2likes7downloads
9 commits on main
f329efa5mo ago

Fix verification metadata

chikingsley
cb1fb325mo ago

Add selected audio archive

chikingsley
479682f5mo ago

Add rejected ledger

chikingsley
4d08db85mo ago

Add alignment ledger

chikingsley
806fb025mo ago

Add verification metadata

chikingsley
fa8bb2c5mo ago

Add summary metadata

chikingsley
0a9c6bd5mo ago

Add aligned metadata

chikingsley
a24ec5a5mo ago

Add dataset card

chikingsley
c8615f65mo ago

initial commit

chikingsley