Reza2kn/nasle-mana-clean-chunked-30s-avasanj
Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here. Splits Split Rows Audio Columns labeled 4,981 41.41 hours audio, label to_transcribe 11,127 92.72 hours audio The labeled split contains the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s-avasanj.
Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks
Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here.
Splits
The labeled split contains the text-linked recordings that passed the text/audio quality gates. These recordings were read by the same female narrator according to the source grouping; narrator identity was not independently verified. The to_transcribe split contains playable audio without labels and may contain additional speakers.
Construction and quality gates
- 1,512 source records were considered; 1,021 were materialized and 491 were excluded.
- Target chunk length was 30 seconds; every exported WAV is at most 45 seconds.
- For
labeled, sentence boundaries were obtained from AvaSanj CTC forced alignment over Negara-G2P phone sequences. Source text is concatenated only from complete aligned sentences; it is not generated or paraphrased. - AvaSanj compact-CER acceptance threshold was 0.15; Shenava long-form windows were used as an additional verification signal.
- For
to_transcribe, boundaries use detected VAD pauses only. Sentence boundaries are not verified because no source text is available. - Short residual tails are retained when merging them would exceed the maximum; therefore some chunks are shorter than the nominal 15-second cut-selection minimum.
- Audio is mono, 16 kHz, PCM WAV and is embedded in Parquet as an
Audiofeature so the Hugging Face Dataset Viewer can play it.
Provenance and limitations
The source website and article authors retain their rights. This dataset does not assert a license for the underlying recordings or texts; users are responsible for obtaining permission and complying with applicable copyright, privacy, and research-use requirements. The text labels are alignment candidates extracted from matching posts, not independently human-verified transcripts. to_transcribe requires ASR transcription before supervised speech training.
The complete construction manifest and summary are included in the repository for auditability.
