datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ganjoor-recitations
Ganjoor Persian Poetry Recitations (Full)
Every published audio recitation on Ganjoor / AVA
paired with its transcription — 30,133 clips, 1,276 hours of audio.
Audio is stored full-length and unchunked, and every clip carries a single
clean transcription in text, so it's ready for ASR / TTS training as-is.
Columns
column
description
audio
full-length mp3 (native sample rate), embedded and playable
text
full transcription of the clip… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/ganjoor-recitations.ganjoor-recitations-chunked
🗂️ ganjoor-recitations-chunked
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Ganjoor recitation chunked ASR dataset.
قطعههای تلاوت و خوانش گنجور برای آموزش و ارزیابی گفتار ادبی، شعر و خوانش رسمی فارسی.
🧩 Role
Persian speech dataset
مجموعهدادهٔ گفتار فارسی
📦 Snapshot
64 files; approximately 118.09 GB
64 فایل؛ حدود 118.09 GB
🧱 Packaging
61 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/ganjoor-recitations-chunked.librispeech_asr_dummy
Dataset Card for librispeech_asr_dummy
Dataset Summary
This is a truncated version of the LibriSpeech dataset. It contains 20 samples from each of the splits. To view the full dataset, visit: https://huggingface.co/datasets/librispeech_asr
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been… See the full description on the dataset page: https://huggingface.co/datasets/sanchit-gandhi/librispeech_asr_dummy.ganjoor-chunk-smoke
Ganjoor Recitations — Clean-Cut Chunks
Training-ready ~15-20 s segments derived from
Reza2kn/ganjoor-recitations
(full-length Persian poetry recitations from ganjoor.net). Two columns only: audio (16 kHz mono)
and text — same schema as the source, just many more rows of shorter clips + matching labels.
How it was chunked (never mid-word)
Per recitation (tools/ganjoor_chunk_job.py):
Forced-align gold text to audio with torchaudio MMS_FA (uroman -> per-word times +… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/ganjoor-chunk-smoke.
