datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ganjoor-chunk-smoke
Ganjoor Recitations — Clean-Cut Chunks
Training-ready ~15-20 s segments derived from
Reza2kn/ganjoor-recitations
(full-length Persian poetry recitations from ganjoor.net). Two columns only: audio (16 kHz mono)
and text — same schema as the source, just many more rows of shorter clips + matching labels.
How it was chunked (never mid-word)
Per recitation (tools/ganjoor_chunk_job.py):
Forced-align gold text to audio with torchaudio MMS_FA (uroman -> per-word times +… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/ganjoor-chunk-smoke.stage1a_smoke_data
stage1a_smoke_data — AuT-ready 128-mel TFRecords (en/zh)
Smoke-scale training data for Stage 1A input audio alignment of a Qwen3-ASR-AuT → MLP →
frozen-VL-LLM omni model. Audio is pre-extracted 128-bin log-mel (the Qwen3-ASR AuT frontend:
WhisperFeatureExtractor, 16 kHz, hop 160, n_fft 400) so training only needs to run the frozen AuT
encoder — no raw-audio decoding at train time.
113,396 samples across 4 sources, stored as GZIP-compressed TFRecords (one file per source shard).… See the full description on the dataset page: https://huggingface.co/datasets/Letian2003/stage1a_smoke_data.wordts_smoke
wordts_smoke
This dataset contains lossless FLAC chunks derived from 30 Nigerian-language,
Nigerian English, and Nigerian Pidgin recordings. Access requests require
manual approval by the repository owner.
Contents
Chunks: 2,217
Chunk audio duration: 3.673 hours
Source transcript rows represented: 4,328
Standalone audio-tag chunks: 0
Non-music tag rows attached to nearest same-speaker speech:
102
Music annotation boundaries: 227
Non-speech annotations: 553 (also… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/wordts_smoke.
