tada
Datasets
All datasets matching “tada”tadabur-align-references
tadabur-align-references
Precomputed reference embeddings powering tadabur-align — word-level timestamp extraction for Quranic recitation via DTW alignment transfer (no ASR).
What this is
For 5,481 of the Quran's 6,236 ayahs, this dataset holds frame-level tadabur-embedding features for up to 8 reference reciters, plus each reference's word-level timestamps and internal-pause intervals. No audio is included — only model outputs and timing data. tadabur-align… See the full description on the dataset page: https://huggingface.co/datasets/FaisaI/tadabur-align-references.tadabur
Tadabur: A Large-Scale Quran Audio Dataset
The most comprehensive and richly annotated Qur'anic recitation corpus to date
Faisal Alherran
✦ Overview
Tadabur is a large-scale, high-diversity Qur'anic speech dataset designed to advance research in Qur'anic Automatic Speech Recognition (ASR), reciter modeling, tajwīd-aware speech processing, and prosodic analysis. It is the most comprehensive publicly available collection of… See the full description on the dataset page: https://huggingface.co/datasets/FaisaI/tadabur.radiotalk-us-audio-tada-clean
RadioTalk US Audio (Clean)
Synthesized clean-speech audio for ~100k US air-traffic-control conversation scenarios. One row per turn, embedded 24 kHz mono PCM_16 WAV.
This is the clean variant. A VHF-AM-channel-degraded variant is published as twangodev/radiotalk-us-audio-tada-noisy.
Quick start
from datasets import load_dataset
ds = load_dataset("twangodev/radiotalk-us-audio-tada-clean", split="train", streaming=True)
row = next(iter(ds))
print(row["text"]… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-audio-tada-clean.tadabur-lora-data-fullmidf-egangotri-sanskrit
MIDF/eGangotri Sanskrit Manuscripts
Reviewed line-segmentation annotations
Segmentation v1.1 contains 2,879 reviewed
pages with images, curved PAGE XML baselines, and editable geometry. Its 1,916
training pages contain 18,990 lines. A 60-page panel supports checkpoint
selection, while 734 pages from three unseen manuscripts support broader
validation. The test data contains 220 pages from the unseen M00638 manuscript
and nine fixed adaptation pages from the… See the full description on the dataset page: https://huggingface.co/datasets/tadad/midf-egangotri-sanskrit.kat57-ocr-bench-500
Kat57 OCR outputs
Raw outputs from 16 OCR models on the same deterministic 500-card sample of Lund University Library's Kat57 catalogue-card collection.
Each model is stored as a separate dataset configuration. Every configuration retains the source card identifiers, image, PAGE XML reference transcription, model output, and inference metadata so the results can be rescored without rerunning inference.
Source sample
CER/WER results and limitations
ocr-bench… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ocr-bench-500.
