datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quranic-asr-cloud-rawdata
Quranic ASR Provider Benchmark Results
Professional benchmark artifacts for comparing commercial and official ASR providers on the Quranic ASR benchmark hosted at Quran-Lab/quranic-asr-benchmark.
This repository contains metadata, normalized result tables, raw provider responses, unchanged run scripts, scoring outputs, Tarteel streaming probes, and reports. It does not duplicate the source audio.
What Is Included
Area
Path
Purpose
Benchmark split… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-cloud-rawdata.tantraloka-dyczkowski-raw
Two views of the same 5,146 verses
structured is the full record — 24 columns over all 37 āhnikas, including extracted
entities, cross-references, parallel passages and audio alignment. Several of those
columns hold JSON documents running to thousands of characters, which is what makes the
dataset viewer unreadable in a browser: a row is a wall of text.
reading (the default config) is a projection of the same rows onto the eight columns a
reader wants — volume, chapter, verse… See the full description on the dataset page: https://huggingface.co/datasets/Anamavajra-Labs/tantraloka-dyczkowski-raw.Podcast-Transcripts-Raw
Podcast Transcripts
Speaker-diarized transcripts from ~100k+ podcast episodes and YouTube videos (~45k hours of audio).
This dataset will be gated. Only people who are part of our team may access.
Splits
Config
Rows
Description
shows
124
Channels / podcast feeds (name, description, hosts, links)
episodes
102,374
Episode/video metadata (title, description, guests, tags, dates)
transcripts
102,374
ASR transcripts with diarized segments + YouTube… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/Podcast-Transcripts-Raw.
