datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tadabur
Tadabur: A Large-Scale Quran Audio Dataset
The most comprehensive and richly annotated Qur'anic recitation corpus to date
Faisal Alherran
✦ Overview
Tadabur is a large-scale, high-diversity Qur'anic speech dataset designed to advance research in Qur'anic Automatic Speech Recognition (ASR), reciter modeling, tajwīd-aware speech processing, and prosodic analysis. It is the most comprehensive publicly available collection of… See the full description on the dataset page: https://huggingface.co/datasets/FaisaI/tadabur.radiotalk-us-audio-tada-clean
RadioTalk US Audio (Clean)
Synthesized clean-speech audio for ~100k US air-traffic-control conversation scenarios. One row per turn, embedded 24 kHz mono PCM_16 WAV.
This is the clean variant. A VHF-AM-channel-degraded variant is published as twangodev/radiotalk-us-audio-tada-noisy.
Quick start
from datasets import load_dataset
ds = load_dataset("twangodev/radiotalk-us-audio-tada-clean", split="train", streaming=True)
row = next(iter(ds))
print(row["text"]… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-audio-tada-clean.tadabur-lora-data-fullmidf-egangotri-sanskrit
MIDF/eGangotri Sanskrit Manuscripts
Reviewed line-segmentation annotations
Segmentation v1.1 contains 2,879 reviewed
pages with images, curved PAGE XML baselines, and editable geometry. Its 1,916
training pages contain 18,990 lines. A 60-page panel supports checkpoint
selection, while 734 pages from three unseen manuscripts support broader
validation. The test data contains 220 pages from the unseen M00638 manuscript
and nine fixed adaptation pages from the… See the full description on the dataset page: https://huggingface.co/datasets/tadad/midf-egangotri-sanskrit.radiotalk-us-audio-tada-noisy
RadioTalk US Audio (Noisy)
VHF AM aviation channel-degraded variants of synthesized US air-traffic-control speech. One row per (clean turn × variant), embedded 8 kHz mono PCM_16 WAV.
This is the noisy variant of twangodev/radiotalk-us-audio-tada-clean — same transcripts and voices, passed through a probabilistic channel-simulation pipeline calibrated to the ATCO2 corpus SNR distribution (mean ~8 dB, range -5 to +30 dB) and shaped to ITU-R M.1084 / DO-186B aero voice passband… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-audio-tada-noisy.kat57-ocr-bench-500
Kat57 OCR outputs
Raw outputs from 16 OCR models on the same deterministic 500-card sample of Lund University Library's Kat57 catalogue-card collection.
Each model is stored as a separate dataset configuration. Every configuration retains the source card identifiers, image, PAGE XML reference transcription, model output, and inference metadata so the results can be rescored without rerunning inference.
Source sample
CER/WER results and limitations
ocr-bench… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ocr-bench-500.kat57-ground-truth
Kat57 ground truth
Hugging Face conversion of Lund University Library's
Kat57 ground-truth release: 10,695 scanned catalogue
cards with manually corrected PAGE XML transcriptions.
The cards come from Catalogue -1957, Lund University Library's alphabetical
catalogue of holdings published through 1957. They contain a mixture of
typewritten and handwritten text in several languages.
Fields
image: original PNG card scan
reference: line transcriptions joined in PAGE… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ground-truth.tadabur
Tadabur: A Large-Scale Quran Audio Dataset
The most comprehensive and richly annotated Qur'anic recitation corpus to date
Faisal Alherran
✦ Overview
Tadabur is a large-scale, high-diversity Qur'anic speech dataset designed to advance research in Qur'anic Automatic Speech Recognition (ASR), reciter modeling, tajwīd-aware speech processing, and prosodic analysis. It is the most comprehensive publicly available collection of… See the full description on the dataset page: https://huggingface.co/datasets/MShakir7137/tadabur.koraboTadA-BenchTadA-Bench
A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein Engineering
Jin Gao1,
Juntu Zhao1,
Zirui Zeng1,
Jiaqi Shen1,
Junhao Shi1,
Dukun Zhao1,
Yuming Lu1,†,
Dequan Wang1,2,†
1Shanghai Jiao Tong University ·
2Shanghai Innovation Institute
†Corresponding authors
Dataset Summary
TadA-Bench is a fixed-data wet-lab replay benchmark built from 31 rounds of
TadA directed evolution.… See the full description on the dataset page: https://huggingface.co/datasets/JinGao/TadA-Bench.tadaimaokaeri
Bangumi Image Base of Tadaima, Okaeri
This is the image base of bangumi Tadaima, Okaeri, we detected 21 characters, 3782 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/tadaimaokaeri.french-fiction-16-18th-century
French Fiction of the 16th–18th Centuries
A Hugging Face conversion of Pierre-Carl Langlais's French Fiction of the 16–18th century deposit for the BigLAM community. It contains historical French OCR, bibliographic metadata, a genre-labeled and lemmatized subset, and the source R model.
The Zenodo deposit is the source of record. This conversion preserves its OCR, metadata, work assignments, and labels without scholarly correction.
Structure
Configuration… See the full description on the dataset page: https://huggingface.co/datasets/tadad/french-fiction-16-18th-century.kat57-ocr-bench-500-results
Kat57 OCR benchmark — CER/WER
Strict reference-based evaluation of 16 OCR models on a deterministic 500-card sample from Lund University Library's Kat57 catalogue-card collection. This result set contains only Character Error Rate (CER) and Word Error Rate (WER); it does not contain VLM judging or ELO ratings.
The sample was drawn with seed 57 from tadad/kat57-ground-truth and is published as tadad/kat57-ground-truth-500. The OCR outputs are retained in… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ocr-bench-500-results.twkm-maya-glyphs
TWKM Maya Glyphs
An ML-ready, relational snapshot of the University of Bonn's Text Database and Dictionary of Classic Mayan (TWKM) digital sign catalogue. It packages the catalogue as seven Parquet configurations and embeds the explicitly CC BY 4.0 standardized graph drawings in the graphs configuration.
This is an independent preservation and interoperability package, not an official TWKM publication. The source of authority for sign classification and interpretation remains… See the full description on the dataset page: https://huggingface.co/datasets/tadad/twkm-maya-glyphs.kat57-ground-truth-500
Kat57 500-card benchmark sample
A deterministic 500-card sample of
tadad/kat57-ground-truth for comparing OCR systems
against Kat57's human-corrected transcriptions.
The sample is drawn from all 10,695 source rows by ranking each stable id with
SHA-256 over seed + NUL + id, selecting the lowest 500 digests, and restoring
source order. Sampling seed: 57. Pinned source revision: 2f4b7e6a8f8746631c0628280dd3f40be2b997f6.
Every selected row has a non-empty reference; all original… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ground-truth-500.diorisis-ancient-greek
Diorisis Ancient Greek Corpus
A Hugging Face conversion of Alessandro Vatri and Barbara McGillivray's
Diorisis Ancient Greek Corpus for the
BigLAM community. Diorisis contains 820 literary texts from Homer through the fifth century CE,
with automatic lemma, part-of-speech, and morphological annotations.
The conversion combines the original XML headers with the
JSON corpus and a checksum-pinned snapshot of
the author's public per-file corrections. It retains Beta Code and adds… See the full description on the dataset page: https://huggingface.co/datasets/tadad/diorisis-ancient-greek.kat57-ocr-bench-results
Kat57 OCR smoke benchmark results
Exact ground-truth scoring for a 50-card Tesseract integration run over tadad/kat57-ground-truth-smoke. OCR outputs are published in the tesseract config of tadad/kat57-ocr-bench.
Model
CER
WER
Evaluated
Empty outputs
Error sentinels
Skipped references
Tesseract 5
0.4656
0.8605
50
1
0
0
The corpus totals are 4,629 character edits over 9,941 reference characters and 1,221 word edits over 1,419 reference words. Scoring used… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ocr-bench-results.kat57-ocr-bench
Document OCR using Tesseract
This dataset contains OCR results from images in tadad/kat57-ground-truth-smoke using Tesseract, the classical open-source CPU OCR engine — a cheap, no-GPU baseline alongside the VLM OCR recipes.
Processing Details
Source Dataset: tadad/kat57-ground-truth-smoke
Engine: Tesseract 5.3.0
Language(s): eng
Number of Samples: 50
Processing Time: 1.1 min
Processing Date: 2026-09-03 18:39 UTC
Configuration
Image Column: image… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ocr-bench.kat57-ground-truth-smoke
Kat57 ground-truth smoke subset
A 50-card integration subset of tadad/kat57-ground-truth, drawn from the first published Parquet shard.
This subset exists to test OCR pipelines and exact CER/WER scoring without downloading the full 36 GB collection. It is ordered by source identifier and is not a representative benchmark sample; substantive Kat57 claims should use a documented sample across the full collection.
All fields are preserved from the full conversion, including the… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ground-truth-smoke.tadabur-lora-dataTeluguRiddles
Summary
TeluguRiddles is an open source dataset of instruct-style records generated by webscraping multiple riddles websites. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview
TeluguRiddles is a corpus of… See the full description on the dataset page: https://huggingface.co/datasets/tadakaluri/TeluguRiddles.TA-Dataset-ColonelBlottogsm8k_testTAdam8bitImbalanceresearch_papersresearch_papers_for_knowledge_transferta-dataset
