CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FaisaI /tadabur Tadabur: A Large-Scale Quran Audio Dataset The most comprehensive and richly annotated Qur'anic recitation corpus to date Faisal Alherran &nbsp; &nbsp; &nbsp; ✦ Overview Tadabur is a large-scale, high-diversity Qur'anic speech dataset designed to advance research in Qur'anic Automatic Speech Recognition (ASR), reciter modeling, tajwīd-aware speech processing, and prosodic analysis. It is the most comprehensive publicly available collection of… See the full description on the dataset page: https://huggingface.co/datasets/FaisaI/tadabur.audioaudio-classification100K<n<1M23 likes3.3k downloads2mo agoHugging Face02twangodev /radiotalk-us-audio-tada-clean RadioTalk US Audio (Clean) Synthesized clean-speech audio for ~100k US air-traffic-control conversation scenarios. One row per turn, embedded 24 kHz mono PCM_16 WAV. This is the clean variant. A VHF-AM-channel-degraded variant is published as twangodev/radiotalk-us-audio-tada-noisy. Quick start from datasets import load_dataset ds = load_dataset("twangodev/radiotalk-us-audio-tada-clean", split="train", streaming=True) row = next(iter(ds)) print(row["text"]… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-audio-tada-clean.audiotext-to-speech1M<n<10M0 likes1.2k downloads2mo agoHugging Face03AhmedSaman /tadabur-lora-data-fullaudio10K<n<100K0 likes1.1k downloads26d agoHugging Face04tadad /midf-egangotri-sanskrit MIDF/eGangotri Sanskrit Manuscripts Reviewed line-segmentation annotations Segmentation v1.1 contains 2,879 reviewed pages with images, curved PAGE XML baselines, and editable geometry. Its 1,916 training pages contain 18,990 lines. A 60-page panel supports checkpoint selection, while 734 pages from three unseen manuscripts support broader validation. The test data contains 220 pages from the unseen M00638 manuscript and nine fixed adaptation pages from the… See the full description on the dataset page: https://huggingface.co/datasets/tadad/midf-egangotri-sanskrit.documentimage-to-text100K<n<1M0 likes1k downloads7d agoHugging Face05twangodev /radiotalk-us-audio-tada-noisy RadioTalk US Audio (Noisy) VHF AM aviation channel-degraded variants of synthesized US air-traffic-control speech. One row per (clean turn × variant), embedded 8 kHz mono PCM_16 WAV. This is the noisy variant of twangodev/radiotalk-us-audio-tada-clean — same transcripts and voices, passed through a probabilistic channel-simulation pipeline calibrated to the ATCO2 corpus SNR distribution (mean ~8 dB, range -5 to +30 dB) and shaped to ITU-R M.1084 / DO-186B aero voice passband… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-audio-tada-noisy.audioautomatic-speech-recognition1M<n<10M0 likes594 downloads2mo agoHugging Face06tadad /kat57-ocr-bench-500 Kat57 OCR outputs Raw outputs from 16 OCR models on the same deterministic 500-card sample of Lund University Library's Kat57 catalogue-card collection. Each model is stored as a separate dataset configuration. Every configuration retains the source card identifiers, image, PAGE XML reference transcription, model output, and inference metadata so the results can be rescored without rerunning inference. Source sample CER/WER results and limitations ocr-bench… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ocr-bench-500.image1K<n<10K0 likes499 downloads19d agoHugging Face07tadad /kat57-ground-truth Kat57 ground truth Hugging Face conversion of Lund University Library's Kat57 ground-truth release: 10,695 scanned catalogue cards with manually corrected PAGE XML transcriptions. The cards come from Catalogue -1957, Lund University Library's alphabetical catalogue of holdings published through 1957. They contain a mixture of typewritten and handwritten text in several languages. Fields image: original PNG card scan reference: line transcriptions joined in PAGE… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ground-truth.image10K<n<100K1 likes445 downloads19d agoHugging Face08MShakir7137 /tadabur Tadabur: A Large-Scale Quran Audio Dataset The most comprehensive and richly annotated Qur'anic recitation corpus to date Faisal Alherran &nbsp; &nbsp; &nbsp; ✦ Overview Tadabur is a large-scale, high-diversity Qur'anic speech dataset designed to advance research in Qur'anic Automatic Speech Recognition (ASR), reciter modeling, tajwīd-aware speech processing, and prosodic analysis. It is the most comprehensive publicly available collection of… See the full description on the dataset page: https://huggingface.co/datasets/MShakir7137/tadabur.audioaudio-classification100K<n<1M0 likes261 downloads2mo agoHugging Face09tadao2024 /korabotextn<1K0 likes259 downloads1d agoHugging Face10JinGao /TadA-BenchTadA-Bench A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein Engineering Jin Gao1, Juntu Zhao1, Zirui Zeng1, Jiaqi Shen1, Junhao Shi1, Dukun Zhao1, Yuming Lu1,†, Dequan Wang1,2,† 1Shanghai Jiao Tong University &nbsp; · &nbsp; 2Shanghai Innovation Institute †Corresponding authors Dataset Summary TadA-Bench is a fixed-data wet-lab replay benchmark built from 31 rounds of TadA directed evolution.… See the full description on the dataset page: https://huggingface.co/datasets/JinGao/TadA-Bench.text1M<n<10M0 likes258 downloads4mo agoHugging Face11BangumiBase /tadaimaokaeri Bangumi Image Base of Tadaima, Okaeri This is the image base of bangumi Tadaima, Okaeri, we detected 21 characters, 3782 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/tadaimaokaeri.image1K<n<10K0 likes255 downloads2y agoHugging Face12tadad /french-fiction-16-18th-century French Fiction of the 16th–18th Centuries A Hugging Face conversion of Pierre-Carl Langlais's French Fiction of the 16–18th century deposit for the BigLAM community. It contains historical French OCR, bibliographic metadata, a genre-labeled and lemmatized subset, and the source R model. The Zenodo deposit is the source of record. This conversion preserves its OCR, metadata, work assignments, and labels without scholarly correction. Structure Configuration… See the full description on the dataset page: https://huggingface.co/datasets/tadad/french-fiction-16-18th-century.tabulartext-classification100K<n<1M0 likes163 downloads18d agoHugging Face13tadad /kat57-ocr-bench-500-results Kat57 OCR benchmark — CER/WER Strict reference-based evaluation of 16 OCR models on a deterministic 500-card sample from Lund University Library's Kat57 catalogue-card collection. This result set contains only Character Error Rate (CER) and Word Error Rate (WER); it does not contain VLM judging or ELO ratings. The sample was drawn with seed 57 from tadad/kat57-ground-truth and is published as tadad/kat57-ground-truth-500. The OCR outputs are retained in… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ocr-bench-500-results.tabular1K<n<10K0 likes162 downloads19d agoHugging Face14tadad /twkm-maya-glyphs TWKM Maya Glyphs An ML-ready, relational snapshot of the University of Bonn's Text Database and Dictionary of Classic Mayan (TWKM) digital sign catalogue. It packages the catalogue as seven Parquet configurations and embeds the explicitly CC BY 4.0 standardized graph drawings in the graphs configuration. This is an independent preservation and interoperability package, not an official TWKM publication. The source of authority for sign classification and interpretation remains… See the full description on the dataset page: https://huggingface.co/datasets/tadad/twkm-maya-glyphs.imageimage-classification10K<n<100K0 likes133 downloads18d agoHugging Face15tadad /kat57-ground-truth-500 Kat57 500-card benchmark sample A deterministic 500-card sample of tadad/kat57-ground-truth for comparing OCR systems against Kat57's human-corrected transcriptions. The sample is drawn from all 10,695 source rows by ranking each stable id with SHA-256 over seed + NUL + id, selecting the lowest 500 digests, and restoring source order. Sampling seed: 57. Pinned source revision: 2f4b7e6a8f8746631c0628280dd3f40be2b997f6. Every selected row has a non-empty reference; all original… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ground-truth-500.imagen<1K0 likes79 downloads19d agoHugging Face16tadad /diorisis-ancient-greek Diorisis Ancient Greek Corpus A Hugging Face conversion of Alessandro Vatri and Barbara McGillivray's Diorisis Ancient Greek Corpus for the BigLAM community. Diorisis contains 820 literary texts from Homer through the fifth century CE, with automatic lemma, part-of-speech, and morphological annotations. The conversion combines the original XML headers with the JSON corpus and a checksum-pinned snapshot of the author's public per-file corrections. It retains Beta Code and adds… See the full description on the dataset page: https://huggingface.co/datasets/tadad/diorisis-ancient-greek.tabulartoken-classification100K<n<1M0 likes73 downloads18d agoHugging Face17tadad /kat57-ocr-bench-results Kat57 OCR smoke benchmark results Exact ground-truth scoring for a 50-card Tesseract integration run over tadad/kat57-ground-truth-smoke. OCR outputs are published in the tesseract config of tadad/kat57-ocr-bench. Model CER WER Evaluated Empty outputs Error sentinels Skipped references Tesseract 5 0.4656 0.8605 50 1 0 0 The corpus totals are 4,629 character edits over 9,941 reference characters and 1,221 word edits over 1,419 reference words. Scoring used… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ocr-bench-results.tabularn<1K0 likes67 downloads19d agoHugging Face18tadad /kat57-ocr-bench Document OCR using Tesseract This dataset contains OCR results from images in tadad/kat57-ground-truth-smoke using Tesseract, the classical open-source CPU OCR engine — a cheap, no-GPU baseline alongside the VLM OCR recipes. Processing Details Source Dataset: tadad/kat57-ground-truth-smoke Engine: Tesseract 5.3.0 Language(s): eng Number of Samples: 50 Processing Time: 1.1 min Processing Date: 2026-09-03 18:39 UTC Configuration Image Column: image… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ocr-bench.imagen<1K0 likes59 downloads19d agoHugging Face19tadad /kat57-ground-truth-smoke Kat57 ground-truth smoke subset A 50-card integration subset of tadad/kat57-ground-truth, drawn from the first published Parquet shard. This subset exists to test OCR pipelines and exact CER/WER scoring without downloading the full 36 GB collection. It is ordered by source identifier and is not a representative benchmark sample; substantive Kat57 claims should use a documented sample across the full collection. All fields are preserved from the full conversion, including the… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ground-truth-smoke.imagen<1K0 likes47 downloads19d agoHugging Face20AhmedSaman /tadabur-lora-dataaudio10K<n<100K0 likes39 downloads29d agoHugging Face21tadakaluri /TeluguRiddles Summary TeluguRiddles is an open source dataset of instruct-style records generated by webscraping multiple riddles websites. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview TeluguRiddles is a corpus of… See the full description on the dataset page: https://huggingface.co/datasets/tadakaluri/TeluguRiddles.texttext-generationn<1K0 likes12 downloads10mo agoHugging Face22LeonGuertler /TA-Dataset-ColonelBlottotextn<1K0 likes5 downloads1y agoHugging Face23tadashi-asaoka /gsm8k_testtextn<1K0 likes4 downloads2y agoHugging Face24ngtranai09 /TAdam8bitImbalancetabularn<1K0 likes4 downloads11mo agoHugging Face25ta-datalab /research_paperstabular10K<n<100K0 likes2 downloads2y agoHugging Face26ta-datalab /research_papers_for_knowledge_transfertextn<1K0 likes2 downloads1y agoHugging Face27a920177242 /ta-datasettextn<1K0 likes2 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.