CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lucazhou2000 /sciencemysterybench-transcriptsimagen<1K0 likes2k downloads7d agoHugging Face02ashraq /youtube-transcriptionThis is YouTube video transcription dataset built from YTTTS Speech Collection for semantic search.image1K<n<10K5 likes59 downloads4y agoHugging Face03QasimHussain /spatial-transcriptomics-atlas-demo Spatial Transcriptomics Atlas: Human Lymph Node Architecture and Immune Microenvironment Integrated multi-modal analysis of spatial gene expression in human lymph node tissue. Panels depict high-resolution histology (A), annotated tissue domains (B), gene detection density (C), expression patterns of top spatially variable genes (D--F), and neighborhood enrichment statistics (G). Abstract This repository presents a reproducible computational… See the full description on the dataset page: https://huggingface.co/datasets/QasimHussain/spatial-transcriptomics-atlas-demo.imagen<1K0 likes57 downloads2mo agoHugging Face04BDRC /monlamai-transcriptions Tibetan OCR — MonlamAI transcriptions 3,072 page images of Tibetan dbu-med (u-med) manuscripts with page-level Unicode transcriptions, contributed by MonlamAI over BDRC manuscript scans and aligned page by page. This is the full page equivalent of the line-segmented datasets available on openpecha/OCR-Betsug and openpecha/OCR-Drutsa, with the following changes: filter out cases where not all line in a page are transcribed (so not all the lines in the original datasets are… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/monlamai-transcriptions.imageimage-to-text1K<n<10K0 likes57 downloads1mo agoHugging Face05BDRC /palri-parkhang-transcriptions Tibetan OCR — Palri Parkhang 11,133 page images of Tibetan text with page-level Unicode transcriptions, mostly dbu-med (u-med) manuscripts with a small uchen portion. The transcriptions were produced by Palri Parkhang, an input project led by Chris Tomlinson (former BDRC's CTO) in Nepal in 2006-2013 and aligned to BDRC scans; the material is largely Nyingma collected works and gter ma cycles. The transcriptions prioritized legibility over fidelity to the enscribed text and… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/palri-parkhang-transcriptions.imageimage-to-text10K<n<100K0 likes54 downloads1mo agoHugging Face06BDRC /berkeley-transcriptions Tibetan OCR — Berkeley 8,866 page images of Tibetan text with page-level Unicode transcriptions, mostly dbu-med (u-med) manuscripts with a small woodblock (uchen) portion. These works were transcribed by Geshe Dangsong Namgyal, from 2016 to 2026, in his role as a Data System Analyst at University of California, Berkeley. This work was initiated by Prof. Kurt Keutzer with the longstanding hope of producing training data for Handwritten Text Recognition for dbu med and OCR of… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/berkeley-transcriptions.imageimage-to-text1K<n<10K0 likes52 downloads1mo agoHugging Face07pinecone /yt-transcriptionsimage10K<n<100K1 likes43 downloads4y agoHugging Face08BSC-CSSH /AMSMB-line-transcription Dataset Card Dataset for line-level handwritten text recognition on medieval historical manuscripts, consisting of 3,369 lines (images of text lines with the associated transcription and metadata) from 100 digitized documents written by at least 80 different hands and spanning three centuries (from 1208 to 1499). This dataset is derived from the AMSMB dataset, which contains the full-page images of the digitized manuscripts and their associated transcriptions in the PageXML format.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-CSSH/AMSMB-line-transcription.imageimage-to-text1K<n<10K0 likes32 downloads1y agoHugging Face09justinsunqiu /multilingual_transcriptions_summarized_by_english_backtranslated_finalimage10K<n<100K0 likes21 downloads1y agoHugging Face10jdabello /yt_transcriptionsimage10K<n<100K0 likes20 downloads3y agoHugging Face11QasimHussain /transcriptome-health-dashboard-demo Transcriptome Health Dashboard v2.0 A professional-grade RNA-Seq quality control and analysis pipeline implementing biologically-rigorous normalization, interactive visualizations, and comprehensive sample QC metrics. Principal Component Analysis of 424 TCGA-LIHC samples visualizing transcriptomic structure. Overview This pipeline performs comprehensive quality control analysis for bulk RNA-Seq datasets, implementing industry-standard bioinformatics… See the full description on the dataset page: https://huggingface.co/datasets/QasimHussain/transcriptome-health-dashboard-demo.imagen<1K0 likes14 downloads2mo agoHugging Face12justinsunqiu /multilingual_transcriptions_finalimage1K<n<10K0 likes13 downloads1y agoHugging Face13justinsunqiu /multilingual_transcriptions_fullimage1K<n<10K0 likes11 downloads1y agoHugging Face14justinsunqiu /multilingual_transcriptions_translated_rawimage1K<n<10K0 likes9 downloads1y agoHugging Face15justinsunqiu /multilingual_transcriptions_cleanedimage1K<n<10K0 likes8 downloads1y agoHugging Face16justinsunqiu /multilingual_transcriptions_translated_english_finalimage1K<n<10K0 likes8 downloads1y agoHugging Face17justinsunqiu /multilingual_transcriptions_summarizedimage1K<n<10K0 likes7 downloads1y agoHugging Face18justinsunqiu /transcription_changesimagen<1K0 likes7 downloads1y agoHugging Face19justinsunqiu /multilingual_transcriptions_summarized_by_native_nonnativeimage1K<n<10K0 likes6 downloads1y agoHugging Face20justinsunqiu /multilingual_transcriptions_summarized_by_type_finalimage1K<n<10K0 likes6 downloads1y agoHugging Face21justinsunqiu /multilingual_transcriptionsimage1K<n<10K0 likes5 downloads2y agoHugging Face22justinsunqiu /transcription_changes_classifiedimagen<1K0 likes5 downloads1y agoHugging Face23justinsunqiu /multilingual_transcriptions_rawimage1K<n<10K0 likes3 downloads2y agoHugging Face24curiousmrk /transcription-coding-wiki-500kgated Transcription Dataset: Code & Wiki (390K) Text-to-image rendered dataset for training vision-language models to read code and text from images. Schema Column Type Description image Image Rendered grayscale JPEG prompt string Transcription instruction (varied) response string Ground truth text language string python/javascript/java/c++/rust/go/english domain string code or english length_bucket string short/medium/long/gundam resolution string… See the full description on the dataset page: https://huggingface.co/datasets/curiousmrk/transcription-coding-wiki-500k.image100K<n<1M0 likes2 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.