datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sciencemysterybench-transcriptsyoutube-transcriptionThis is YouTube video transcription dataset built from YTTTS Speech Collection for semantic search.spatial-transcriptomics-atlas-demo
Spatial Transcriptomics Atlas: Human Lymph Node Architecture and Immune Microenvironment
Integrated multi-modal analysis of spatial gene expression in human lymph node tissue. Panels depict high-resolution histology (A), annotated tissue domains (B), gene detection density (C), expression patterns of top spatially variable genes (D--F), and neighborhood enrichment statistics (G).
Abstract
This repository presents a reproducible computational… See the full description on the dataset page: https://huggingface.co/datasets/QasimHussain/spatial-transcriptomics-atlas-demo.monlamai-transcriptions
Tibetan OCR — MonlamAI transcriptions
3,072 page images of Tibetan dbu-med (u-med) manuscripts with page-level Unicode
transcriptions, contributed by MonlamAI over BDRC manuscript
scans and aligned page by page.
This is the full page equivalent of the line-segmented datasets available on openpecha/OCR-Betsug and openpecha/OCR-Drutsa, with the following changes:
filter out cases where not all line in a page are transcribed (so not all the lines in the original datasets are… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/monlamai-transcriptions.palri-parkhang-transcriptions
Tibetan OCR — Palri Parkhang
11,133 page images of Tibetan text with page-level Unicode transcriptions, mostly
dbu-med (u-med) manuscripts with a small uchen portion.
The transcriptions were
produced by Palri Parkhang, an input project led by Chris Tomlinson (former BDRC's CTO) in Nepal in 2006-2013 and aligned to BDRC scans; the material is largely
Nyingma collected works and gter ma cycles. The transcriptions prioritized legibility over fidelity to the enscribed text and… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/palri-parkhang-transcriptions.berkeley-transcriptions
Tibetan OCR — Berkeley
8,866 page images of Tibetan text with page-level Unicode transcriptions, mostly
dbu-med (u-med) manuscripts with a small woodblock (uchen) portion.
These works were transcribed by Geshe Dangsong Namgyal, from 2016 to 2026, in his role as a Data System Analyst at University of California, Berkeley.
This work was initiated by Prof. Kurt Keutzer with the longstanding hope of producing training data for Handwritten Text Recognition for dbu med and OCR of… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/berkeley-transcriptions.yt-transcriptionsAMSMB-line-transcription
Dataset Card
Dataset for line-level handwritten text recognition on medieval historical manuscripts, consisting of 3,369 lines (images of text lines with the associated transcription and metadata) from 100 digitized documents written by at least 80 different hands and spanning three centuries (from 1208 to 1499). This dataset is derived from the AMSMB dataset, which contains the full-page images of the digitized manuscripts and their associated transcriptions in the PageXML format.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-CSSH/AMSMB-line-transcription.multilingual_transcriptions_summarized_by_english_backtranslated_finalyt_transcriptionstranscriptome-health-dashboard-demo
Transcriptome Health Dashboard v2.0
A professional-grade RNA-Seq quality control and analysis pipeline implementing biologically-rigorous normalization, interactive visualizations, and comprehensive sample QC metrics.
Principal Component Analysis of 424 TCGA-LIHC samples visualizing transcriptomic structure.
Overview
This pipeline performs comprehensive quality control analysis for bulk RNA-Seq datasets, implementing industry-standard bioinformatics… See the full description on the dataset page: https://huggingface.co/datasets/QasimHussain/transcriptome-health-dashboard-demo.multilingual_transcriptions_finalmultilingual_transcriptions_fullmultilingual_transcriptions_translated_rawmultilingual_transcriptions_cleanedmultilingual_transcriptions_translated_english_finalmultilingual_transcriptions_summarizedtranscription_changesmultilingual_transcriptions_summarized_by_native_nonnativemultilingual_transcriptions_summarized_by_type_finalmultilingual_transcriptionstranscription_changes_classifiedmultilingual_transcriptions_rawtranscription-coding-wiki-500k
Transcription Dataset: Code & Wiki (390K)
Text-to-image rendered dataset for training vision-language models to read code and text from images.
Schema
Column
Type
Description
image
Image
Rendered grayscale JPEG
prompt
string
Transcription instruction (varied)
response
string
Ground truth text
language
string
python/javascript/java/c++/rust/go/english
domain
string
code or english
length_bucket
string
short/medium/long/gundam
resolution
string… See the full description on the dataset page: https://huggingface.co/datasets/curiousmrk/transcription-coding-wiki-500k.
