CoolFace
Datasetpublicgated

curiousmrk/transcription-coding-wiki-500k

Transcription Dataset: Code & Wiki (390K) Text-to-image rendered dataset for training vision-language models to read code and text from images. Schema Column Type Description image Image Rendered grayscale JPEG prompt string Transcription instruction (varied) response string Ground truth text language string python/javascript/java/c++/rust/go/english domain string code or english length_bucket string short/medium/long/gundam resolution… See the full description on the dataset page: https://huggingface.co/datasets/curiousmrk/transcription-coding-wiki-500k.

sourceHugging Faceupdated 9mo agoView on Hugging Face
0likes2downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.