curiousmrk/transcription-coding-wiki-500k
Transcription Dataset: Code & Wiki (390K) Text-to-image rendered dataset for training vision-language models to read code and text from images. Schema Column Type Description image Image Rendered grayscale JPEG prompt string Transcription instruction (varied) response string Ground truth text language string python/javascript/java/c++/rust/go/english domain string code or english length_bucket string short/medium/long/gundam resolution… See the full description on the dataset page: https://huggingface.co/datasets/curiousmrk/transcription-coding-wiki-500k.
02
No card is published for this repository, or it could not be fetched from Hugging Face right now.
