CoolFace
Datasetpublic

florapeterpaul3/geeklink-ocr-benchmark

GeekLink OCR Benchmark Published by GeekLink, a Mac/Windows app for extracting, translating, and burning in video subtitles. This dataset is an open benchmark we use to evaluate our own OCR model against other engines — not the app itself. If you're looking for the GeekLink app, see geeklink.dev. Source code and eval scripts: github.com/GeekLinkDev/geeklink-ocr-benchmark. A benchmark for burned-in video subtitle OCR: 1,140 subtitle images across 6 languages (English, Spanish… See the full description on the dataset page: https://huggingface.co/datasets/florapeterpaul3/geeklink-ocr-benchmark.

sourceHugging Facecc0-1.0updated 27d agoView on Hugging Face
0likes66downloads
4 commits on main
c989a6d27d ago

docs: clarify GeekLink is the app, this dataset is just a benchmark

Flora
5d7e1911mo ago

Fix clipped subtitle text (437 samples), extend to 1140 samples

florapeterpaul3
85ede8a1mo ago

Initial upload: 600-sample burned-in subtitle OCR benchmark

florapeterpaul3
63397131mo ago

initial commit

florapeterpaul3