florapeterpaul3/geeklink-ocr-benchmark
GeekLink OCR Benchmark Published by GeekLink, a Mac/Windows app for extracting, translating, and burning in video subtitles. This dataset is an open benchmark we use to evaluate our own OCR model against other engines — not the app itself. If you're looking for the GeekLink app, see geeklink.dev. Source code and eval scripts: github.com/GeekLinkDev/geeklink-ocr-benchmark. A benchmark for burned-in video subtitle OCR: 1,140 subtitle images across 6 languages (English, Spanish… See the full description on the dataset page: https://huggingface.co/datasets/florapeterpaul3/geeklink-ocr-benchmark.
docs: clarify GeekLink is the app, this dataset is just a benchmark
Fix clipped subtitle text (437 samples), extend to 1140 samples
Initial upload: 600-sample burned-in subtitle OCR benchmark
initial commit
