CoolFace
Datasetpublic

Srijan-Chakraborty/OCR-Finetuning-EN-Dataset

OCR-Finetuning-EN-Dataset A large-scale English OCR fine-tuning dataset containing synthetic and real-world text images for training modern OCR recognition models. The dataset is distributed in Apache Parquet format with embedded image data, making it fully compatible with the Hugging Face datasets library and the Hugging Face Dataset Viewer. Features ✅ 167,330 OCR image-text pairs ✅ Images embedded directly inside Parquet files ✅ Compatible with Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/Srijan-Chakraborty/OCR-Finetuning-EN-Dataset.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes106downloads
7 commits on main
a0165cd3mo ago

Upload folder using huggingface_hub

Srijan-Chakraborty
28b8d0b3mo ago

Update README.md

Srijan-Chakraborty
21366613mo ago

Update README.md

Srijan-Chakraborty
20945043mo ago

Update README.md

Srijan-Chakraborty
c27ee273mo ago

Delete Readme.md

Srijan-Chakraborty
3bf92b83mo ago

Upload folder using huggingface_hub

Srijan-Chakraborty
bb6057e3mo ago

initial commit

Srijan-Chakraborty