Srijan-Chakraborty/OCR-Finetuning-EN-Dataset
OCR-Finetuning-EN-Dataset A large-scale English OCR fine-tuning dataset containing synthetic and real-world text images for training modern OCR recognition models. The dataset is distributed in Apache Parquet format with embedded image data, making it fully compatible with the Hugging Face datasets library and the Hugging Face Dataset Viewer. Features ✅ 167,330 OCR image-text pairs ✅ Images embedded directly inside Parquet files ✅ Compatible with Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/Srijan-Chakraborty/OCR-Finetuning-EN-Dataset.
Upload folder using huggingface_hub
Update README.md
Update README.md
Update README.md
Delete Readme.md
Upload folder using huggingface_hub
initial commit
