CoolFace
20 results

sard

riotu-lab /SARD SARD: Synthetic Arabic Recognition Dataset Overview SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts. Key Features… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/SARD.image-to-text100K<n<1M14 likes33k downloads4mo agoHugging Facesardinelab /MF2tabularvisual-question-answeringn<1K9 likes673 downloads1y agoHugging Facevrinda2712 /SARD SARD: Synthetic Arabic Recognition Dataset Overview SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts. Key Features Massive… See the full description on the dataset page: https://huggingface.co/datasets/vrinda2712/SARD.image-to-text100K<n<1M0 likes442 downloads8mo agoHugging Facecaoxuhao /SARD SARD: Synthetic Arabic Recognition Dataset Overview SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts. Key Features Massive… See the full description on the dataset page: https://huggingface.co/datasets/caoxuhao/SARD.image-to-text100K<n<1M0 likes295 downloads4mo agoHugging Faceriotu-lab /SARD-Extended SARD: Synthetic Arabic Recognition Dataset Overview SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts. Key Features Massive… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/SARD-Extended.image-to-text9 likes290 downloads4mo agoHugging Facesardinelab /DocBlocks Dataset Card for DocBlocks DocBlocks is a high-quality, multilingual document-level machine translation (MT) dataset designed to fine-tune large language models (LLMs) on long-context translation tasks. Unlike traditional sentence-level datasets, it contains full documents with natural discourse structures and contextual alignment, helping models maintain coherence, consistency, and high translation quality across longer texts. Curated by: Instituto Superior Técnico, Instituto de… See the full description on the dataset page: https://huggingface.co/datasets/sardinelab/DocBlocks.texttranslation100K<n<1M4 likes132 downloads1y agoHugging Face