sard
Datasets
All datasets matching “sard”SARD
SARD: Synthetic Arabic Recognition Dataset
Overview
SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts.
Key Features… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/SARD.MF2SARD
SARD: Synthetic Arabic Recognition Dataset
Overview
SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts.
Key Features
Massive… See the full description on the dataset page: https://huggingface.co/datasets/vrinda2712/SARD.SARD
SARD: Synthetic Arabic Recognition Dataset
Overview
SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts.
Key Features
Massive… See the full description on the dataset page: https://huggingface.co/datasets/caoxuhao/SARD.SARD-Extended
SARD: Synthetic Arabic Recognition Dataset
Overview
SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts.
Key Features
Massive… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/SARD-Extended.DocBlocks
Dataset Card for DocBlocks
DocBlocks is a high-quality, multilingual document-level machine translation (MT) dataset designed to fine-tune large language models (LLMs) on long-context translation tasks. Unlike traditional sentence-level datasets, it contains full documents with natural discourse structures and contextual alignment, helping models maintain coherence, consistency, and high translation quality across longer texts.
Curated by: Instituto Superior Técnico, Instituto de… See the full description on the dataset page: https://huggingface.co/datasets/sardinelab/DocBlocks.
