datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AraDiCE
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
Overview
The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic.
As part of the supplemental materials, we have selected a few datasets (see below) for the reader to review. We will make the full… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE.Gutenberg-Arabic-OCR-HTML-Pages
Gutenberg Arabic HTML-Page Dataset
📖 Dataset Description
The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language.
The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.testccs_synthetic_translated_arabic_processedccs_synthetic_translated_arabicThe columns inside the dataset as follows:
index
url
caption_en
caption_ar
The dataset size is 12556500 rows × 4 columns
ArabicConceptualCaptions3M
Arabic Translated Conceptual Captions Dataset
Overview
This dataset consists of conceptual captions translated into Arabic using the Google Translate API. It serves as a resource for researchers and developers interested in exploring the vision-language tasks and biases introduced during the translation process.
Dataset Information
Source Dataset: Conceptual Captions
Translation Tool: Google Translate API
Translation Language: English to Arabic… See the full description on the dataset page: https://huggingface.co/datasets/LinaAlhuri/ArabicConceptualCaptions3M.
