datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Arabic-VLM-Full-Pearl
💎 The Arabic VLM Dataset (Full Pearl Edition)
This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper.
Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.ALLaVA-4V-Arabic
ALLaVA-4V for Arabic
This is the Arabic version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Arabic through ChatGPT and instructed ChatGPT not to translate content related to OCR.
The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V.
Citation
If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of Hong… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Arabic.aljazeera-news-arabic
Al Jazeera Arabic News Articles Dataset
A dataset of 9,758 Arabic news articles scraped from Al Jazeera Arabic (aljazeera.net), covering the period from September 30, 2025 to March 14, 2026.
Dataset Description
Each record contains the full article text, title, publication date, topic labels, and optional image metadata. The articles span 183 unique topic tags across politics, sports, economy, religion, and more.
Supported Tasks
Text Classification… See the full description on the dataset page: https://huggingface.co/datasets/Skiittoo/aljazeera-news-arabic.
