CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01QCRI /ImageEval-ArabicNLP26 ImageEval-ArabicNLP26 👁️ ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026. It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation. The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.audio10K<n<100K4 likes2.5k downloads26d agoHugging Face02Arabic-NLP-2026 /context-aware-arabic-to-english-model-with-register Context-Aware Arabic Dialect Translation Dataset This repository contains the dataset and code for the paper "Context-Aware Dialectal Arabic Machine Translation with Interactive Region and Register Selection" (Anonymous Submission). Contents context_aware_en_ar_v2.ipynb: The main Google Colab notebook used for training and evaluation. balanced_dataset_ready.csv: The full augmented dataset (57,600 sentence pairs) produced by our RBDA pipeline. train_dataset.csv: The… See the full description on the dataset page: https://huggingface.co/datasets/Arabic-NLP-2026/context-aware-arabic-to-english-model-with-register.text10K<n<100K1 likes70 downloads9mo agoHugging Face03BDas /ArabicNLPDatasetThe dataset, prepared in Arabic, includes 10.000 tests, 10.000 validations and 80000 train data. The data is composed of customer comments and created from e-commerce sites.text-classification100K<n<1M2 likes31 downloads2y agoHugging Face04ArabicNLPWorld /arabic-russian-scientific-translationsgated Arabic–Russian Scientific Translation Corpus Description This dataset provides parallel translations of scientific and medical texts from Arabic (original) and English (source) into Russian, generated by two state‑of‑the‑art language models: Gemma 3:4B (Google) LLaMA 3.1:8B (Meta) The corpus is built from four established Arabic–English corpora (see Sources below) and is intended for machine translation, model evaluation, and linguistic research.… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-russian-scientific-translations.texttranslation10K<n<100K0 likes26 downloads3mo agoHugging Face05ArabicNLPWorld /canonical-islamic-corpusgated 🕌 Canonical Islamic Corpus (Quran + Hadith) Description Comprehensive corpus of authentic Islamic texts: 6,236 verses of the Holy Quran from Tanzil (Simple Clean) 315,913 unique hadith matns from four curated collectionsTotal: 322,149 texts with rich metadata. Prepared by Mullosharaf Arabov for IslamicEval 2026 Shared Task (Subtask 2: Hallucination Detection). 📊 Corpus Statistics Metric Value Total entries 322,149 Quran verses 6… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/canonical-islamic-corpus.tabulartext-generation100K<n<1M0 likes21 downloads3mo agoHugging Face06ArabicNLPWorld /arabic-wikipedia-wikibooks-corpusgated Arabic Wiki Corpus - Parquet Format Dataset Description This dataset contains CLEANED Arabic text from Wikipedia and Wikibooks in Parquet format (efficient, fast, columnar storage). Key Features ✅ Parquet format - faster loading, smaller size, columnar storage ✅ Fully cleaned - no markup, no HTML, no references ✅ Arabic normalized - alef, ya, ta marbuta normalized ✅ Ready for ML/NLP/LLM - use directly without preprocessing Statistics Metric… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-wikipedia-wikibooks-corpus.tabulartext-generation1M<n<10M1 likes15 downloads4mo agoHugging Face07ArabicNLPWorld /arabic-russian-parallel-corpusgated Arabic‑Russian Parallel Corpus A parallel corpus for Arabic–Russian language pairs. Each record contains an Arabic sentence/phrase, its Russian translation, and the source of the pair.The dataset has been cleaned, deduplicated, and source names normalized to lowercase. 📊 Dataset Statistics Overview Metric Value Total entries 116,393 Unique Arabic strings 116,124 Unique Russian strings 116,152 Unique sources 6 Data… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-russian-parallel-corpus.tabular100K<n<1M0 likes15 downloads3mo agoHugging Face08ArabicNLPWorld /holy-qurangated 🕌 The Holy Quran — Tanzil Simple Clean Description Complete text of the Holy Quran (6,236 ayahs) in Parquet format.Source: Tanzil Project — Simple Clean text.Prepared by Mullosharaf Arabov for IslamicEval 2026 Shared Task (Subtask 2: Hallucination Detection). 📊 Dataset Statistics Metric Value Total ayahs 6,236 Total surahs 114 Total juz 30 Total words 82,627 Total letters 332,837 Avg words/ayah 13.2 Avg letters/ayah 53.4… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/holy-quran.tabulartext-generation1K<n<10K0 likes9 downloads3mo agoHugging Face09ArabicNLPWorld /arabic-nlp-corpusgated Dataset Card for Arabic NLP Corpus Legal Notice & Rights ⚠️ Important Legal Information This dataset contains only bibliographic metadata (titles, abstracts, author names, affiliations, DOIs, citation counts, etc.) aggregated from publicly available sources. The copyright and intellectual property rights to the original scholarly texts (including the full text of papers) belong to their respective authors, publishers, or institutions. ArabicNLPWorld does not claim… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-nlp-corpus.tabulartext-classification1K<n<10K0 likes9 downloads1mo agoHugging Face10ArabicNLPWorld /arabic-islamic-hallucination-syntheticgated 🧪 IslamicEval 2026 Synthetic Subtask 2 Dataset Description Synthetic training dataset for IslamicEval 2026 Subtask 2 (Hallucination Identification).Created from the Canonical Islamic Corpus by applying realistic Arabic/Islamic distortions. Correct examples: original canonical texts. Incorrect examples: texts with one or more realistic errors (word replacements, swaps, deletions, character changes, negation removal, number changes, etc.). 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-islamic-hallucination-synthetic.tabulartext-classification100K<n<1M1 likes7 downloads3mo agoHugging Face11ArabicNLPWorld /arabic-russian-translation-corpusgated Dataset Card: Arabic-Russian Translation Corpus Dataset Details Dataset Description This is a large-scale Arabic–Russian parallel corpus containing 15,467,945 sentence pairs aggregated from diverse sources: open subtitle and document corpora (OPUS), TED Talks, lexicographic dictionaries, religious texts (Quran, hadith collections, Bible), conversational phrasebooks, community-contributed examples, and news articles from political and diplomatic open… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-russian-translation-corpus.tabulartranslation10M<n<100M0 likes3 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.