CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01QCRI /ImageEval-ArabicNLP26 ImageEval-ArabicNLP26 👁️ ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026. It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation. The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.audio10K<n<100K4 likes2.5k downloads24d agoHugging Face02sboughorbel /tinystories_dataset_arabictabular1M<n<10M1 likes412 downloads2y agoHugging Face03oddadmix /arabic-guardrail Arabic Guardrail — 250,842 rows, 12 classes Defensive dataset for training Arabic prompt-safety classifiers. Each row is an incoming user message and which of 12 safety classes it belongs to. بالعربية: مجموعة بيانات عربية لتدريب نماذج تصنّف الرسائل الواردة قبل وصولها للمساعد الذكي. Arabic guardrails were a gap. Hugging Face searches for Arabic jailbreak / safety / prompt-injection datasets return zero results, and the one Arabic guardrail model… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-guardrail.tabulartext-classification100K<n<1M0 likes192 downloads24d agoHugging Face04oddadmix /arabic-prompt-routing Arabic Prompt Routing — توجيه عربي صفري 233,720 rows. Each row is a text, a set of free-text categories, and which one it belongs to. Categories are arbitrary Arabic — the point is a model that routes into a label set it has never seen. بالعربية: مجموعة بيانات عربية لتوجيه النصوص إلى فئات يكتبها المستخدم بلغة طبيعية. الفئات ليست ثابتة، والهدف نموذج يوجّه إلى فئات لم يرها أثناء التدريب. Arabic counterpart to the task in LiquidAI/LFM2.5-Encoder-350M-Prompt-Router. split… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-prompt-routing.tabularzero-shot-classification100K<n<1M0 likes122 downloads20d agoHugging Face05humain-ai /AraMath 🗂️ Dataset: AraMath 📋 Description: AraMath comprises 605 multiple-choice questions (MCQs) adapted from ArMath (Alghamdi et al., 2022), a dataset of mathematical word problems where each solution corresponds to a solvable equation. We reformulated the original problems into MCQ format to support structured evaluation. 🌐 Language(s): Arabic | 🧠 Task Category: Multiple Choice Questions 🔗 Hugging Face Usage Example from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/humain-ai/AraMath.tabularmultiple-choicen<1K1 likes119 downloads1y agoHugging Face06syamjithnk /arabic-corpus-audit Arabic Corpus Integrity Audit Author: Syamjith NK Date: 9 September 2026 · corrected 13 September 2026 Tool: arabic-lint 0.5.0 Correction, 13 September 2026. An earlier version of this card said the labels in Yousefmd/arabic_ocr_dataset were stored in visual order, and called that the full reshape + bidi signature. That was wrong. Only the shaping step ran; the words are in logical order and plain NFKC recovers them. What was measured, and stands, is that the labels store… See the full description on the dataset page: https://huggingface.co/datasets/syamjithnk/arabic-corpus-audit.tabulartext-classificationn<1K0 likes107 downloads10d agoHugging Face07arashakb /openvla-oft-libero-calib-datatabularn<1K0 likes91 downloads5mo agoHugging Face08oddadmix /arabic-rule-checking Arabic Rule Checking — قواعد ونصوص عربية بأحكام محسوبة 172,488 labelled (text, rule) pairs in Arabic. Each row asks one question: does this text satisfy this rule? The answer is مطابق or مخالف. بالعربية: مجموعة بيانات عربية للتحقق من مطابقة النصوص لقواعد مكتوبة بلغة طبيعية. كل صف يحتوي على نص وقاعدة وحكم محسوب آليًا، وليس رأي نموذج. split pairs texts train 159,240 48,030 validation 13,248 2,002 Built from 50,062 generated Arabic texts across 12 document types… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rule-checking.tabulartext-classification100K<n<1M0 likes89 downloads22d agoHugging Face09webxos /arachnid_RL ___ ___ ___ ___ ___ ___ /\ \ /\ \ /\ \ /\__\ /\ \ /\ \ _____ /::\ \ /::\ \ /::\ \ /:/ / \:\ \ \:\ \ ___ /::\ \ /:/\:\ \ /:/\:\__\ /:/\:\ \ /:/ / \:\ \ \:\ \ /\__\ /:/\:\ \ /:/ /::\ \ /:/ /:/ / /:/ /::\ \ /:/ /… See the full description on the dataset page: https://huggingface.co/datasets/webxos/arachnid_RL.tabularreinforcement-learning1K<n<10K2 likes80 downloads7mo agoHugging Face10albaz2000 /arabic-itsm-dataset Arabic ITSM Dataset A synthetic dataset of 10,000 Arabic IT support tickets, labeled with a structured 3-level ITSM taxonomy, generated using LLMs, and validated programmatically before release. Tickets are written in Egyptian Arabic (عامية مصرية) and cover the full range of helpdesk scenarios: access issues, network problems, hardware faults, software errors, security incidents, and service requests. Arabic technical vocabulary is mixed with English terms as they naturally… See the full description on the dataset page: https://huggingface.co/datasets/albaz2000/arabic-itsm-dataset.tabulartext-classification10K<n<100K0 likes79 downloads20d agoHugging Face11oddadmix /arabic-rag-chat-8k-eval arabic-rag-chat-8k-eval Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models: the test split, every model's raw replies, every judge verdict, and the rendered report for each. Thirteen judged models, all scored on the same 1,651 prompts by the same judge at temperature 0.0, so the comparison below is like-for-like and can be recomputed offline without a GPU or a judge server. This is the measurement half of oddadmix/100M-8192-Nawah-dsv4; the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.tabularquestion-answeringn<1K0 likes72 downloads1mo agoHugging Face12Shaer-AI-2 /shaer-eval-raw-gpt2-small-arabic-poetry Shaer Evaluation Results Models: gpt2_small_arabic_poetry Source dataset: Shaer-AI/shaer-sft-test-generations-k5 Rows: 3481 Validation passed: True Scored rows included: True Dataset repo: Shaer-AI/shaer-eval-raw-gpt2-small-arabic-poetry Files generations.jsonl: raw generation rows generations.csv: raw generation rows in CSV generations_scored.jsonl: raw rows plus meter/count evaluation validation.json: validation summary generations_scored.csv: scored rows in… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/shaer-eval-raw-gpt2-small-arabic-poetry.tabular1K<n<10K0 likes55 downloads4mo agoHugging Face13oddadmix /arabic-rag-support-25K Arabic RAG customer-support scenarios (27,927 rows) Synthetic Modern Standard Arabic customer-support scenarios for training small RAG answerers, distilled from unsloth/gemma-4-31B-it-NVFP4 on a local vLLM. Built as the training set for oddadmix/Nawah-50M-RAG-Support. Each row: a customer question + the knowledge-base chunks of one fictional company (products, prices, policies, FAQ entries) + the ideal grounded agent answer. One generation request invents one company KB and 4 QA… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-support-25K.tabularquestion-answering10K<n<100K0 likes49 downloads1mo agoHugging Face14medyas /arabic-ocr-printed-500k Arabic Printed OCR Lines — Synthetic, 500k A general-purpose printed Arabic text-line recognition corpus: 500,000 train + 2,000 val line images with labels, built to fine-tune line-recognition models (PaddleOCR PP-OCR rec CTC/MultiHead, TrOCR, etc.). Real line-crop printed-Arabic data does not exist at this scale on the Hub, so this corpus is rendered synthetically with diverse fonts + real Arabic text and a documented label/decoding contract. Why this exists… See the full description on the dataset page: https://huggingface.co/datasets/medyas/arabic-ocr-printed-500k.tabularimage-to-textn<1K0 likes43 downloads3mo agoHugging Face15nizarun /English_Arabic_Translation_Pairs English · العربية English→Arabic Technical & Reasoning Translation Dataset High-quality English → Modern Standard Arabic translation pairs focused on native-English educational, scientific, and reasoning content. English source text is drawn from real corpora (FineWeb-Edu and three NVIDIA reasoning datasets); Arabic translations are produced by DeepSeek-v4-flash under a strict translation-only prompt that preserves notation, numbers, formulas, code, and citations. Pairs… See the full description on the dataset page: https://huggingface.co/datasets/nizarun/English_Arabic_Translation_Pairs.tabulartranslation100K<n<1M0 likes39 downloads2mo agoHugging Face16Shams03 /Ara-Egy-Medical-QAtabularquestion-answering10K<n<100K0 likes35 downloads6mo agoHugging Face17oddadmix /arabic-rag-chat-grpo-5K Arabic multi-turn RAG conversations — GRPO pool (5,259 conversations) The reinforcement-learning half of oddadmix/arabic-rag-chat-30K: same generator, same validator, same schema, disjoint companies. It exists so GRPO explores fresh knowledge bases instead of taking a second pass over material the SFT already memorised. conversations turns companies this pool 5,259 14,018 309 Company-disjointness is exact and verified: this pool shares zero company_id values… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-grpo-5K.tabularquestion-answering1K<n<10K1 likes35 downloads1mo agoHugging Face18madilcy /arabic-medical-qa-MERGED-MAQA-MMMLU-MItabular10K<n<100K0 likes34 downloads1y agoHugging Face19ISTNetworks /arabic_alpaca_modeltabularquestion-answering1K<n<10K2 likes33 downloads3y agoHugging Face20abdo1819 /arabic-english-code-switching-review-annotations Review Annotations for Arabic-English Code-Switching Speech This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts. The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index. Coverage and outcomes The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.tabularautomatic-speech-recognition10K<n<100K0 likes33 downloads2mo agoHugging Face21LordTenson /arabic_eou_sada_dataset Arabic EOU SADA Dataset (Saudi Dialect) 414,053 conversational Arabic utterances annotated for End-of-Utterance (EOU) detectionStrong focus on natural Saudi dialect (خليجي / نجدي / حجازي) Task Binary classification: label = 1 → End of speaker turn (EOU) label = 0 → Speaker will continue Columns text: Arabic transcription label: 0 or 1 silence_after_seconds: pause duration after this segment split: train | validation | test (already included)… See the full description on the dataset page: https://huggingface.co/datasets/LordTenson/arabic_eou_sada_dataset.tabular100K<n<1M0 likes31 downloads10mo agoHugging Face22Arailym-tleubayeva /LegalRAG Kazakh Legal Text Chunks Dataset Summary Kazakh Legal Text Chunks is a processed corpus of official legal texts of the Republic of Kazakhstan, prepared for retrieval-augmented generation (RAG), legal information retrieval, and grounded legal question answering in the Kazakh language. The dataset contains structure-preserving text chunks derived from publicly available legal and normative documents. It is intended for research and development in: legal retrieval, legal QA… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/LegalRAG.tabularquestion-answering10K<n<100K0 likes31 downloads7mo agoHugging Face23Shaer-AI-2 /shaer-eval-raw-gpt2-medium-arabic-poetry Raw Shaer Continuation Generations - gpt2_medium_arabic_poetry This dataset contains cumulative raw continuation generations for gpt2_medium_arabic_poetry. The main table is data/test.jsonl; it includes source/prompt/reference fields plus the model output in generated_text and raw_generated_text. Current uploaded progress target: 3481 rows out of 3481. Scored upload: 1. When scored upload is true, reward/evaluation columns are included in the main table. Artifacts such as… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/shaer-eval-raw-gpt2-medium-arabic-poetry.tabular1K<n<10K0 likes31 downloads4mo agoHugging Face24Aratako /Bluemoon_Top50MB_Sorted_Fixed_ja Bluemoon_Top50MB_Sorted_Fixed_ja SicariusSicariiStuff/Bluemoon_Top50MB_Sorted_Fixedを、GENIAC-Team-Ozaki/karakuri-lm-8x7b-chat-v0.1-awqを用いて日本語に翻訳したロールプレイ学習用データセットです。 LLMの推論にはDeepInfraというサービスを使いました。 翻訳の詳細 3-shots promptingでの翻訳 mistralのtokenizerで出力が8000トークンを超えるまで翻訳 元データセットにある非常に長い対話は上記条件で途中のターンで翻訳を終了しています。 LLM特有の同じ出力が繰り返される現象に遭遇した場合、その時点で該当レコードの翻訳を終了 この結果1ターン未満となったレコード(157件)を削除… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Bluemoon_Top50MB_Sorted_Fixed_ja.tabulartext-generationn<1K3 likes26 downloads2y agoHugging Face25Arabic-Clip-Archive /Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-SigLIP-512_validationimage1K<n<10K0 likes23 downloads3y agoHugging Face26oddadmix /arabic-rag-chat-30K Arabic multi-turn RAG customer-support conversations (31,294 conversations) Synthetic Modern Standard Arabic customer-support conversations for training small Arabic RAG assistants. The bulk was distilled from gemini-3.1-flash-lite via the Batch API; a first 2.4% came from unsloth/gemma-4-31B-it-NVFP4 on a local vLLM server before the run was moved off-GPU. Both teachers were given the same prompts and the same validator. Each row is one conversation of 1-5 rounds over one… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-30K.tabularquestion-answering10K<n<100K0 likes23 downloads1mo agoHugging Face27QCRI /AraDICE-HellaSwag AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs Overview The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. In this repository, we present the HellaSwag split of the data. Evaluation We have used lm-harness eval framework to for… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-HellaSwag.tabular10K<n<100K0 likes19 downloads1y agoHugging Face28aramdov /Amazon-2022-quarterliestabularn<1K0 likes15 downloads3y agoHugging Face29ChristopheKhalil /ARAFAgated ARAFA: Arabic Fact-Checking Dataset (MSA) ARAFA is a large publicly available Arabic fact-checking dataset in Modern Standard Arabic (MSA). It contains 181,976 claim–evidence pairs labeled as Supported, Refuted, or Not Enough Information (NEI), automatically generated and validated using large language models from Arabic Wikipedia. The dataset is designed for training and benchmarking Arabic fact-checking models. It covers a wide range of domains, including history, politics… See the full description on the dataset page: https://huggingface.co/datasets/ChristopheKhalil/ARAFA.tabulartext-classification100K<n<1M0 likes13 downloads14h agoHugging Face30Arabic-Clip-Archive /Arabic_dataset_1M_translated_jsonl_format_ViT-B-16-plus-240This translation done using https://huggingface.co/Helsinki-NLP/opus-mt-en-ar image100K<n<1M0 likes12 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.