datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mixed-Arabic-Datasets-Repo
Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"
The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts
Dataset Description
The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.smolkalam-arabic-conversational-sft
SmolKalam
SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets.
Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.FineWeb-Edu-Arabic-24M
English
العربية
FineWeb-Edu Arabic 24M
An Arabic-only pretraining corpus of 24,794,425 complete documents, translated from the sample-350BT configuration of FineWeb-Edu. It contains 34.86 billion Arabic tokenizer tokens and preserves the original FineWeb-Edu document IDs, source scores, and detailed translation diagnostics.
Highlight
Saudi architecture shaped by place. A well-translated tour of how builders in Najd, the Gulf coast, Hejaz, and Asir adapted local… See the full description on the dataset page: https://huggingface.co/datasets/nizarun/FineWeb-Edu-Arabic-24M.Mixed-Arabic-Datasets-Repo
Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"
The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts
Dataset Description
The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Mixed-Arabic-Datasets-Repo.Synthetic-JP-EN-Coding-Dataset-801k
Synthetic-JP-EN-Coding-Dataset-801k
Magpieによって作成したコードSFTデータセットであるAratako/Synthetic-JP-EN-Coding-Dataset-Magpie-69kを元に、Evol-Instructのような手法を用いて複数のinstructionとresonseを生成し拡張して作成した、日英混合801262件のコードSFT用合成データセットです。
日本語: 173849件
英語: 627413件
元のinstructionの作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。
nvidia/Nemotron-4-340B-Instruct
microsoft/Phi-3-medium-4k-instruct
mistralai/Mixtral-8x22B-Instruct-v0.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-EN-Coding-Dataset-801k.arabic-dialect-corpus
Arabic Dialect Corpus
A comprehensive collection of Arabic dialectal text, standardized for Natural Language Processing (NLP) model training, evaluation, and linguistic analysis. This corpus has been meticulously processed to ensure high-quality tokenization and consistent metadata.
Dataset Statistics
Metric
Value
Total Records
127,180
Total Tokens
5,802,324
Average Tokens per Record
45.62
Dialect Categories
5
Changelog… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/arabic-dialect-corpus.arabic_tashkil_dataset
Arabic Tashkil (Diacritization) Dataset 📖✨
Dataset Summary
This is a massive, high-quality, Gold-Standard dataset designed explicitly for training Arabic Automatic Diacritization (Tashkil) AI models (such as ByT5, AraT5, or Custom Transformers).
The dataset contains 1,494,228 heavily vocalized pages (~2.47 GB of data) extracted from Classical Arabic and Islamic texts sourced from Thahabi.org.
To ensure the highest possible ground-truth quality, every single page… See the full description on the dataset page: https://huggingface.co/datasets/freococo/arabic_tashkil_dataset.KazakhLawCorpus-clean
KazakhLawCorpus-clean
Dataset Summary
KazakhLawCorpus-clean is a cleaned, Kazakh-only corpus of legislative documents from the Republic of Kazakhstan. It is a processed derivative of the original Arailym-tleubayeva/KazakhLawCorpus dataset.
The original dataset repository was downloaded from Hugging Face and used as the source for this release. Its laws_metadata.csv file contained 223,245 legislative records with multilingual fields and source-oriented metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhLawCorpus-clean.agent-safety-bench
Agent Safety Bench (ASB)
ASB is a benchmark for evaluating the safety of tool-using LLM agents. Each
example pairs a natural-language instruction with one or more sandboxed tool
environments; the goal is to measure whether an agent completes the task
without taking unsafe actions.
This repository hosts the task data for ASB. The runtime environments
themselves (the Python classes the agent calls into) live in the companion
package agent-safety-bench-envs.
It ships two configs:… See the full description on the dataset page: https://huggingface.co/datasets/aradhye/agent-safety-bench.Mixed-Arabic-Dataset-Main
Dataset Card for "Mixed-Arabic-Dataset"
Mixed Arabic Datasets (MAD)
The Mixed Arabic Datasets (MAD) project provides a comprehensive collection of diverse Arabic-language datasets, sourced from various repositories, platforms, and domains. These datasets cover a wide range of text types, including books, articles, Wikipedia content, stories, and more.
MAD Repo vs. MAD Main
MAD Repo
Versatility: In the MAD Repository (MAD Repo), datasets are made… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Dataset-Main.egyptian-arabic-fake-reviews
🕵️♂️🇪🇬 FREAD: Fake Reviews Egyptian Arabic Dataset
Author: IbrahimAmin, Ismail Fakhr, M. Waleed Fakhr, Rasha Kashef License: MIT Paper: Boosting Arabic Fake Reviews Detection by Integrating Textual and Metadata Features: A Transformer-Based Model Languages: Arabic (Egyptian Dialect)
📚 Dataset Summary
FREAD is designed for detecting fake reviews in Arabic using both textual content and behavioral metadata. It contains 60,000 reviews (50K train / 10K test)… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimAmin/egyptian-arabic-fake-reviews.Math_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.arabic-rag-chat-8k-eval
arabic-rag-chat-8k-eval
Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models:
the test split, every model's raw replies, every judge verdict, and the rendered
report for each. Thirteen judged models, all scored on the same 1,651 prompts
by the same judge at temperature 0.0, so the comparison below is like-for-like
and can be recomputed offline without a GPU or a judge server.
This is the measurement half of
oddadmix/100M-8192-Nawah-dsv4;
the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.arabic-lexicons
islamlab — The Classical Arabic Lexicons
197,731 entries from 136 lexical works — the dictionaries,
the gharīb collections and the technical glossaries — cut so that the
headword is its own column and the article is its own text.
A dictionary sits in the corpus like any other book, but nobody reads one that
way. This is the same material arranged for the thing people actually do with
it: look a word up.
work
author
entries
شمس العلوم ودواء كلام العرب من الكلوم
نشوان… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/arabic-lexicons.arabic-rag-chat-30K
Arabic multi-turn RAG customer-support conversations (31,294 conversations)
Synthetic Modern Standard Arabic customer-support conversations for training
small Arabic RAG assistants. The bulk was distilled from gemini-3.1-flash-lite
via the Batch API; a first 2.4% came from unsloth/gemma-4-31B-it-NVFP4
on a local vLLM server before the run was moved off-GPU. Both teachers were given
the same prompts and the same validator. Each row is one conversation of 1-5
rounds over one… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-30K.arabic-rag-support-25K
Arabic RAG customer-support scenarios (27,927 rows)
Synthetic Modern Standard Arabic customer-support scenarios for training small
RAG answerers, distilled from unsloth/gemma-4-31B-it-NVFP4 on a local vLLM.
Built as the training set for oddadmix/Nawah-50M-RAG-Support.
Each row: a customer question + the knowledge-base chunks of one fictional
company (products, prices, policies, FAQ entries) + the ideal grounded agent
answer. One generation request invents one company KB and 4 QA… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-support-25K.General_Facts_in_English_Arabic_Egyptian_Arabic
🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized)
The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.iterative-dpo-data-for-SimPO-iter2
iterative-dpo-data-for-SimPO-iter2
概要
合成instructionデータであるAratako/Magpie-Tanuki-Instruction-Selected-Evolved-26.5kを元に以下のような手順で作成した日本語Preferenceデータセットです。
開発途中のモデルであるAratako/Llama-Gemma-2-27b-CPO_SimPO-iter1を用いて、temperature=1で回答を5回生成
5個の回答それぞれに対して、Qwen/Qwen2.5-72B-Instruct-GPTQ-Int8を用いて0~5点のスコア付けを実施
1つのinstructionに対する5個の回答について、最もスコアが高いものをchosenに、低いものをrejectedに配置
全て同じスコアの場合や、最も良いスコアが2点以下の場合は除外
ライセンス
本データセットは回答の作成に利用したモデルの関係で以下のライセンスの影響を受けます。
META LLAMA 3.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/iterative-dpo-data-for-SimPO-iter2.SFT-Dataset-For-Self-Taught-Evaluators-iter1arabic-rag-chat-grpo-5K
Arabic multi-turn RAG conversations — GRPO pool (5,259 conversations)
The reinforcement-learning half of
oddadmix/arabic-rag-chat-30K:
same generator, same validator, same schema, disjoint companies. It exists
so GRPO explores fresh knowledge bases instead of taking a second pass over
material the SFT already memorised.
conversations
turns
companies
this pool
5,259
14,018
309
Company-disjointness is exact and verified: this pool shares zero
company_id values… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-grpo-5K.dataset-for-annotation-v2arabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.LegalRAG
Kazakh Legal Text Chunks
Dataset Summary
Kazakh Legal Text Chunks is a processed corpus of official legal texts of the Republic of Kazakhstan, prepared for retrieval-augmented generation (RAG), legal information retrieval, and grounded legal question answering in the Kazakh language.
The dataset contains structure-preserving text chunks derived from publicly available legal and normative documents. It is intended for research and development in:
legal retrieval,
legal QA… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/LegalRAG.iterative-dpo-data-for-ORPO-iter3
iterative-dpo-data-for-ORPO-iter3
概要
合成instructionデータであるAratako/Self-Instruct-Qwen2.5-72B-Instruct-60kを元に以下のような手順で作成した日本語Preferenceデータセットです。
開発途中のモデルであるAratako/Llama-Gemma-2-27b-CPO_SimPO-iter2を用いて、temperature=1で回答を5回生成
5個の回答それぞれに対して、Qwen/Qwen2.5-72B-Instruct-GPTQ-Int8を用いて0~5点のスコア付けを実施
1つのinstructionに対する5個の回答について、最もスコアが高いものをchosenに、低いものをrejectedに配置
全て同じスコアの場合や、最も良いスコアが2点以下の場合は除外
ライセンス
本データセットは回答の作成に利用したモデルの関係で以下のライセンスの影響を受けます。
META LLAMA 3.1 COMMUNITY… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/iterative-dpo-data-for-ORPO-iter3.Bluemoon_Top50MB_Sorted_Fixed_ja
Bluemoon_Top50MB_Sorted_Fixed_ja
SicariusSicariiStuff/Bluemoon_Top50MB_Sorted_Fixedを、GENIAC-Team-Ozaki/karakuri-lm-8x7b-chat-v0.1-awqを用いて日本語に翻訳したロールプレイ学習用データセットです。
LLMの推論にはDeepInfraというサービスを使いました。
翻訳の詳細
3-shots promptingでの翻訳
mistralのtokenizerで出力が8000トークンを超えるまで翻訳
元データセットにある非常に長い対話は上記条件で途中のターンで翻訳を終了しています。
LLM特有の同じ出力が繰り返される現象に遭遇した場合、その時点で該当レコードの翻訳を終了
この結果1ターン未満となったレコード(157件)を削除… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Bluemoon_Top50MB_Sorted_Fixed_ja.arabic-financial-qa_eval
Arabic Financial Q&A Evaluation Dataset
Validation and test splits for evaluating models on Arabic Financial Q&A with analytical and causal reasoning.
Dataset Structure
Format: Simple prompt-answer pairs
Language: Arabic
Domain: Financial reports analysis
Task: Analytical question answering
Fields
id: Unique identifier
prompt: Full prompt with report and question
question: The analytical question
report: The financial report content
answer:… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/arabic-financial-qa_eval.canonical-islamic-corpus
🕌 Canonical Islamic Corpus (Quran + Hadith)
Description
Comprehensive corpus of authentic Islamic texts:
6,236 verses of the Holy Quran from Tanzil (Simple Clean)
315,913 unique hadith matns from four curated collectionsTotal: 322,149 texts with rich metadata.
Prepared by Mullosharaf Arabov for IslamicEval 2026 Shared Task (Subtask 2: Hallucination Detection).
📊 Corpus Statistics
Metric
Value
Total entries
322,149
Quran verses
6… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/canonical-islamic-corpus.Essay-quetions-auto-grading-arabicDataset Overview
The Open Orca Enhanced Dataset is meticulously designed to improve the performance of automated essay grading models using deep learning techniques. This dataset integrates robust data instances from the FLAN collection, augmented with responses generated by GPT-3.5 or GPT-4, creating a diverse and context-rich resource for training models.
Dataset Structure
The dataset is structured in a tabular format, with the following key fields:
id: A unique identifier for each data… See the full description on the dataset page: https://huggingface.co/datasets/mohamedemam/Essay-quetions-auto-grading-arabic.neo_ara_v2arabic-player-stats
👤 YallaShoot — إحصاءات اللاعبين العرب
مجموعة بيانات تضم إحصاءات تفصيلية للاعبي كرة القدم العرب والمحترفين في الدوريات العربية، مُعدَّة لتدريب نماذج تحليل الأداء الرياضي.
📌 وصف مجموعة البيانات
تشمل هذه المجموعة بيانات موسمية تفصيلية للاعبين في:
🇸🇦 دوري روشن السعودي للمحترفين
🇪🇬 الدوري المصري الممتاز
🌍 المنتخبات الوطنية العربية
🏆 المحترفون العرب في الدوريات الأوروبية
📂 هيكل البيانات
العمود
النوع
الوصف
player_id
string
معرّف اللاعب
name_ar… See the full description on the dataset page: https://huggingface.co/datasets/yallashoot/arabic-player-stats.
