CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01M-A-D /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B38 likes2.2k downloads3y agoHugging Face02AdaMLLab /smolkalam-arabic-conversational-sft SmolKalam SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets. Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.tabulartext-generation1M<n<10M3 likes890 downloads1mo agoHugging Face03nizarun /FineWeb-Edu-Arabic-24M English العربية FineWeb-Edu Arabic 24M An Arabic-only pretraining corpus of 24,794,425 complete documents, translated from the sample-350BT configuration of FineWeb-Edu. It contains 34.86 billion Arabic tokenizer tokens and preserves the original FineWeb-Edu document IDs, source scores, and detailed translation diagnostics. Highlight Saudi architecture shaped by place. A well-translated tour of how builders in Najd, the Gulf coast, Hejaz, and Asir adapted local… See the full description on the dataset page: https://huggingface.co/datasets/nizarun/FineWeb-Edu-Arabic-24M.tabulartext-generation10M<n<100M0 likes464 downloads26d agoHugging Face04yrrhall /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B0 likes340 downloads4mo agoHugging Face05Aratako /Synthetic-JP-EN-Coding-Dataset-801k Synthetic-JP-EN-Coding-Dataset-801k Magpieによって作成したコードSFTデータセットであるAratako/Synthetic-JP-EN-Coding-Dataset-Magpie-69kを元に、Evol-Instructのような手法を用いて複数のinstructionとresonseを生成し拡張して作成した、日英混合801262件のコードSFT用合成データセットです。 日本語: 173849件 英語: 627413件 元のinstructionの作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。 nvidia/Nemotron-4-340B-Instruct microsoft/Phi-3-medium-4k-instruct mistralai/Mixtral-8x22B-Instruct-v0.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-EN-Coding-Dataset-801k.tabulartext-generation100K<n<1M17 likes169 downloads2y agoHugging Face06dataflare /arabic-dialect-corpus Arabic Dialect Corpus A comprehensive collection of Arabic dialectal text, standardized for Natural Language Processing (NLP) model training, evaluation, and linguistic analysis. This corpus has been meticulously processed to ensure high-quality tokenization and consistent metadata. Dataset Statistics Metric Value Total Records 127,180 Total Tokens 5,802,324 Average Tokens per Record 45.62 Dialect Categories 5 Changelog… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/arabic-dialect-corpus.tabulartext-generation100K<n<1M1 likes158 downloads8mo agoHugging Face07freococo /arabic_tashkil_dataset Arabic Tashkil (Diacritization) Dataset 📖✨ Dataset Summary This is a massive, high-quality, Gold-Standard dataset designed explicitly for training Arabic Automatic Diacritization (Tashkil) AI models (such as ByT5, AraT5, or Custom Transformers). The dataset contains 1,494,228 heavily vocalized pages (~2.47 GB of data) extracted from Classical Arabic and Islamic texts sourced from Thahabi.org. To ensure the highest possible ground-truth quality, every single page… See the full description on the dataset page: https://huggingface.co/datasets/freococo/arabic_tashkil_dataset.tabulartext-generation1M<n<10M0 likes155 downloads3mo agoHugging Face08Arailym-tleubayeva /KazakhLawCorpus-clean KazakhLawCorpus-clean Dataset Summary KazakhLawCorpus-clean is a cleaned, Kazakh-only corpus of legislative documents from the Republic of Kazakhstan. It is a processed derivative of the original Arailym-tleubayeva/KazakhLawCorpus dataset. The original dataset repository was downloaded from Hugging Face and used as the source for this release. Its laws_metadata.csv file contained 223,245 legislative records with multilingual fields and source-oriented metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhLawCorpus-clean.tabulartext-retrieval100K<n<1M1 likes107 downloads14d agoHugging Face09aradhye /agent-safety-bench Agent Safety Bench (ASB) ASB is a benchmark for evaluating the safety of tool-using LLM agents. Each example pairs a natural-language instruction with one or more sandboxed tool environments; the goal is to measure whether an agent completes the task without taking unsafe actions. This repository hosts the task data for ASB. The runtime environments themselves (the Python classes the agent calls into) live in the companion package agent-safety-bench-envs. It ships two configs:… See the full description on the dataset page: https://huggingface.co/datasets/aradhye/agent-safety-bench.tabulartext-generation1K<n<10K1 likes96 downloads5mo agoHugging Face10M-A-D /Mixed-Arabic-Dataset-Main Dataset Card for "Mixed-Arabic-Dataset" Mixed Arabic Datasets (MAD) The Mixed Arabic Datasets (MAD) project provides a comprehensive collection of diverse Arabic-language datasets, sourced from various repositories, platforms, and domains. These datasets cover a wide range of text types, including books, articles, Wikipedia content, stories, and more. MAD Repo vs. MAD Main MAD Repo Versatility: In the MAD Repository (MAD Repo), datasets are made… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Dataset-Main.tabulartext-generation100K<n<1M7 likes95 downloads3y agoHugging Face11IbrahimAmin /egyptian-arabic-fake-reviews 🕵️‍♂️🇪🇬 FREAD: Fake Reviews Egyptian Arabic Dataset Author: IbrahimAmin, Ismail Fakhr, M. Waleed Fakhr, Rasha Kashef License: MIT Paper: Boosting Arabic Fake Reviews Detection by Integrating Textual and Metadata Features: A Transformer-Based Model Languages: Arabic (Egyptian Dialect) 📚 Dataset Summary FREAD is designed for detecting fake reviews in Arabic using both textual content and behavioral metadata. It contains 60,000 reviews (50K train / 10K test)… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimAmin/egyptian-arabic-fake-reviews.tabulartext-classification10K<n<100K2 likes83 downloads2mo agoHugging Face12miscovery /Math_CoT_Arabic_English_Reasoning Math CoT Arabic English Dataset A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI. Overview Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.tabularquestion-answering1K<n<10K17 likes73 downloads1y agoHugging Face13oddadmix /arabic-rag-chat-8k-eval arabic-rag-chat-8k-eval Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models: the test split, every model's raw replies, every judge verdict, and the rendered report for each. Thirteen judged models, all scored on the same 1,651 prompts by the same judge at temperature 0.0, so the comparison below is like-for-like and can be recomputed offline without a GPU or a judge server. This is the measurement half of oddadmix/100M-8192-Nawah-dsv4; the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.tabularquestion-answeringn<1K0 likes72 downloads1mo agoHugging Face14islamlab /arabic-lexicons islamlab — The Classical Arabic Lexicons 197,731 entries from 136 lexical works — the dictionaries, the gharīb collections and the technical glossaries — cut so that the headword is its own column and the article is its own text. A dictionary sits in the corpus like any other book, but nobody reads one that way. This is the same material arranged for the thing people actually do with it: look a word up. work author entries شمس العلوم ودواء كلام العرب من الكلوم نشوان… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/arabic-lexicons.tabulartext-retrieval100K<n<1M0 likes52 downloads1mo agoHugging Face15oddadmix /arabic-rag-chat-30K Arabic multi-turn RAG customer-support conversations (31,294 conversations) Synthetic Modern Standard Arabic customer-support conversations for training small Arabic RAG assistants. The bulk was distilled from gemini-3.1-flash-lite via the Batch API; a first 2.4% came from unsloth/gemma-4-31B-it-NVFP4 on a local vLLM server before the run was moved off-GPU. Both teachers were given the same prompts and the same validator. Each row is one conversation of 1-5 rounds over one… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-30K.tabularquestion-answering10K<n<100K0 likes50 downloads1mo agoHugging Face16oddadmix /arabic-rag-support-25K Arabic RAG customer-support scenarios (27,927 rows) Synthetic Modern Standard Arabic customer-support scenarios for training small RAG answerers, distilled from unsloth/gemma-4-31B-it-NVFP4 on a local vLLM. Built as the training set for oddadmix/Nawah-50M-RAG-Support. Each row: a customer question + the knowledge-base chunks of one fictional company (products, prices, policies, FAQ entries) + the ideal grounded agent answer. One generation request invents one company KB and 4 QA… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-support-25K.tabularquestion-answering10K<n<100K0 likes49 downloads1mo agoHugging Face17miscovery /General_Facts_in_English_Arabic_Egyptian_Arabic 🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized) The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages: 🌍 English 🇸🇦 Modern Standard Arabic (MSA) 🇪🇬 Egyptian Arabic (Dialect) Each entry includes: The question and answer A category and sub-category Language tag (en, ar, ar_eg) Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.tabularquestion-answering10K<n<100K12 likes44 downloads1y agoHugging Face18Aratako /iterative-dpo-data-for-SimPO-iter2 iterative-dpo-data-for-SimPO-iter2 概要 合成instructionデータであるAratako/Magpie-Tanuki-Instruction-Selected-Evolved-26.5kを元に以下のような手順で作成した日本語Preferenceデータセットです。 開発途中のモデルであるAratako/Llama-Gemma-2-27b-CPO_SimPO-iter1を用いて、temperature=1で回答を5回生成 5個の回答それぞれに対して、Qwen/Qwen2.5-72B-Instruct-GPTQ-Int8を用いて0~5点のスコア付けを実施 1つのinstructionに対する5個の回答について、最もスコアが高いものをchosenに、低いものをrejectedに配置 全て同じスコアの場合や、最も良いスコアが2点以下の場合は除外 ライセンス 本データセットは回答の作成に利用したモデルの関係で以下のライセンスの影響を受けます。 META LLAMA 3.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/iterative-dpo-data-for-SimPO-iter2.tabulartext-generation10K<n<100K1 likes42 downloads2y agoHugging Face19Aratako /SFT-Dataset-For-Self-Taught-Evaluators-iter1tabulartext-generation10K<n<100K1 likes36 downloads2y agoHugging Face20oddadmix /arabic-rag-chat-grpo-5K Arabic multi-turn RAG conversations — GRPO pool (5,259 conversations) The reinforcement-learning half of oddadmix/arabic-rag-chat-30K: same generator, same validator, same schema, disjoint companies. It exists so GRPO explores fresh knowledge bases instead of taking a second pass over material the SFT already memorised. conversations turns companies this pool 5,259 14,018 309 Company-disjointness is exact and verified: this pool shares zero company_id values… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-grpo-5K.tabularquestion-answering1K<n<10K1 likes36 downloads1mo agoHugging Face21Aratako /dataset-for-annotation-v2tabulartext-generation100K<n<1M0 likes32 downloads2y agoHugging Face22CNTXTAI0 /arabic_dialects_question_and_answerData Content The file provided: Q/A Reasoning dataset contains the following columns: ID # : Denotes the reference ID for: a. Question b. Answer to the question c. Hint d. Reasoning e. Word count for items a to d above Dialects: Contains the following dialects in separate columns: a. English b. MSA c. Emirati d. Egyptian e. Levantine Syria f. Levantine Jordan g. Levantine Palestine h. Levantine Lebanon Data Generation Process The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.tabularquestion-answeringn<1K6 likes32 downloads2y agoHugging Face23Arailym-tleubayeva /LegalRAG Kazakh Legal Text Chunks Dataset Summary Kazakh Legal Text Chunks is a processed corpus of official legal texts of the Republic of Kazakhstan, prepared for retrieval-augmented generation (RAG), legal information retrieval, and grounded legal question answering in the Kazakh language. The dataset contains structure-preserving text chunks derived from publicly available legal and normative documents. It is intended for research and development in: legal retrieval, legal QA… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/LegalRAG.tabularquestion-answering10K<n<100K0 likes31 downloads7mo agoHugging Face24Aratako /iterative-dpo-data-for-ORPO-iter3 iterative-dpo-data-for-ORPO-iter3 概要 合成instructionデータであるAratako/Self-Instruct-Qwen2.5-72B-Instruct-60kを元に以下のような手順で作成した日本語Preferenceデータセットです。 開発途中のモデルであるAratako/Llama-Gemma-2-27b-CPO_SimPO-iter2を用いて、temperature=1で回答を5回生成 5個の回答それぞれに対して、Qwen/Qwen2.5-72B-Instruct-GPTQ-Int8を用いて0~5点のスコア付けを実施 1つのinstructionに対する5個の回答について、最もスコアが高いものをchosenに、低いものをrejectedに配置 全て同じスコアの場合や、最も良いスコアが2点以下の場合は除外 ライセンス 本データセットは回答の作成に利用したモデルの関係で以下のライセンスの影響を受けます。 META LLAMA 3.1 COMMUNITY… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/iterative-dpo-data-for-ORPO-iter3.tabulartext-generation10K<n<100K3 likes27 downloads2y agoHugging Face25Aratako /Bluemoon_Top50MB_Sorted_Fixed_ja Bluemoon_Top50MB_Sorted_Fixed_ja SicariusSicariiStuff/Bluemoon_Top50MB_Sorted_Fixedを、GENIAC-Team-Ozaki/karakuri-lm-8x7b-chat-v0.1-awqを用いて日本語に翻訳したロールプレイ学習用データセットです。 LLMの推論にはDeepInfraというサービスを使いました。 翻訳の詳細 3-shots promptingでの翻訳 mistralのtokenizerで出力が8000トークンを超えるまで翻訳 元データセットにある非常に長い対話は上記条件で途中のターンで翻訳を終了しています。 LLM特有の同じ出力が繰り返される現象に遭遇した場合、その時点で該当レコードの翻訳を終了 この結果1ターン未満となったレコード(157件)を削除… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Bluemoon_Top50MB_Sorted_Fixed_ja.tabulartext-generationn<1K3 likes26 downloads2y agoHugging Face26SahmBenchmark /arabic-financial-qa_eval Arabic Financial Q&A Evaluation Dataset Validation and test splits for evaluating models on Arabic Financial Q&A with analytical and causal reasoning. Dataset Structure Format: Simple prompt-answer pairs Language: Arabic Domain: Financial reports analysis Task: Analytical question answering Fields id: Unique identifier prompt: Full prompt with report and question question: The analytical question report: The financial report content answer:… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/arabic-financial-qa_eval.tabularquestion-answeringn<1K0 likes24 downloads9mo agoHugging Face27ArabicNLPWorld /canonical-islamic-corpusgated 🕌 Canonical Islamic Corpus (Quran + Hadith) Description Comprehensive corpus of authentic Islamic texts: 6,236 verses of the Holy Quran from Tanzil (Simple Clean) 315,913 unique hadith matns from four curated collectionsTotal: 322,149 texts with rich metadata. Prepared by Mullosharaf Arabov for IslamicEval 2026 Shared Task (Subtask 2: Hallucination Detection). 📊 Corpus Statistics Metric Value Total entries 322,149 Quran verses 6… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/canonical-islamic-corpus.tabulartext-generation100K<n<1M0 likes23 downloads3mo agoHugging Face28mohamedemam /Essay-quetions-auto-grading-arabicDataset Overview The Open Orca Enhanced Dataset is meticulously designed to improve the performance of automated essay grading models using deep learning techniques. This dataset integrates robust data instances from the FLAN collection, augmented with responses generated by GPT-3.5 or GPT-4, creating a diverse and context-rich resource for training models. Dataset Structure The dataset is structured in a tabular format, with the following key fields: id: A unique identifier for each data… See the full description on the dataset page: https://huggingface.co/datasets/mohamedemam/Essay-quetions-auto-grading-arabic.tabulartext-classification10K<n<100K0 likes22 downloads2y agoHugging Face29giseldo /neo_ara_v2tabulartext-generation10K<n<100K1 likes22 downloads1y agoHugging Face30yallashoot /arabic-player-stats 👤 YallaShoot — إحصاءات اللاعبين العرب مجموعة بيانات تضم إحصاءات تفصيلية للاعبي كرة القدم العرب والمحترفين في الدوريات العربية، مُعدَّة لتدريب نماذج تحليل الأداء الرياضي. 📌 وصف مجموعة البيانات تشمل هذه المجموعة بيانات موسمية تفصيلية للاعبين في: 🇸🇦 دوري روشن السعودي للمحترفين 🇪🇬 الدوري المصري الممتاز 🌍 المنتخبات الوطنية العربية 🏆 المحترفون العرب في الدوريات الأوروبية 📂 هيكل البيانات العمود النوع الوصف player_id string معرّف اللاعب name_ar… See the full description on the dataset page: https://huggingface.co/datasets/yallashoot/arabic-player-stats.tabulartable-question-answeringn<1K0 likes22 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.