CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mesolitica /Malay-Dialect-Instructions Malay dialect instruction including coding Negeri Sembilan QA public transport QA, Coding CUDA coding, Kedah QA infra QA, Coding Rust coding, Kelantan QA Najib Razak QA, Coding Go coding, Perak QA Anwar Ibrahim QA, Coding SQL coding, Pahang QA Pendatang asing QA, Coding Typescript coding, Terengganu… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malay-Dialect-Instructions.texttext-generation10K<n<100K5 likes2.1k downloads2y agoHugging Face02ISLAM-PO /arab-dialects-20-countries-3m Arab Dialects Dataset - 20 Countries A large-scale Arabic dialects dataset covering 20 Arab countries, 7 content types per country, 3,000,000 records, 140 JSONL files, 12.07 GB. UTF-8 JSONL, ready for Hugging Face Datasets. 1. Contents 1. Contents 2. Dataset Summary 3. Repository Map 4. Countries Table (20 folders) 5. Data Types Table (7 files) 6. Record Schema 7. Loading and Usage 8. Generation and Reproduction 9. Considerations and Limitations 10. Contributors… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arab-dialects-20-countries-3m.texttext-generation1M<n<10M0 likes586 downloads19d agoHugging Face03surrey-nlp /dialect-preferences DiaLLM — Pooled Preference Dataset (Implicit Thread) Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main). 45,690 preference pairs, pooling all three variety-specific sets (Australian, Northern British, Indian) without variety targeting. Used for implicit-thread DPO training, where the three varieties are pooled rather than targeted individually, preserving the variety-agnostic objective of that thread.… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/dialect-preferences.tabulartext-generation10K<n<100K0 likes233 downloads29d agoHugging Face04HeshamHaroon /saudi-dialect-conversations Saudi Najdi Dialect Conversations A curated dataset of 3,545 multi-turn conversations in Saudi Najdi Arabic dialect (the dialect spoken in Riyadh, Qassim, and central Najd region). Designed for Supervised Fine-Tuning (SFT) of Arabic language models. Dataset Details Metric Value Total conversations 3,545 Total turns 22,536 Average turns per conversation 6.4 Complexity distribution Simple: 31%, Intermediate: 38%, Advanced: 31% Topics covered 18 categories… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/saudi-dialect-conversations.texttext-generation1K<n<10K20 likes210 downloads7mo agoHugging Face05dialect-ai /shironaam Dataset Card for Shironaam Corpus Dataset Summary Automatic headline generation systems have the potential to assist editors in finding interesting headlines to attract visitors or readers. However, the performance of headline generation systems remains challenging due to the unavailability of sufficient parallel data for low-resource languages like Bengali. We provide Shironaam, a large-scale news headline generation dataset of a low-resource language i.e., Bengali… See the full description on the dataset page: https://huggingface.co/datasets/dialect-ai/shironaam.texttext-generation100K<n<1M6 likes192 downloads3y agoHugging Face06skilledu /Malay-Dialect-Instructions Malay dialect instruction including coding Negeri Sembilan QA public transport QA, Coding CUDA coding, Kedah QA infra QA, Coding Rust coding, Kelantan QA Najib Razak QA, Coding Go coding, Perak QA Anwar Ibrahim QA, Coding SQL coding, Pahang QA Pendatang asing QA… See the full description on the dataset page: https://huggingface.co/datasets/skilledu/Malay-Dialect-Instructions.texttext-generation10K<n<100K0 likes182 downloads4mo agoHugging Face07dataflare /arabic-dialect-corpus Arabic Dialect Corpus A comprehensive collection of Arabic dialectal text, standardized for Natural Language Processing (NLP) model training, evaluation, and linguistic analysis. This corpus has been meticulously processed to ensure high-quality tokenization and consistent metadata. Dataset Statistics Metric Value Total Records 127,180 Total Tokens 5,802,324 Average Tokens per Record 45.62 Dialect Categories 5 Changelog… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/arabic-dialect-corpus.tabulartext-generation100K<n<1M1 likes158 downloads8mo agoHugging Face08salmane11 /Dialect2SQL Dialect2SQL Dataset Description Dialect2SQL is a novel dataset designed for the Text-to-SQL task in Arabic dialects, with a particular focus on Moroccan Darija.It provides natural language questions written in Darija, paired with corresponding SQL queries and database schemas.The dataset enables research on low-resource natural language interfaces to databases (NLIDB) in non-standard Arabic varieties. Dataset Summary Dialect2SQL aims to bridge the gap… See the full description on the dataset page: https://huggingface.co/datasets/salmane11/Dialect2SQL.text-generation1K<n<10K0 likes88 downloads8mo agoHugging Face09fr3on /arabic-dialect-corpus 🇪🇬🇸🇦 Arabic Dialect Corpus (Egyptian & Saudi) Dataset Description This dataset contains 150K+ natural, informal Arabic text samples scraped from high-engagement YouTube discussions. It specifically targets Egyptian (EG) and Saudi (SA) dialects, filling a critical gap in resources for training LLMs on colloquial Arabic (Ammiya) rather than just Modern Standard Arabic (MSA). Languages Primary Dialects: Egyptian Arabic (EG) - Cairene and regional Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/fr3on/arabic-dialect-corpus.texttext-generation1M<n<10M1 likes87 downloads8mo agoHugging Face10statworx /swiss-dialects Dataset Card for ArchiMod Corpus Dataset Summary The ArchiMob corpus represents German linguistic varieties spoken within the territory of Switzerland. This corpus is the first electronic resource containing long samples of transcribed text in Swiss German, intended for studying the spatial distribution of morphosyntactic features and for natural language processing. Languages Swiss-German Dataset Structure Data Instances { 'sentence':… See the full description on the dataset page: https://huggingface.co/datasets/statworx/swiss-dialects.texttext-generation1K<n<10K1 likes69 downloads4y agoHugging Face11Pawitt /dialectical-reasoningA specialised dialectical reasoning dataset. contain { Thesis:, Antithesis:, Synthesis: }. Domain are math, science, creative writing texttext-generation100K<n<1M1 likes37 downloads10mo agoHugging Face12HeshamHaroon /arabic-dialect-dpo Arabic Dialect DPO Dataset - Egyptian & Saudi The first large-scale Arabic dialect preference dataset for DPO/ORPO/GRPO alignment training. Contains 22,538 preference triples across two major Arabic dialects: Egyptian (Masry) and Saudi (Najdi). Dataset Summary Config Dialect Rows Language Code egyptian Egyptian Arabic (مصري) 11,038 ar-EG saudi Saudi Arabic (سعودي نجدي) 11,500 ar-SA Total 22,538 Usage from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-dialect-dpo.texttext-generation10K<n<100K0 likes35 downloads7mo agoHugging Face13CNTXTAI0 /arabic_dialects_question_and_answerData Content The file provided: Q/A Reasoning dataset contains the following columns: ID # : Denotes the reference ID for: a. Question b. Answer to the question c. Hint d. Reasoning e. Word count for items a to d above Dialects: Contains the following dialects in separate columns: a. English b. MSA c. Emirati d. Egyptian e. Levantine Syria f. Levantine Jordan g. Levantine Palestine h. Levantine Lebanon Data Generation Process The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.tabularquestion-answeringn<1K6 likes30 downloads2y agoHugging Face14yrrhall /arabic-dialect-corpus 🇪🇬🇸🇦 Arabic Dialect Corpus (Egyptian & Saudi) Dataset Description This dataset contains 150K+ natural, informal Arabic text samples scraped from high-engagement YouTube discussions. It specifically targets Egyptian (EG) and Saudi (SA) dialects, filling a critical gap in resources for training LLMs on colloquial Arabic (Ammiya) rather than just Modern Standard Arabic (MSA). Languages Primary Dialects: Egyptian Arabic (EG) - Cairene and regional… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/arabic-dialect-corpus.texttext-generation1M<n<10M0 likes29 downloads4mo agoHugging Face15morgendigital /dialect-at-tirol Dataset of Tyrolean Dialect (Austria) This dataset contains 200+ words used in Tirol (Austria), together with their German translation and (optional) meaning. text-generationn<1K1 likes27 downloads3y agoHugging Face16premio-ai /TheArabicPile_Dialects The Arabic Pile Introduction: The Arabic Pile is a comprehensive dataset meticulously designed to parallel the structure of The Pile and The Nordic Pile. Focused on the Arabic language, the dataset encompasses a vast array of linguistic nuances, incorporating both Modern Standard Arabic (MSA) and various Levantine, North African, and Egyptian dialects. Tailored for the training and fine-tuning of large language models, the dataset consists of 13 subsets, each uniquely… See the full description on the dataset page: https://huggingface.co/datasets/premio-ai/TheArabicPile_Dialects.texttext-generation100K<n<1M0 likes26 downloads3y agoHugging Face17Rabe3 /saudi-dialect-rag Saudi Dialect RAG Fine-Tuning Dataset A RAG-formatted fine-tuning dataset for Saudi Arabic dialect, built from HeshamHaroon/saudi-dialect-conversations. Format Each example follows the LlamaFactory Alpaca format: Field Description instruction System prompt + MSA context paragraph + optional conversation history + question input Always empty string output Assistant reply in Saudi dialect How it was built Loaded source multi-turn Saudi… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/saudi-dialect-rag.textquestion-answering10K<n<100K0 likes25 downloads7mo agoHugging Face18levantdata /jordanian-dialect-sample-v1 Jordanian Dialect Sample (v1) Levant AI is building dialect-authentic Arabic training data for the Levantine region, starting with Jordanian/Shami dialect — collected natively, not translated from Modern Standard Arabic or English. This is our first public sample: 30 examples spanning three categories that reflect real gaps in current Arabic AI training data: Categories general_conversation (10 examples) — everyday natural Jordanian dialect exchanges… See the full description on the dataset page: https://huggingface.co/datasets/levantdata/jordanian-dialect-sample-v1.texttext-generationn<1K0 likes23 downloads2mo agoHugging Face19furquan /dialectic-preferences-bias-aae-sae-parallel Dialectic Preferences Bias Dataset Dataset Description Overview This dataset is part of a research study examining dialectic preference bias in Large Language Models (LLMs). It contains paired sentences in African American English (AAE) and Standard American English (SAE), used to analyze potential biases in language models' treatment of different dialects. The dataset contains two columns: african_american_english: Text samples in African American English… See the full description on the dataset page: https://huggingface.co/datasets/furquan/dialectic-preferences-bias-aae-sae-parallel.texttext-classification1K<n<10K0 likes22 downloads2y agoHugging Face20hikewa /dialectic-reasoning-traces Dialectic Reasoning Traces 255 scored dialectic reasoning traces for training models on integrative resolution under conflicting frames. Instead of list-format pros/cons or generic hedging, these traces teach models to identify real tension, make conditional commitments, and reach specific resolutions. Version Note This dataset contains only v1 traces — the clean, non-fabricating training data. An earlier version on this repo included augmented data from later pipeline… See the full description on the dataset page: https://huggingface.co/datasets/hikewa/dialectic-reasoning-traces.text-generationn<1K0 likes19 downloads6mo agoHugging Face21andreiski /dialectic-sft-against-only-750 Dialectic SFT — Against-Only (750) 750 supervised fine-tuning conversations that teach a model the structured "dialectical" output format: a set of candidate positions [pN] followed by against-claims [cN] against pM: that critique those positions. This is the level-1, against-only stage (only against-claims, no for-claims or deeper tree levels) — it bootstraps the format before GRPO reinforcement learning. Row count 750 rows. Schema One JSON object… See the full description on the dataset page: https://huggingface.co/datasets/andreiski/dialectic-sft-against-only-750.texttext-generationn<1K0 likes19 downloads4mo agoHugging Face22yusufbaykaloglu /Turkish-Dialectical-Reasoning-Dataset-Sokrates-ToT Turkish Dialectical Reasoning Dataset (Sokrates-ToT) The Turkish Dialectical Reasoning Dataset (Sokrates-ToT) is a collection structured in a Tree-of-Thought (ToT) format, based on a multi-persona and dialectical reasoning framework.Inspired by Socrates' method of dialogue, it facilitates deep analysis of complex and multidimensional issues by having AI personas with different expertise interact and ultimately reach a final synthesis. Purpose of the Dataset This… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-Dialectical-Reasoning-Dataset-Sokrates-ToT.textquestion-answering1K<n<10K0 likes17 downloads1y agoHugging Face23Whomstt /irish-english-dialecttexttext-generationn<1K0 likes8 downloads8mo agoHugging Face24farabi-lab /Understanding_Dialect_Textgated 🇰🇿 Kazakh Dialect Analysis and Standardization Dataset Dataset Summary Kazakh Dialect Analysis and Standardization Dataset is a Kazakh-language linguistics dataset designed for dialect identification, dialectal feature analysis, and normalization into standard Kazakh. The dataset contains instruction-style prompts asking the model to identify dialectal or regional language features in a given Kazakh text and explain how they can be converted into standard… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Understanding_Dialect_Text.texttext-generationn<1K0 likes6 downloads2mo agoHugging Face25andreiski /dialectic-rl-questions-10k Dialectic RL Questions (10k) 10,000 real-world dilemma / debate prompts used as the GRPO training prompts for a dialectical-debate model. Each prompt is an open-ended question (advice dilemmas, opinion debates, and general user requests) that the model is trained to answer by generating multiple positions and against-claims in a structured "dialectical" format. The prompts are drawn from public real-world sources: Reddit AITA (r/AmItheAsshole), SHP (Stanford Human Preferences, a… See the full description on the dataset page: https://huggingface.co/datasets/andreiski/dialectic-rl-questions-10k.texttext-generation10K<n<100K0 likes5 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.