yrrhall/Mixed-Arabic-Datasets-Repo
Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Mixed-Arabic-Datasets-Repo.
0340
1---2language:3- ar4size_categories:5- 1B<n<10B6task_categories:7- text-classification8- question-answering9- translation10- summarization11- conversational12- text-generation13- text2text-generation14- fill-mask15pretty_name: Mixed Arabic Datasets (MAD) Corpus16dataset_info:17- config_name: Ara--Ali-C137--Hindawi-Books-dataset18 features:19 - name: BookLink20 dtype: string21 - name: BookName22 dtype: string23 - name: AuthorName24 dtype: string25 - name: AboutBook26 dtype: string27 - name: ChapterLink28 dtype: string29 - name: ChapterName30 dtype: string31 - name: ChapterText32 dtype: string33 - name: AboutAuthor34 dtype: string35 splits:36 - name: train37 num_bytes: 136485425938 num_examples: 4982139 download_size: 49467800240 dataset_size: 136485425941- config_name: Ara--Goud--Goud-sum42 features:43 - name: article44 dtype: string45 - name: headline46 dtype: string47 - name: categories48 dtype: string49 splits:50 - name: train51 num_bytes: 28829654452 num_examples: 13928853 download_size: 14773577654 dataset_size: 28829654455- config_name: Ara--J-Mourad--MNAD.v156 features:57 - name: Title58 dtype: string59 - name: Body60 dtype: string61 - name: Category62 dtype: string63 splits:64 - name: train65 num_bytes: 110192198066 num_examples: 41856367 download_size: 52715412268 dataset_size: 110192198069- config_name: Ara--JihadZa--IADD70 features:71 - name: Sentence72 dtype: string73 - name: Region74 dtype: string75 - name: DataSource76 dtype: string77 - name: Country78 dtype: string79 splits:80 - name: train81 num_bytes: 1916707082 num_examples: 13580483 download_size: 864449184 dataset_size: 1916707085- config_name: Ara--LeMGarouani--MAC-corpus86 features:87 - name: tweets88 dtype: string89 - name: type90 dtype: string91 - name: class92 dtype: string93 splits:94 - name: train95 num_bytes: 194564696 num_examples: 1808797 download_size: 86619898 dataset_size: 194564699- config_name: Ara--MBZUAI--Bactrian-X100 features:101 - name: instruction102 dtype: string103 - name: input104 dtype: string105 - name: id106 dtype: string107 - name: output108 dtype: string109 splits:110 - name: train111 num_bytes: 66093524112 num_examples: 67017113 download_size: 33063779114 dataset_size: 66093524115- config_name: Ara--OpenAssistant--oasst1116 features:117 - name: message_id118 dtype: string119 - name: parent_id120 dtype: string121 - name: user_id122 dtype: string123 - name: created_date124 dtype: string125 - name: text126 dtype: string127 - name: role128 dtype: string129 - name: lang130 dtype: string131 - name: review_count132 dtype: int32133 - name: review_result134 dtype: bool135 - name: deleted136 dtype: bool137 - name: rank138 dtype: float64139 - name: synthetic140 dtype: bool141 - name: model_name142 dtype: 'null'143 - name: detoxify144 dtype: 'null'145 - name: message_tree_id146 dtype: string147 - name: tree_state148 dtype: string149 - name: emojis150 struct:151 - name: count152 sequence: int32153 - name: name154 sequence: string155 - name: labels156 struct:157 - name: count158 sequence: int32159 - name: name160 sequence: string161 - name: value162 sequence: float64163 - name: __index_level_0__164 dtype: int64165 splits:166 - name: train167 num_bytes: 58168168 num_examples: 56169 download_size: 30984170 dataset_size: 58168171- config_name: Ara--Wikipedia172 features:173 - name: id174 dtype: string175 - name: url176 dtype: string177 - name: title178 dtype: string179 - name: text180 dtype: string181 splits:182 - name: train183 num_bytes: 3052201469184 num_examples: 1205403185 download_size: 1316212231186 dataset_size: 3052201469187- config_name: Ara--bigscience--xP3188 features:189 - name: inputs190 dtype: string191 - name: targets192 dtype: string193 splits:194 - name: train195 num_bytes: 4727881680196 num_examples: 2148955197 download_size: 2805060725198 dataset_size: 4727881680199- config_name: Ara--cardiffnlp--tweet_sentiment_multilingual200 features:201 - name: text202 dtype: string203 - name: label204 dtype:205 class_label:206 names:207 '0': negative208 '1': neutral209 '2': positive210 splits:211 - name: train212 num_bytes: 306108213 num_examples: 1839214 - name: validation215 num_bytes: 53276216 num_examples: 324217 - name: test218 num_bytes: 141536219 num_examples: 870220 download_size: 279900221 dataset_size: 500920222- config_name: Ara--miracl--miracl223 features:224 - name: query_id225 dtype: string226 - name: query227 dtype: string228 - name: positive_passages229 list:230 - name: docid231 dtype: string232 - name: text233 dtype: string234 - name: title235 dtype: string236 - name: negative_passages237 list:238 - name: docid239 dtype: string240 - name: text241 dtype: string242 - name: title243 dtype: string244 splits:245 - name: train246 num_bytes: 32012083247 num_examples: 3495248 download_size: 15798509249 dataset_size: 32012083250- config_name: Ara--mustapha--QuranExe251 features:252 - name: text253 dtype: string254 - name: resource_name255 dtype: string256 - name: verses_keys257 dtype: string258 splits:259 - name: train260 num_bytes: 133108687261 num_examples: 49888262 download_size: 58769417263 dataset_size: 133108687264- config_name: Ara--pain--Arabic-Tweets265 features:266 - name: text267 dtype: string268 splits:269 - name: train270 num_bytes: 41639770853271 num_examples: 202700438272 download_size: 22561651700273 dataset_size: 41639770853274- config_name: Ara--saudinewsnet275 features:276 - name: source277 dtype: string278 - name: url279 dtype: string280 - name: date_extracted281 dtype: string282 - name: title283 dtype: string284 - name: author285 dtype: string286 - name: content287 dtype: string288 splits:289 - name: train290 num_bytes: 103654009291 num_examples: 31030292 download_size: 49117164293 dataset_size: 103654009294- config_name: Ary--AbderrahmanSkiredj1--Darija-Wikipedia295 features:296 - name: text297 dtype: string298 splits:299 - name: train300 num_bytes: 8104410301 num_examples: 4862302 download_size: 3229966303 dataset_size: 8104410304- config_name: Ary--Ali-C137--Darija-Stories-Dataset305 features:306 - name: ChapterName307 dtype: string308 - name: ChapterLink309 dtype: string310 - name: Author311 dtype: string312 - name: Text313 dtype: string314 - name: Tags315 dtype: int64316 splits:317 - name: train318 num_bytes: 476926644319 num_examples: 6142320 download_size: 241528641321 dataset_size: 476926644322- config_name: Ary--Wikipedia323 features:324 - name: id325 dtype: string326 - name: url327 dtype: string328 - name: title329 dtype: string330 - name: text331 dtype: string332 splits:333 - name: train334 num_bytes: 10007364335 num_examples: 6703336 download_size: 4094377337 dataset_size: 10007364338- config_name: Arz--Wikipedia339 features:340 - name: id341 dtype: string342 - name: url343 dtype: string344 - name: title345 dtype: string346 - name: text347 dtype: string348 splits:349 - name: train350 num_bytes: 1364641408351 num_examples: 1617770352 download_size: 306420318353 dataset_size: 1364641408354configs:355- config_name: Ara--Ali-C137--Hindawi-Books-dataset356 data_files:357 - split: train358 path: Ara--Ali-C137--Hindawi-Books-dataset/train-*359- config_name: Ara--Goud--Goud-sum360 data_files:361 - split: train362 path: Ara--Goud--Goud-sum/train-*363- config_name: Ara--J-Mourad--MNAD.v1364 data_files:365 - split: train366 path: Ara--J-Mourad--MNAD.v1/train-*367- config_name: Ara--JihadZa--IADD368 data_files:369 - split: train370 path: Ara--JihadZa--IADD/train-*371- config_name: Ara--LeMGarouani--MAC-corpus372 data_files:373 - split: train374 path: Ara--LeMGarouani--MAC-corpus/train-*375- config_name: Ara--MBZUAI--Bactrian-X376 data_files:377 - split: train378 path: Ara--MBZUAI--Bactrian-X/train-*379- config_name: Ara--OpenAssistant--oasst1380 data_files:381 - split: train382 path: Ara--OpenAssistant--oasst1/train-*383- config_name: Ara--Wikipedia384 data_files:385 - split: train386 path: Ara--Wikipedia/train-*387- config_name: Ara--bigscience--xP3388 data_files:389 - split: train390 path: Ara--bigscience--xP3/train-*391- config_name: Ara--cardiffnlp--tweet_sentiment_multilingual392 data_files:393 - split: train394 path: Ara--cardiffnlp--tweet_sentiment_multilingual/train-*395 - split: validation396 path: Ara--cardiffnlp--tweet_sentiment_multilingual/validation-*397 - split: test398 path: Ara--cardiffnlp--tweet_sentiment_multilingual/test-*399- config_name: Ara--miracl--miracl400 data_files:401 - split: train402 path: Ara--miracl--miracl/train-*403- config_name: Ara--mustapha--QuranExe404 data_files:405 - split: train406 path: Ara--mustapha--QuranExe/train-*407- config_name: Ara--pain--Arabic-Tweets408 data_files:409 - split: train410 path: Ara--pain--Arabic-Tweets/train-*411- config_name: Ara--saudinewsnet412 data_files:413 - split: train414 path: Ara--saudinewsnet/train-*415- config_name: Ary--AbderrahmanSkiredj1--Darija-Wikipedia416 data_files:417 - split: train418 path: Ary--AbderrahmanSkiredj1--Darija-Wikipedia/train-*419- config_name: Ary--Ali-C137--Darija-Stories-Dataset420 data_files:421 - split: train422 path: Ary--Ali-C137--Darija-Stories-Dataset/train-*423- config_name: Ary--Wikipedia424 data_files:425 - split: train426 path: Ary--Wikipedia/train-*427- config_name: Arz--Wikipedia428 data_files:429 - split: train430 path: Arz--Wikipedia/train-*431---432# Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"433 434**The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts**435 436## Dataset Description437 438The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we are trying to centralize these dispersed resources into a single, comprehensive repository.439 440Encompassing a wide spectrum of content, ranging from social media conversations to literary masterpieces, MAD captures the rich tapestry of Arabic communication, including both standard Arabic and regional dialects.441 442This corpus offers comprehensive insights into the linguistic diversity and cultural nuances of Arabic expression.443 444## Usage 445 446If you want to use this dataset you pick one among the available configs:447 448`Ara--MBZUAI--Bactrian-X` | `Ara--OpenAssistant--oasst1` | `Ary--AbderrahmanSkiredj1--Darija-Wikipedia`449 450`Ara--Wikipedia` | `Ary--Wikipedia` | `Arz--Wikipedia`451 452`Ary--Ali-C137--Darija-Stories-Dataset` | `Ara--Ali-C137--Hindawi-Books-dataset` | ``453 454Example of usage:455 456```python457dataset = load_dataset('M-A-D/Mixed-Arabic-Datasets-Repo', 'Ara--MBZUAI--Bactrian-X')458```459 460If you loaded multiple datasets and wanted to merge them together then you can simply laverage `concatenate_datasets()` from `datasets`461 462```pyhton463dataset3 = concatenate_datasets([dataset1['train'], dataset2['train']])464```465 466Note : proccess the datasets before merging in order to make sure you have a new dataset that is consistent467 468## Dataset Size469 470The Mixed Arabic Datasets (MAD) is a dynamic and evolving collection, with its size fluctuating as new datasets are added or removed. As MAD continuously expands, it becomes a living resource that adapts to the ever-changing landscape of Arabic language datasets.471 472**Dataset List**473 474MAD draws from a diverse array of sources, each contributing to its richness and breadth. While the collection is constantly evolving, some of the datasets that are poised to join MAD in the near future include:475 476- [✔] OpenAssistant/oasst1 (ar portion) : [Dataset Link](https://huggingface.co/datasets/OpenAssistant/oasst1)477- [✔] MBZUAI/Bactrian-X (ar portion) : [Dataset Link](https://huggingface.co/datasets/MBZUAI/Bactrian-X/viewer/ar/train)478- [✔] AbderrahmanSkiredj1/Darija-Wikipedia : [Dataset Link](https://huggingface.co/datasets/AbderrahmanSkiredj1/moroccan_darija_wikipedia_dataset)479- [✔] Arabic Wikipedia : [Dataset Link](https://huggingface.co/datasets/wikipedia)480- [✔] Moroccan Arabic Wikipedia : [Dataset Link](https://huggingface.co/datasets/wikipedia)481- [✔] Egyptian Arabic Wikipedia : [Dataset Link](https://huggingface.co/datasets/wikipedia)482- [✔] Darija Stories Dataset : [Dataset Link](https://huggingface.co/datasets/Ali-C137/Darija-Stories-Dataset)483- [✔] Hindawi Books Dataset : [Dataset Link](https://huggingface.co/datasets/Ali-C137/Hindawi-Books-dataset)484- [] uonlp/CulturaX - ar : [Dataset Link](https://huggingface.co/datasets/uonlp/CulturaX/viewer/ar/train)485- [✔] Pain/ArabicTweets : [Dataset Link](https://huggingface.co/datasets/pain/Arabic-Tweets)486- [] Abu-El-Khair Corpus : [Dataset Link](https://huggingface.co/datasets/arabic_billion_words)487- [✔] QuranExe : [Dataset Link](https://huggingface.co/datasets/mustapha/QuranExe)488- [✔] MNAD : [Dataset Link](https://huggingface.co/datasets/J-Mourad/MNAD.v1)489- [✔] IADD : [Dataset Link](https://raw.githubusercontent.com/JihadZa/IADD/main/IADD.json)490- [] OSIAN : [Dataset Link](https://wortschatz.uni-leipzig.de/en/download/Arabic#ara-tn_newscrawl-OSIAN_2018)491- [✔] MAC corpus : [Dataset Link](https://raw.githubusercontent.com/LeMGarouani/MAC/main/MAC%20corpus.csv)492- [✔] Goud.ma-Sum : [Dataset Link](https://huggingface.co/datasets/Goud/Goud-sum)493- [✔] SaudiNewsNet : [Dataset Link](https://huggingface.co/datasets/saudinewsnet)494- [✔] Miracl : [Dataset Link](https://huggingface.co/datasets/miracl/miracl)495- [✔] CardiffNLP/TweetSentimentMulti : [Dataset Link](https://huggingface.co/datasets/cardiffnlp/tweet_sentiment_multilingual)496- [] OSCAR-2301 : [Dataset Link](https://huggingface.co/datasets/oscar-corpus/OSCAR-2301/viewer/ar/train)497- [] mc4 : [Dataset Link](https://huggingface.co/datasets/mc4/viewer/ar/train)498- [✔] bigscience/xP3 : [Dataset Link](https://huggingface.co/datasets/bigscience/xP3/viewer/ar/train)499- [] Muennighoff/xP3x : [Dataset Link](https://huggingface.co/datasets/Muennighoff/xP3x)500- [] Ai_Society : [Dataset Link](https://huggingface.co/datasets/camel-ai/ai_society_translated)501 502 503## Potential Use Cases504 505The Mixed Arabic Datasets (MAD) holds the potential to catalyze a multitude of groundbreaking applications:506 507- **Linguistic Analysis:** Employ MAD to conduct in-depth linguistic studies, exploring dialectal variances, language evolution, and grammatical structures.508- **Topic Modeling:** Dive into diverse themes and subjects through the extensive collection, revealing insights into emerging trends and prevalent topics.509- **Sentiment Understanding:** Decode sentiments spanning Arabic dialects, revealing cultural nuances and emotional dynamics.510- **Sociocultural Research:** Embark on a sociolinguistic journey, unraveling the intricate connection between language, culture, and societal shifts.511 512## Dataset Access513 514MAD's access mechanism is unique: while it doesn't carry a general license itself, each constituent dataset within the corpus retains its individual license. By accessing the dataset details through the provided links in the "Dataset List" section above, users can understand the specific licensing terms for each dataset.515 516### Join Us on Discord517 518For discussions, contributions, and community interactions, join us on Discord! [](https://discord.gg/2NpJ9JGm)519 520### How to Contribute521 522Want to contribute to the Mixed Arabic Datasets project? Follow our comprehensive guide on Google Colab for step-by-step instructions: [Contribution Guide](https://colab.research.google.com/drive/1kOIRoicgCOV8TPvASAI_2uMY7rpXnqzJ?usp=sharing).523 524**Note**: If you'd like to test a contribution before submitting it, feel free to do so on the [MAD Test Dataset](https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Dataset-test).525 526## Citation527 528```529@dataset{ 530title = {Mixed Arabic Datasets (MAD)},531author = {MAD Community},532howpublished = {Dataset},533url = {https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo},534year = {2023},535}536```