CoolFace
Datasetpublic

yrrhall/Mixed-Arabic-Datasets-Repo

Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Mixed-Arabic-Datasets-Repo.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes340downloads
README.md536 linesDownload Raw Back to root
1---2language:3- ar4size_categories:5- 1B<n<10B6task_categories:7- text-classification8- question-answering9- translation10- summarization11- conversational12- text-generation13- text2text-generation14- fill-mask15pretty_name: Mixed Arabic Datasets (MAD) Corpus16dataset_info:17- config_name: Ara--Ali-C137--Hindawi-Books-dataset18  features:19  - name: BookLink20    dtype: string21  - name: BookName22    dtype: string23  - name: AuthorName24    dtype: string25  - name: AboutBook26    dtype: string27  - name: ChapterLink28    dtype: string29  - name: ChapterName30    dtype: string31  - name: ChapterText32    dtype: string33  - name: AboutAuthor34    dtype: string35  splits:36  - name: train37    num_bytes: 136485425938    num_examples: 4982139  download_size: 49467800240  dataset_size: 136485425941- config_name: Ara--Goud--Goud-sum42  features:43  - name: article44    dtype: string45  - name: headline46    dtype: string47  - name: categories48    dtype: string49  splits:50  - name: train51    num_bytes: 28829654452    num_examples: 13928853  download_size: 14773577654  dataset_size: 28829654455- config_name: Ara--J-Mourad--MNAD.v156  features:57  - name: Title58    dtype: string59  - name: Body60    dtype: string61  - name: Category62    dtype: string63  splits:64  - name: train65    num_bytes: 110192198066    num_examples: 41856367  download_size: 52715412268  dataset_size: 110192198069- config_name: Ara--JihadZa--IADD70  features:71  - name: Sentence72    dtype: string73  - name: Region74    dtype: string75  - name: DataSource76    dtype: string77  - name: Country78    dtype: string79  splits:80  - name: train81    num_bytes: 1916707082    num_examples: 13580483  download_size: 864449184  dataset_size: 1916707085- config_name: Ara--LeMGarouani--MAC-corpus86  features:87  - name: tweets88    dtype: string89  - name: type90    dtype: string91  - name: class92    dtype: string93  splits:94  - name: train95    num_bytes: 194564696    num_examples: 1808797  download_size: 86619898  dataset_size: 194564699- config_name: Ara--MBZUAI--Bactrian-X100  features:101  - name: instruction102    dtype: string103  - name: input104    dtype: string105  - name: id106    dtype: string107  - name: output108    dtype: string109  splits:110  - name: train111    num_bytes: 66093524112    num_examples: 67017113  download_size: 33063779114  dataset_size: 66093524115- config_name: Ara--OpenAssistant--oasst1116  features:117  - name: message_id118    dtype: string119  - name: parent_id120    dtype: string121  - name: user_id122    dtype: string123  - name: created_date124    dtype: string125  - name: text126    dtype: string127  - name: role128    dtype: string129  - name: lang130    dtype: string131  - name: review_count132    dtype: int32133  - name: review_result134    dtype: bool135  - name: deleted136    dtype: bool137  - name: rank138    dtype: float64139  - name: synthetic140    dtype: bool141  - name: model_name142    dtype: 'null'143  - name: detoxify144    dtype: 'null'145  - name: message_tree_id146    dtype: string147  - name: tree_state148    dtype: string149  - name: emojis150    struct:151    - name: count152      sequence: int32153    - name: name154      sequence: string155  - name: labels156    struct:157    - name: count158      sequence: int32159    - name: name160      sequence: string161    - name: value162      sequence: float64163  - name: __index_level_0__164    dtype: int64165  splits:166  - name: train167    num_bytes: 58168168    num_examples: 56169  download_size: 30984170  dataset_size: 58168171- config_name: Ara--Wikipedia172  features:173  - name: id174    dtype: string175  - name: url176    dtype: string177  - name: title178    dtype: string179  - name: text180    dtype: string181  splits:182  - name: train183    num_bytes: 3052201469184    num_examples: 1205403185  download_size: 1316212231186  dataset_size: 3052201469187- config_name: Ara--bigscience--xP3188  features:189  - name: inputs190    dtype: string191  - name: targets192    dtype: string193  splits:194  - name: train195    num_bytes: 4727881680196    num_examples: 2148955197  download_size: 2805060725198  dataset_size: 4727881680199- config_name: Ara--cardiffnlp--tweet_sentiment_multilingual200  features:201  - name: text202    dtype: string203  - name: label204    dtype:205      class_label:206        names:207          '0': negative208          '1': neutral209          '2': positive210  splits:211  - name: train212    num_bytes: 306108213    num_examples: 1839214  - name: validation215    num_bytes: 53276216    num_examples: 324217  - name: test218    num_bytes: 141536219    num_examples: 870220  download_size: 279900221  dataset_size: 500920222- config_name: Ara--miracl--miracl223  features:224  - name: query_id225    dtype: string226  - name: query227    dtype: string228  - name: positive_passages229    list:230    - name: docid231      dtype: string232    - name: text233      dtype: string234    - name: title235      dtype: string236  - name: negative_passages237    list:238    - name: docid239      dtype: string240    - name: text241      dtype: string242    - name: title243      dtype: string244  splits:245  - name: train246    num_bytes: 32012083247    num_examples: 3495248  download_size: 15798509249  dataset_size: 32012083250- config_name: Ara--mustapha--QuranExe251  features:252  - name: text253    dtype: string254  - name: resource_name255    dtype: string256  - name: verses_keys257    dtype: string258  splits:259  - name: train260    num_bytes: 133108687261    num_examples: 49888262  download_size: 58769417263  dataset_size: 133108687264- config_name: Ara--pain--Arabic-Tweets265  features:266  - name: text267    dtype: string268  splits:269  - name: train270    num_bytes: 41639770853271    num_examples: 202700438272  download_size: 22561651700273  dataset_size: 41639770853274- config_name: Ara--saudinewsnet275  features:276  - name: source277    dtype: string278  - name: url279    dtype: string280  - name: date_extracted281    dtype: string282  - name: title283    dtype: string284  - name: author285    dtype: string286  - name: content287    dtype: string288  splits:289  - name: train290    num_bytes: 103654009291    num_examples: 31030292  download_size: 49117164293  dataset_size: 103654009294- config_name: Ary--AbderrahmanSkiredj1--Darija-Wikipedia295  features:296  - name: text297    dtype: string298  splits:299  - name: train300    num_bytes: 8104410301    num_examples: 4862302  download_size: 3229966303  dataset_size: 8104410304- config_name: Ary--Ali-C137--Darija-Stories-Dataset305  features:306  - name: ChapterName307    dtype: string308  - name: ChapterLink309    dtype: string310  - name: Author311    dtype: string312  - name: Text313    dtype: string314  - name: Tags315    dtype: int64316  splits:317  - name: train318    num_bytes: 476926644319    num_examples: 6142320  download_size: 241528641321  dataset_size: 476926644322- config_name: Ary--Wikipedia323  features:324  - name: id325    dtype: string326  - name: url327    dtype: string328  - name: title329    dtype: string330  - name: text331    dtype: string332  splits:333  - name: train334    num_bytes: 10007364335    num_examples: 6703336  download_size: 4094377337  dataset_size: 10007364338- config_name: Arz--Wikipedia339  features:340  - name: id341    dtype: string342  - name: url343    dtype: string344  - name: title345    dtype: string346  - name: text347    dtype: string348  splits:349  - name: train350    num_bytes: 1364641408351    num_examples: 1617770352  download_size: 306420318353  dataset_size: 1364641408354configs:355- config_name: Ara--Ali-C137--Hindawi-Books-dataset356  data_files:357  - split: train358    path: Ara--Ali-C137--Hindawi-Books-dataset/train-*359- config_name: Ara--Goud--Goud-sum360  data_files:361  - split: train362    path: Ara--Goud--Goud-sum/train-*363- config_name: Ara--J-Mourad--MNAD.v1364  data_files:365  - split: train366    path: Ara--J-Mourad--MNAD.v1/train-*367- config_name: Ara--JihadZa--IADD368  data_files:369  - split: train370    path: Ara--JihadZa--IADD/train-*371- config_name: Ara--LeMGarouani--MAC-corpus372  data_files:373  - split: train374    path: Ara--LeMGarouani--MAC-corpus/train-*375- config_name: Ara--MBZUAI--Bactrian-X376  data_files:377  - split: train378    path: Ara--MBZUAI--Bactrian-X/train-*379- config_name: Ara--OpenAssistant--oasst1380  data_files:381  - split: train382    path: Ara--OpenAssistant--oasst1/train-*383- config_name: Ara--Wikipedia384  data_files:385  - split: train386    path: Ara--Wikipedia/train-*387- config_name: Ara--bigscience--xP3388  data_files:389  - split: train390    path: Ara--bigscience--xP3/train-*391- config_name: Ara--cardiffnlp--tweet_sentiment_multilingual392  data_files:393  - split: train394    path: Ara--cardiffnlp--tweet_sentiment_multilingual/train-*395  - split: validation396    path: Ara--cardiffnlp--tweet_sentiment_multilingual/validation-*397  - split: test398    path: Ara--cardiffnlp--tweet_sentiment_multilingual/test-*399- config_name: Ara--miracl--miracl400  data_files:401  - split: train402    path: Ara--miracl--miracl/train-*403- config_name: Ara--mustapha--QuranExe404  data_files:405  - split: train406    path: Ara--mustapha--QuranExe/train-*407- config_name: Ara--pain--Arabic-Tweets408  data_files:409  - split: train410    path: Ara--pain--Arabic-Tweets/train-*411- config_name: Ara--saudinewsnet412  data_files:413  - split: train414    path: Ara--saudinewsnet/train-*415- config_name: Ary--AbderrahmanSkiredj1--Darija-Wikipedia416  data_files:417  - split: train418    path: Ary--AbderrahmanSkiredj1--Darija-Wikipedia/train-*419- config_name: Ary--Ali-C137--Darija-Stories-Dataset420  data_files:421  - split: train422    path: Ary--Ali-C137--Darija-Stories-Dataset/train-*423- config_name: Ary--Wikipedia424  data_files:425  - split: train426    path: Ary--Wikipedia/train-*427- config_name: Arz--Wikipedia428  data_files:429  - split: train430    path: Arz--Wikipedia/train-*431---432# Dataset Card for "Mixed Arabic Datasets (MAD) Corpus"433 434**The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts**435 436## Dataset Description437 438The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we are trying to centralize these dispersed resources into a single, comprehensive repository.439 440Encompassing a wide spectrum of content, ranging from social media conversations to literary masterpieces, MAD captures the rich tapestry of Arabic communication, including both standard Arabic and regional dialects.441 442This corpus offers comprehensive insights into the linguistic diversity and cultural nuances of Arabic expression.443 444## Usage 445 446If you want to use this dataset you pick one among the available configs:447 448`Ara--MBZUAI--Bactrian-X` | `Ara--OpenAssistant--oasst1` | `Ary--AbderrahmanSkiredj1--Darija-Wikipedia`449 450`Ara--Wikipedia` | `Ary--Wikipedia` | `Arz--Wikipedia`451 452`Ary--Ali-C137--Darija-Stories-Dataset` | `Ara--Ali-C137--Hindawi-Books-dataset` | ``453 454Example of usage:455 456```python457dataset = load_dataset('M-A-D/Mixed-Arabic-Datasets-Repo', 'Ara--MBZUAI--Bactrian-X')458```459 460If you loaded multiple datasets and wanted to merge them together then you can simply laverage `concatenate_datasets()` from `datasets`461 462```pyhton463dataset3 = concatenate_datasets([dataset1['train'], dataset2['train']])464```465 466Note : proccess the datasets before merging in order to make sure you have a new dataset that is consistent467 468## Dataset Size469 470The Mixed Arabic Datasets (MAD) is a dynamic and evolving collection, with its size fluctuating as new datasets are added or removed. As MAD continuously expands, it becomes a living resource that adapts to the ever-changing landscape of Arabic language datasets.471 472**Dataset List**473 474MAD draws from a diverse array of sources, each contributing to its richness and breadth. While the collection is constantly evolving, some of the datasets that are poised to join MAD in the near future include:475 476- [✔] OpenAssistant/oasst1 (ar portion) : [Dataset Link](https://huggingface.co/datasets/OpenAssistant/oasst1)477- [✔] MBZUAI/Bactrian-X (ar portion) : [Dataset Link](https://huggingface.co/datasets/MBZUAI/Bactrian-X/viewer/ar/train)478- [✔] AbderrahmanSkiredj1/Darija-Wikipedia : [Dataset Link](https://huggingface.co/datasets/AbderrahmanSkiredj1/moroccan_darija_wikipedia_dataset)479- [✔] Arabic Wikipedia : [Dataset Link](https://huggingface.co/datasets/wikipedia)480- [✔] Moroccan Arabic Wikipedia : [Dataset Link](https://huggingface.co/datasets/wikipedia)481- [✔] Egyptian Arabic Wikipedia : [Dataset Link](https://huggingface.co/datasets/wikipedia)482- [✔] Darija Stories Dataset : [Dataset Link](https://huggingface.co/datasets/Ali-C137/Darija-Stories-Dataset)483- [✔] Hindawi Books Dataset : [Dataset Link](https://huggingface.co/datasets/Ali-C137/Hindawi-Books-dataset)484- [] uonlp/CulturaX - ar : [Dataset Link](https://huggingface.co/datasets/uonlp/CulturaX/viewer/ar/train)485- [✔] Pain/ArabicTweets : [Dataset Link](https://huggingface.co/datasets/pain/Arabic-Tweets)486- [] Abu-El-Khair Corpus : [Dataset Link](https://huggingface.co/datasets/arabic_billion_words)487- [✔] QuranExe : [Dataset Link](https://huggingface.co/datasets/mustapha/QuranExe)488- [✔] MNAD : [Dataset Link](https://huggingface.co/datasets/J-Mourad/MNAD.v1)489- [✔] IADD : [Dataset Link](https://raw.githubusercontent.com/JihadZa/IADD/main/IADD.json)490- [] OSIAN : [Dataset Link](https://wortschatz.uni-leipzig.de/en/download/Arabic#ara-tn_newscrawl-OSIAN_2018)491- [✔] MAC corpus : [Dataset Link](https://raw.githubusercontent.com/LeMGarouani/MAC/main/MAC%20corpus.csv)492- [✔] Goud.ma-Sum : [Dataset Link](https://huggingface.co/datasets/Goud/Goud-sum)493- [✔] SaudiNewsNet : [Dataset Link](https://huggingface.co/datasets/saudinewsnet)494- [✔] Miracl : [Dataset Link](https://huggingface.co/datasets/miracl/miracl)495- [✔] CardiffNLP/TweetSentimentMulti : [Dataset Link](https://huggingface.co/datasets/cardiffnlp/tweet_sentiment_multilingual)496- [] OSCAR-2301 : [Dataset Link](https://huggingface.co/datasets/oscar-corpus/OSCAR-2301/viewer/ar/train)497- [] mc4 : [Dataset Link](https://huggingface.co/datasets/mc4/viewer/ar/train)498- [✔] bigscience/xP3 : [Dataset Link](https://huggingface.co/datasets/bigscience/xP3/viewer/ar/train)499- [] Muennighoff/xP3x : [Dataset Link](https://huggingface.co/datasets/Muennighoff/xP3x)500- [] Ai_Society : [Dataset Link](https://huggingface.co/datasets/camel-ai/ai_society_translated)501 502 503## Potential Use Cases504 505The Mixed Arabic Datasets (MAD) holds the potential to catalyze a multitude of groundbreaking applications:506 507- **Linguistic Analysis:** Employ MAD to conduct in-depth linguistic studies, exploring dialectal variances, language evolution, and grammatical structures.508- **Topic Modeling:** Dive into diverse themes and subjects through the extensive collection, revealing insights into emerging trends and prevalent topics.509- **Sentiment Understanding:** Decode sentiments spanning Arabic dialects, revealing cultural nuances and emotional dynamics.510- **Sociocultural Research:** Embark on a sociolinguistic journey, unraveling the intricate connection between language, culture, and societal shifts.511 512## Dataset Access513 514MAD's access mechanism is unique: while it doesn't carry a general license itself, each constituent dataset within the corpus retains its individual license. By accessing the dataset details through the provided links in the "Dataset List" section above, users can understand the specific licensing terms for each dataset.515 516### Join Us on Discord517 518For discussions, contributions, and community interactions, join us on Discord! [![Discord](https://img.shields.io/discord/798499298231726101?label=Join%20us%20on%20Discord&logo=discord&logoColor=white&style=for-the-badge)](https://discord.gg/2NpJ9JGm)519 520### How to Contribute521 522Want to contribute to the Mixed Arabic Datasets project? Follow our comprehensive guide on Google Colab for step-by-step instructions: [Contribution Guide](https://colab.research.google.com/drive/1kOIRoicgCOV8TPvASAI_2uMY7rpXnqzJ?usp=sharing).523 524**Note**: If you'd like to test a contribution before submitting it, feel free to do so on the [MAD Test Dataset](https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Dataset-test).525 526## Citation527 528```529@dataset{ 530title = {Mixed Arabic Datasets (MAD)},531author = {MAD Community},532howpublished = {Dataset},533url = {https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo},534year = {2023},535}536```