CoolFace
Datasetpublic

CohereLabs/aya_collection_language_split

This is a re-upload of the aya_collection, and only differs in the structure of upload. While the original aya_collection is structured by folders split according to dataset name, this dataset is split by language. We recommend you use this version of the dataset if you are only interested in downloading all of the Aya collection for a single or smaller set of languages. Dataset Summary The Aya Collection is a massive multilingual collection consisting of 513 million instances… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_collection_language_split.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
122likes22kdownloads
Dataset Card

Aya Header

**This is a re-upload of the [aya_collection](https://huggingface.co/datasets/CohereLabs/aya_collection), and only differs in the structure of upload. While the original [aya_collection](https://huggingface.co/datasets/CohereLabs/aya_collection) is structured by folders split according to dataset name, this dataset is split by language. We recommend you use this version of the dataset if you are only interested in downloading all of the Aya collection for a single or smaller set of languages.**

Dataset Summary

The Aya Collection is a massive multilingual collection consisting of 513 million instances of prompts and completions covering a wide range of tasks. This collection incorporates instruction-style templates from fluent speakers and applies them to a curated list of datasets, as well as translations of instruction-style datasets into 101 languages. Aya Dataset, a human-curated multilingual instruction and response dataset, is also part of this collection. See our paper for more details regarding the collection.

  • Aya Datasets Family: | Name | Explanation | |------|--------------| | aya_dataset | Human-annotated multilingual instruction finetuning dataset, comprising over 204K instances across 65 languages. | | aya_collection | Created by applying instruction-style templates from fluent speakers to 44 datasets, including translations of 19 instruction-style datasets into 101 languages. This collection structured based on dataset level subsets. An alternative version of the collection structured by language subsets is also available.| | aya_collection_language_split | Aya Collection structured based on language level subsets. | | aya_evaluation_suite | A diverse evaluation set for multilingual open-ended generation, featuring 250 culturally grounded prompts in 7 languages, 200 translated prompts in 24 languages, and human-edited versions selected for cross-cultural relevance from English Dolly in 6 languages.| | aya_redteaming| A red-teaming dataset consisting of harmful prompts in 8 languages across 9 different categories of harm with explicit labels for "global" and "local" harm.|

Dataset

The Aya Collection is a comprehensive, large corpus of datasets that can be used by researchers around the world to train multilingual models. Our goal is only to include datasets with permissive licensing for manipulation and redistribution.

The Aya Collection consists of three different sources of data:

  1. 1.Templated data: We collaborated with fluent speakers to create templates that allowed for the automatic expansion of existing datasets into various languages.
  2. 2.Translated data: We translated a hand-selected subset of 19 datasets into 101 languages (114 dialects) using the NLLB 3.3B parameter machine translation model.
  3. 3.Aya Dataset: We release the Aya Dataset as a subset of the overall collection. This is the only dataset in the collection that is human-annotated in its entirety.

Load with Datasets

To load this dataset with Datasets, you'll need to install Datasets as pip install datasets --upgrade and then use the following code:

python
from datasets import load_dataset

dataset = load_dataset("CohereLabs/aya_collection_language_split", "english")

In the above code snippet, "english" refers to a subset of the aya_collection. You can load other subsets by specifying its name at the time of loading the dataset.

Data Instances

An example of a train instance looks as follows:

json
{'id': 246001,
 'inputs': 'The following query in English is taken from the geography category. What could be the answer to the question?\nWhat is the seventh tallest mountain in North America?',
 'targets': 'The answer is Mount Lucania.',
 'dataset_name': 'Mintaka-inst',
 'sub_dataset_name': '-',
 'task_type': 'question-answering',
 'template_id': 3,
 'language': 'eng',
 'split': 'train',
 'script': 'Latn'
}

Data Fields

The data fields are the same among all splits:

  • id: Unique id of the data point
  • inputs: Prompt or input to the language model.
  • targets: Completion or output of the language model.
  • dataset_name: The name of the source dataset that the data point was taken from
  • sub_dataset_name: If the source is a collection, this field indicates which part of that collection the data point was taken from. If it is not a collection, this field is left blank.
  • task_type: The task type that this conversation belongs to.
  • template_id: The id of the template applied to this data point.
  • language: The ISO code of the dialect of the conversation.
  • script: The script of the language.
  • split: Indicates whether the data point is part of the train or the test split.

Statistics

The total number of data points, including the Aya Dataset` is 513,758,189. To view the breakdown of dialect codes and the respective templated and translated data point counts in the Aya Collection , refer to the toggled table below.

<details> <summary> <b> Breakdown of Aya Collection data point counts grouped by dialects </b> </summary>

dialect codelanguagetotal count
aceAchinese8242684
acmArabic4120342
acqArabic4120342
aebArabic4120342
afrAfrikaans4126450
ajpArabic4120342
alsAlbanian4120342
amhAmharic4145669
apcArabic4120342
arbArabic6641429
arsArabic4120342
aryArabic4138418
arzArabic4120342
azbAzerbaijani4120342
azjAzerbaijani4120342
belBelarusian4141615
benBengali4151003
bjnBanjar8242684
bulBulgarian4158064
catCatalan4187242
cebCebuano4120342
cesCzech4299946
ckbKurdish4120342
cymWelsh4120342
danDanish4156652
deuGerman5447064
ellGreek4160633
engEnglish17838105
epoEsperanto4120342
estEstonian4120342
eusBasque4120342
finFinnish4578237
fraFrench4955862
glaScottish Gaelic4120342
gleIrish4120342
glgGalician4120342
gujGujarati4122499
hatHaitian Creole4120342
hauHausa4171738
hebHebrew4223808
hinHindi4380729
hunHungarian4202381
hyeArmenian4127422
iboIgbo4156654
indIndonesian4166051
islIcelandic4120342
itaItalian4526024
javJavanese4121171
jpnJapanese6813519
kanKannada4121498
kasKashmiri4120342
katGeorgian4120342
kazKazakh4120342
khkMongolian4120342
khmKhmer4120342
kirKyrgyz4120342
kmrKurdish4120342
kncKanuri8240684
korKorean4161353
laoLao4120342
litLithuanian4120342
ltzLuxembourgish4120342
lvsLatvian4120342
malMalayalam4124689
marMarathi4124020
minMinangkabau6755788
mkdMacedonian4120342
mltMaltese4120342
mniManipuri4120342
mriMaori4120342
myaBurmese4120342
nldDutch4340523
nnoNorwegian4120342
nobNorwegian4120342
npiNepali4120342
nsoNorthern Sotho4120342
pbtPashto4120342
pesPersian4365862
pltMalagasy4120342
polPolish4452845
porPortuguese4407774
ronRomanian4156701
rusRussian4666262
sinSinhala4120537
slkSlovak4148187
slvSlovenian4146073
smoSamoan4120342
snaShona4124026
sndSindhi4120342
somSomali4123268
sotSouthern Sotho4120342
spaSpanish4499536
srpSerbian4197466
sunSundanese4122550
sweSwedish4196828
swhSwahili4133068
tamTamil4131804
taqTamasheq4120342
telTelugu4598163
tgkTajik4120342
thaThai6245522
turTurkish4180274
ukrUkrainian4309726
urdUrdu4458081
uznUzbek4120342
vieVietnamese4162574
xhoXhosa4123294
yddYiddish4120342
yorYoruba4125249
yueChinese4120342
zho-HansChinese4174870
zho-HantChinese4120342
zsmMalay4134292
zulZulu4121128
arqArabic6046
banBalinese2000
bbcToba Batak2000
bemBemba776
filFilipino220
fonFon845
hrvCroatian9007
kinKinyarwanda11165
lijLigurian6409
madMadurese2000
nijNgaju2000
norNorwegian72352
panPunjabi2156
twiTwi10840
wolWolof785
zhoChinese74972

PS: Templated data also includes Mozambican Portuguese, which doesn't have its own ISO language code.

</details>

<br>

Motivations & Intentions

  • Curation Rationale: Automatic augmentation of existing datasets serves to enhance the available linguistic resources for multiple languages. The list of languages was initially established from mT5 and aligned with the annotators’ language list and NLLB translation model. The datasets were translated directly from English for all languages.

Additional Information

Provenance

  • Methods Used: A combination of crowd-sourced templating and automatic translation was employed to source this dataset.
  • Methodology Details:
  • Source: Existing NLP datasets
  • Dates of Collection: May 2023 - Dec 2023

Dataset Version and Maintenance

  • Maintenance Status: Actively Maintained
  • Version Details:
  • Current version: 1.0
  • Last Update: 02/2024
  • First Release: 02/2024

Authorship

  • Publishing Organization: Cohere Labs
  • Industry Type: Not-for-profit - Tech
  • Contact Details: https://cohere.com/research/aya

Licensing Information

This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.

Citation Information

bibtex
@misc{singh2024aya,
      title={Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning}, 
      author={Shivalika Singh and Freddie Vargus and Daniel Dsouza and Börje F. Karlsson and Abinaya Mahendiran and Wei-Yin Ko and Herumb Shandilya and Jay Patel and Deividas Mataciunas and Laura OMahony and Mike Zhang and Ramith Hettiarachchi and Joseph Wilson and Marina Machado and Luisa Souza Moura and Dominik Krzemiński and Hakimeh Fadaei and Irem Ergün and Ifeoma Okoh and Aisha Alaagib and Oshan Mudannayake and Zaid Alyafeai and Vu Minh Chien and Sebastian Ruder and Surya Guthikonda and Emad A. Alghamdi and Sebastian Gehrmann and Niklas Muennighoff and Max Bartolo and Julia Kreutzer and Ahmet Üstün and Marzieh Fadaee and Sara Hooker},
      year={2024},
      eprint={2402.06619},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}