Llamacha/monolingual-quechua-iic
Dataset Card for Monolingual-Quechua-IIC Dataset Summary We present Monolingual-Quechua-IIC, a monolingual corpus of Southern Quechua, which can be used to build language models using Transformers models. This corpus also includes the Wiki and OSCAR corpora. We used this corpus to build Llama-RoBERTa-Quechua, the first language model for Southern Quechua using Transformers. Supported Tasks and Leaderboards More Information Needed… See the full description on the dataset page: https://huggingface.co/datasets/Llamacha/monolingual-quechua-iic.
Dataset Card for Monolingual-Quechua-IIC
Table of Contents
- Dataset Description
- Dataset Summary
- Supported Tasks and Leaderboards
- Languages
- Dataset Structure
- Data Instances
- Data Fields
- Data Splits
- Dataset Creation
- Curation Rationale
- Source Data
- Annotations
- Personal and Sensitive Information
- Considerations for Using the Data
- Social Impact of Dataset
- Discussion of Biases
- Other Known Limitations
- Additional Information
- Dataset Curators
- Licensing Information
- Citation Information
- Contributions
Dataset Description
- Homepage: https://llamacha.pe
- Paper: Introducing QuBERT: A Large Monolingual Corpus and BERT Model for Southern Quechua
- Point of Contact: Rodolfo Zevallos
- Size of downloaded dataset files: 373.28 MB
Dataset Summary
We present Monolingual-Quechua-IIC, a monolingual corpus of Southern Quechua, which can be used to build language models using Transformers models. This corpus also includes the Wiki and OSCAR corpora. We used this corpus to build Llama-RoBERTa-Quechua, the first language model for Southern Quechua using Transformers.
Supported Tasks and Leaderboards
Languages
Southern Quechua
Dataset Creation
Curation Rationale
Source Data
Initial Data Collection and Normalization
Who are the source language producers?
Annotations
Annotation process
Who are the annotators?
Personal and Sensitive Information
Considerations for Using the Data
Social Impact of Dataset
Discussion of Biases
Other Known Limitations
Additional Information
Dataset Curators
Licensing Information
Apache-2.0
Citation Information
@inproceedings{zevallos2022introducing,
title={Introducing QuBERT: A Large Monolingual Corpus and BERT Model for Southern Quechua},
author={Zevallos, Rodolfo and Ortega, John and Chen, William and Castro, Richard and Bel, Nuria and Toshio, Cesar and Venturas, Renzo and Aradiel, Hilario and Melgarejo, Nelsi},
booktitle={Proceedings of the Third Workshop on Deep Learning for Low-Resource Natural Language Processing},
pages={1--13},
year={2022}
}Contributions
Thanks to @rjzevallos for adding this dataset.
