CoolFace
Datasetpublic

somosnlp-hackathon-2022/Axolotl-Spanish-Nahuatl

Axolotl-Spanish-Nahuatl : Parallel corpus for Spanish-Nahuatl machine translation Dataset Collection In order to get a good translator, we collected and cleaned two of the most complete Nahuatl-Spanish parallel corpora available. Those are Axolotl collected by an expert team at UNAM and Bible UEDIN Nahuatl Spanish crawled by Christos Christodoulopoulos and Mark Steedman from Bible Gateway site. After this, we ended with 12,207 samples from Axolotl due to… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/Axolotl-Spanish-Nahuatl.

sourceHugging Facempl-2.0updated 3y agoView on Hugging Face
16likes88downloads
Dataset Card

Axolotl-Spanish-Nahuatl : Parallel corpus for Spanish-Nahuatl machine translation

Table of Contents

  • —[Dataset Card for [Axolotl-Spanish-Nahuatl]](#dataset-card-for-Axolotl-Spanish-Nahuatl)

Dataset Description

  • —Source 1: http://www.corpus.unam.mx/axolotl
  • —Source 2: http://link.springer.com/article/10.1007/s10579-014-9287-y
  • —Repository:1 https://github.com/ElotlMX/py-elotl
  • —Repository:2 https://github.com/christos-c/bible-corpus/blob/master/bibles/Nahuatl-NT.xml
  • —Paper: https://aclanthology.org/N15-2021.pdf

Dataset Collection

In order to get a good translator, we collected and cleaned two of the most complete Nahuatl-Spanish parallel corpora available. Those are Axolotl collected by an expert team at UNAM and Bible UEDIN Nahuatl Spanish crawled by Christos Christodoulopoulos and Mark Steedman from Bible Gateway site.

After this, we ended with 12,207 samples from Axolotl due to misalignments and duplicated texts in Spanish in both original and nahuatl columns and 7,821 samples from Bible UEDIN for a total of 20028 utterances.

Team members

Applications