MaLA-LM/mala-bilingual-translation-corpus
MaLA Corpus: Massive Language Adaptation Corpus This MaLA-LM/mala-bilingual-translation-corpus is the MaLA bilingual translation corpus, collected and processed from various sources. As a part of MaLA Corpus that aims to enhance massive language adaptation in many languages, it contains bilingual translation data (aka, parallel data and bitexts) in 2,500+ language pairs (500+ languages). Key statistics of all language pairs available at… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-bilingual-translation-corpus.
MaLA Corpus: Massive Language Adaptation Corpus
This **MaLA-LM/mala-bilingual-translation-corpus** is the MaLA bilingual translation corpus, collected and processed from various sources. As a part of **MaLA Corpus** that aims to enhance massive language adaptation in many languages, it contains bilingual translation data (aka, parallel data and bitexts) in 2,500+ language pairs (500+ languages).
Key statistics of all language pairs available at https://github.com/MaLA-LM/LangResourceAtlas/tree/main/mala-parallel
The **MaLA Corpus** (Massive Language Adaptation) is a series of comprehensive, multilingual datasets designed to support the continual pre-training of large language models. This **MaLA-LM/mala-bilingual-translation-corpus** set can also support the training of multilingual translation models.
Key Features
- Language Coverage: Includes data in 2,500+ language pairs.
- Pre-processing: The corpus is cleaned and deduplicated to ensure high-quality training data.
Dataset Creation
This **MaLA-LM/mala-bilingual-translation-corpus** set was created by processing data from various sources, followed by rigorous pre-processing to ensure the quality of the data:
- Cleaning: Noisy and irrelevant data was removed to ensure higher data quality.
- Deduplication: Duplicate entries across multiple sources were eliminated.
- Normalization: The data was normalized, and language codes were standardized to ISO 639-3 to ensure consistency across all sources.
Intended Use
This **MaLA-LM/mala-bilingual-translation-corpus** set is intended for researchers and developers looking to improve the multilingual capabilities of language models. It is especially useful for:
- Continual Pre-training of large language models to enhance the performance in low-resource languages.
- Fine-tuning models on multilingual benchmarks to improve language coverage across a variety of domains.
- Multilingual tasks such as machine translation.
Take-down Policy
We don't own any part of the data. We will comply with legitimate requests by removing the affected sources from the corpora.
Citation
This **MaLA-LM/mala-bilingual-translation-corpus** set was processed by the MaLA-LM project and used to train 🤗MaLA-LM/emma-500-llama3.1-8b-bi and 🤗MaLA-LM/emma-500-llama3-8b-bi. If you find this dataset useful, please cite our paper below.
@inproceedings{ji-etal-2026-data,
title = {Data-Centric Continual Pre-training for 500+ Languages: A New Bilingual Translation Corpus and Multilingual Models},
author = {Ji, Shaoxiong and
Li, Zihao and
Paavola, Jaakko and
Luo, Hengyu and
Tiedemann, J{\"o}rg},
editor = {Liakata, Maria and
Moreira, Viviane P. and
Zhang, Jiajun and
Jurgens, David},
booktitle = {Findings of the {A}ssociation for {C}omputational {L}inguistics: {ACL} 2026},
month = jul,
year = {2026},
address = {San Diego, California, United States},
publisher = {Association for Computational Linguistics},
url = {https://aclanthology.org/2026.findings-acl.937/},
doi = {10.18653/v1/2026.findings-acl.937},
pages = {18776--18807},
isbn = {979-8-89176-395-1}
}
@article{ji2024emma500enhancingmassivelymultilingual,
title={{EMMA}-500: Enhancing Massively Multilingual Adaptation of Large Language Models},
author={Shaoxiong Ji and Zihao Li and Indraneil Paul and Jaakko Paavola and Peiqin Lin and Pinzhen Chen and Dayyán O'Brien and Hengyu Luo and Hinrich Schütze and Jörg Tiedemann and Barry Haddow},
year={2024},
journal={arXiv preprint 2409.17892},
url={https://arxiv.org/abs/2409.17892},
}