language-translation
translation-infrastructure-language-access
When Translation Is Not Access
AI Translation Infrastructure and Communicative Inequality in Under-Resourced Languages
This repository supports a research project on a simple but important question:
When an AI system can translate a language, does that actually mean people can reliably understand and act on the information they receive?
The project focuses on Amharic and Afaan Oromo and examines AI translation in public and institutional communication, where errors can have… See the full description on the dataset page: https://huggingface.co/datasets/endalkchala/translation-infrastructure-language-access.thai-local-language-translation-dataset
Thai Local Language Translation Dataset
Thai Local Language Translation Dataset is a translation dataset for translate Thai Local Language to Thai Central Language. We create the dataset from Thai Dialect Corpus (Thai dialects ASR corpus). We select train set only from Thai Dialect Corpus.
The dataset support Khummuang, Korat, and Pattani.
Reference
Suwanbandit, A., Naowarat, B., Sangpetch, O., Chuangsuwanich, E. (2023) Thai Dialect Corpus and Transfer-based Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-local-language-translation-dataset.English-to-Afar-language-translation
Author
Created by Charif Ayfarah.
Contact: afbarit@gmail.com
License
Licensed under CC BY 4.0. You are free to use, modify, and distribute this dataset, including for commercial purposes, as long as you give appropriate credit.
multimodal_low-resource_language_translation
Dataset Card for Multimodal Low-Resource Language Translation Dataset
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This is the dataset for our paper "From Text to Multi-Modal: Advancing Low-Resource-Language Translation through Synthetic Data Generation and Cross-Modal Alignments" accepted by the workshop LoResMT 2025 of NAACL 2025… See the full description on the dataset page: https://huggingface.co/datasets/qianstats/multimodal_low-resource_language_translation.task1435_ro_sts_parallel_language_translation_ro_to_en
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1435_ro_sts_parallel_language_translation_ro_to_en
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1435_ro_sts_parallel_language_translation_ro_to_en.igbo_language_translation_dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/iamwille/igbo_language_translation_dataset.
