translation
mt5-small-parsinlu-opus-translation_fa_enmt5-base-parsinlu-opus-translation_fa_enenvit5-translationlarge-v2-aug-translationmeta-flores-translation-chinese-english-model-GGUFnayohan_-_llama3-8b-it-translation-sharegpt-en-ko-ggufnb-nn-translationTranslation-EnKo_-_gemma-2-2b-it-general1.2m-trc313eval45-gguf
translations-raw
natgillin/translations-raw
Frozen, canonical raw bitext consolidated from upstream alvations/mtdata-raw* snapshots (since deleted). This is the read-only source-of-truth for downstream quality-filtering pipelines.
31,663 parquet files (1566.8 GB)
49 language pairs under data/<src-tgt>/
Schema: 5 columns — see below
Read-only for downstream pipelines. Do not delete or modify.
Schema
Each parquet has 5 columns:
column
type
description
source
string… See the full description on the dataset page: https://huggingface.co/datasets/natgillin/translations-raw.Quranic-Translation-Audio-Data
Overview
Quranic Translation Audio Data is a highly curated, standardized, and streaming-optimized multilingual audio dataset containing the complete recitation of translation audios and commentaries of the Holy Quran across 51 different translation directories.
Every audio track has been meticulously converted from heavy .mp3 source files into the modern, high-fidelity Opus (.opus) format at a streaming-optimized bitrate of 32kbps. Alongside… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Translation-Audio-Data.mala-bilingual-translation-corpus
MaLA Corpus: Massive Language Adaptation Corpus
This MaLA-LM/mala-bilingual-translation-corpus is the MaLA bilingual translation corpus, collected and processed from various sources.
As a part of MaLA Corpus that aims to enhance massive language adaptation in many languages, it contains bilingual translation data (aka, parallel data and bitexts) in 2,500+ language pairs (500+ languages).
Key statistics of all language pairs available at… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-bilingual-translation-corpus.Weblate-Translations
Dataset Card for Weblate Translations
A dataset containing strings from projects hosted on Weblate and their translations into other languages.
Please consider donating or contributing to Weblate if you find this dataset useful.
To avoid rows with values like "None" and "N/A" being interpreted as missing values, pass the keep_default_na parameter like this:
from datasets import load_dataset
dataset = load_dataset("ayymen/Weblate-Translations", keep_default_na=False)… See the full description on the dataset page: https://huggingface.co/datasets/ayymen/Weblate-Translations.NLG-Machine-Translation
SEA Machine Translation
SEA Machine Translation evaluates a model's ability to translate a document from a source language into a target language coherently and fluently. It is sampled from FLORES 200 for Burmese, Chinese, English, Indonesian, Khmer, Malay, Tamil, Thai, and Vietnamese, and NusaX for Indonesian, Javanese, and Sundanese.
Supported Tasks and Leaderboards
SEA Machine Translation is designed for evaluating chat or instruction-tuned large language models… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Machine-Translation.translation
