datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scb-mt-en-th-2020_mt-opus
Dataset Card for "scb-mt-en-th-2020_mt-opus"
More Information needed
English-Thai scb-mt-en-th-2020 v1.0 and datasets listed in Open Parallel Corpus (OPUS)
This dataset come from A large English–Thai parallel corpus from the web and machine-generated text that released at GitHub.
OPUS-MT-EN-Fixed
OPUS-100-Fixed: Tokenisation-Improved English-Maltese Dataset
Overview
OPUS-100-Fixed is an updated version of the OPUS-100 parallel English-Maltese dataset.
This version addresses tokenisation inconsistencies in the Maltese text using the MLRS tokeniser, aiming to improve machine translation quality.
The "en" column is the same as in the original OPUS-100 data, while the "mt" column has been corrected with the MLRS detokeniser.
Citation
If you use this… See the full description on the dataset page: https://huggingface.co/datasets/MLRS/OPUS-MT-EN-Fixed.opus-mt-ct2opus-mt-arabic-benchmark-2026-03-28
OPUS-MT Arabic-English Translation Benchmark
Experiment Details
Date: 2026-03-28
Models Tested:
Helsinki-NLP/opus-mt-en-ar (English → Arabic)
Helsinki-NLP/opus-mt-ar-en (Arabic → English)
Total Tests: 9
Domain: NLP / Translation
Summary
Metric
Value
MSA Accuracy Rate
100%
Dialectal Accuracy Rate
0%
Avg Latency (MSA)
5.67s
Avg Latency (Dialectal)
0.5s
Key Finding
OPUS-MT handles Modern Standard Arabic (MSA) well but truncates… See the full description on the dataset page: https://huggingface.co/datasets/O96a/opus-mt-arabic-benchmark-2026-03-28.MT-wmt14-500k-opus-mt-en-de
src (source text): WMT14 English-German dataset (train split, 500k sentences)
ref (reference text): original reference
h1,h2,h3,h4,h5: list of 5 translation candidates, each with log-probability score (sorted in descending order)
MT model: Helsinki-NLP/opus-mt-en-de
Translation direction: English → German
Beam search: num_beams=5, num_return_sequences=5)
opus-mt-en-bkm-37opus-mt-en-bkm-60
