datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scb-mt-en-th-2020_mt-opus
Dataset Card for "scb-mt-en-th-2020_mt-opus"
More Information needed
English-Thai scb-mt-en-th-2020 v1.0 and datasets listed in Open Parallel Corpus (OPUS)
This dataset come from A large English–Thai parallel corpus from the web and machine-generated text that released at GitHub.
OPUS-MT-EN-Fixed
OPUS-100-Fixed: Tokenisation-Improved English-Maltese Dataset
Overview
OPUS-100-Fixed is an updated version of the OPUS-100 parallel English-Maltese dataset.
This version addresses tokenisation inconsistencies in the Maltese text using the MLRS tokeniser, aiming to improve machine translation quality.
The "en" column is the same as in the original OPUS-100 data, while the "mt" column has been corrected with the MLRS detokeniser.
Citation
If you use this… See the full description on the dataset page: https://huggingface.co/datasets/MLRS/OPUS-MT-EN-Fixed.opus-mt-ct2opus-mt-arabic-benchmark-2026-03-28
OPUS-MT Arabic-English Translation Benchmark
Experiment Details
Date: 2026-03-28
Models Tested:
Helsinki-NLP/opus-mt-en-ar (English → Arabic)
Helsinki-NLP/opus-mt-ar-en (Arabic → English)
Total Tests: 9
Domain: NLP / Translation
Summary
Metric
Value
MSA Accuracy Rate
100%
Dialectal Accuracy Rate
0%
Avg Latency (MSA)
5.67s
Avg Latency (Dialectal)
0.5s
Key Finding
OPUS-MT handles Modern Standard Arabic (MSA) well but truncates… See the full description on the dataset page: https://huggingface.co/datasets/O96a/opus-mt-arabic-benchmark-2026-03-28.autotrain-data-opus-mt-en-zh_hanz
AutoTrain Dataset for project: opus-mt-en-zh_hanz
Dataset Description
This dataset has been automatically processed by AutoTrain for project opus-mt-en-zh_hanz.
Languages
The BCP-47 code for the dataset's language is en2zh.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"source": "And then I hear something.",
"target": "\u63a5\u7740\u542c\u5230\u4ec0\u4e48\u52a8\u9759\u3002"… See the full description on the dataset page: https://huggingface.co/datasets/darcy01/autotrain-data-opus-mt-en-zh_hanz.opus-mt-tatoeba-conlangThis is just a list of Tatoeba snapshots that I used to fine-tune Opus MT!
Why?
Because we need to like train translation models that support every single language on Earth. Today, we're starting with conlangs, and adding weird languages, and moving out only of Tatoeba and putting more data sources (for more languages)!
Variants (I will update later)
Variant
File Used to Train
Model Used
Languages Supported
Finetuned From
opus-mt-en-jbo
English-Lojban… See the full description on the dataset page: https://huggingface.co/datasets/MihaiPopa-1/opus-mt-tatoeba-conlang.opus-mt-tc-big-wiki-en-koMT-wmt14-500k-opus-mt-en-de
src (source text): WMT14 English-German dataset (train split, 500k sentences)
ref (reference text): original reference
h1,h2,h3,h4,h5: list of 5 translation candidates, each with log-probability score (sorted in descending order)
MT model: Helsinki-NLP/opus-mt-en-de
Translation direction: English → German
Beam search: num_beams=5, num_return_sequences=5)
opus-mt-en-bkm-37farmmind-opus-mt-int8-demo
FarmMind OPUS-MT int8 (demo)
Throwaway demo host for FarmMind offline machine translation — NOT production hosting.
int8-quantized ONNX exports of Helsinki-NLP OPUS-MT, redistributed under CC-BY-4.0
(attribution required).
opus-mt-en-es-onnx/ — English→Spanish, from Helsinki-NLP/opus-mt-en-es
opus-mt-es-en-onnx/ — Spanish→English, from Helsinki-NLP/opus-mt-es-en
Models © Helsinki-NLP (OPUS-MT), licensed CC-BY-4.0. ONNX Runtime components MIT.
opus-mt-en-bkm-60
