CoolFace
20 results

opus-mt

pythainlp /scb-mt-en-th-2020_mt-opus Dataset Card for "scb-mt-en-th-2020_mt-opus" More Information needed English-Thai scb-mt-en-th-2020 v1.0 and datasets listed in Open Parallel Corpus (OPUS) This dataset come from A large English–Thai parallel corpus from the web and machine-generated text that released at GitHub. texttranslation1M<n<10M3 likes95 downloads3y agoHugging FaceMLRS /OPUS-MT-EN-Fixed OPUS-100-Fixed: Tokenisation-Improved English-Maltese Dataset Overview OPUS-100-Fixed is an updated version of the OPUS-100 parallel English-Maltese dataset. This version addresses tokenisation inconsistencies in the Maltese text using the MLRS tokeniser, aiming to improve machine translation quality. The "en" column is the same as in the original OPUS-100 data, while the "mt" column has been corrected with the MLRS detokeniser. Citation If you use this… See the full description on the dataset page: https://huggingface.co/datasets/MLRS/OPUS-MT-EN-Fixed.texttranslation1M<n<10M3 likes58 downloads2y agoHugging FaceTimteamteem /opus-mt-ct2text1M<n<10M1 likes37 downloads4mo agoHugging FaceO96a /opus-mt-arabic-benchmark-2026-03-28 OPUS-MT Arabic-English Translation Benchmark Experiment Details Date: 2026-03-28 Models Tested: Helsinki-NLP/opus-mt-en-ar (English → Arabic) Helsinki-NLP/opus-mt-ar-en (Arabic → English) Total Tests: 9 Domain: NLP / Translation Summary Metric Value MSA Accuracy Rate 100% Dialectal Accuracy Rate 0% Avg Latency (MSA) 5.67s Avg Latency (Dialectal) 0.5s Key Finding OPUS-MT handles Modern Standard Arabic (MSA) well but truncates… See the full description on the dataset page: https://huggingface.co/datasets/O96a/opus-mt-arabic-benchmark-2026-03-28.texttranslationn<1K0 likes30 downloads6mo agoHugging Facedarcy01 /autotrain-data-opus-mt-en-zh_hanz AutoTrain Dataset for project: opus-mt-en-zh_hanz Dataset Description This dataset has been automatically processed by AutoTrain for project opus-mt-en-zh_hanz. Languages The BCP-47 code for the dataset's language is en2zh. Dataset Structure Data Instances A sample from this dataset looks as follows: [ { "source": "And then I hear something.", "target": "\u63a5\u7740\u542c\u5230\u4ec0\u4e48\u52a8\u9759\u3002"… See the full description on the dataset page: https://huggingface.co/datasets/darcy01/autotrain-data-opus-mt-en-zh_hanz.translation0 likes27 downloads4y agoHugging FaceMihaiPopa-1 /opus-mt-tatoeba-conlangThis is just a list of Tatoeba snapshots that I used to fine-tune Opus MT! Why? Because we need to like train translation models that support every single language on Earth. Today, we're starting with conlangs, and adding weird languages, and moving out only of Tatoeba and putting more data sources (for more languages)! Variants (I will update later) Variant File Used to Train Model Used Languages Supported Finetuned From opus-mt-en-jbo English-Lojban… See the full description on the dataset page: https://huggingface.co/datasets/MihaiPopa-1/opus-mt-tatoeba-conlang.translation10K<n<100K0 likes26 downloads3mo agoHugging Face