datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
QuickdrawHDquickmt-train.de-en
quickmt de-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-commoncrawl_wmt13-1-deu-eng
Statmt-europarl_wmt13-7-deu-eng
Statmt-news_commentary_wmt18-13-deu-eng
Statmt-europarl-9-deu-eng
Statmt-europarl-7-deu-eng
Statmt-news_commentary-14-deu-eng
Statmt-news_commentary-15-deu-eng
Statmt-news_commentary-16-deu-eng
Statmt-news_commentary-17-deu-eng
Statmt-news_commentary-18-deu-eng… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.de-en.quickmt-train.it-en
quickmt it-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-europarl-7-ita-eng
Statmt-news_commentary-14-eng-ita
Statmt-news_commentary-15-eng-ita
Statmt-news_commentary-16-eng-ita
Statmt-news_commentary-17-eng-ita
Statmt-news_commentary-18-eng-ita
Statmt-news_commentary-18.1-eng-ita
Statmt-europarl-10-ita-eng
Tilde-eesc-2017-eng-ita
Tilde-ema-2016-eng-ita
Tilde-czechtourism-1-eng-ita… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.it-en.quickmt-train.hi-en
quickmt hi-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
IITB-hien_dev-1.5-hin-eng
Neulab-tedtalks_test-1-eng-hin
Google-wmt24pp-1-eng-hin_IN
IITB-hien_test-1.5-hin-eng
Statmt-news_commentary-14-eng-hin
Statmt-news_commentary-15-eng-hin
Statmt-news_commentary-16-eng-hin
Statmt-news_commentary-17-eng-hin
Statmt-news_commentary-18-eng-hin
Statmt-news_commentary-18.1-eng-hin
Statmt-pmindia-1-eng-hin… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.hi-en.split-text-quickmt-train.zh-enspilit text (sentence)
madlad400-en-backtranslated-ar
madlad400 en Sample Translated into ar
This dataset is a subset of MADLAD-400 translated from en into ar by the quickmt/quickmt-en-ar model (beam size 4) intended to be used for training translation models from ar into en.
quickmt-train.bn-en
quickmt bn-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
OPUS-ccaligned-v1-ben-eng
OPUS-ccmatrix-v1-ben-eng
OPUS-nllb-v1-ben-eng
OPUS-wikimatrix-v1-ben-eng
Statmt-pmindia-1-eng-ben
JoshuaDec-indian_training-1-ben-eng
JoshuaDec-indian_dev-1-ben-eng
JoshuaDec-indian_test-1-ben-eng
JoshuaDec-indian_devtest-1-ben-eng
JoshuaDec-indian_dict-1-ben-eng
Neulab-tedtalks_train-1-eng-ben… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.bn-en.quickmt-train.pt-en
quickmt pt-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-europarl-7-por-eng
Statmt-news_commentary-14-eng-por
Statmt-news_commentary-15-eng-por
Statmt-news_commentary-16-eng-por
Statmt-news_commentary-17-eng-por
Statmt-news_commentary-18-eng-por
Statmt-news_commentary-18.1-eng-por
Statmt-europarl-10-por-eng
Tilde-eesc-2017-eng-por
Tilde-ema-2016-eng-por
Tilde-czechtourism-1-eng-por… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.pt-en.quickmt-train.es-en
quickmt es-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-newstest-2009-eng-spa
Statmt-newstest-2010-eng-spa
Statmt-newstest-2011-eng-spa
Statmt-europarl_wmt13-7-spa-eng
Statmt-europarl-7-spa-eng
Statmt-news_commentary-14-eng-spa
Statmt-news_commentary-15-eng-spa
Statmt-news_commentary-16-eng-spa
Statmt-news_commentary-17-eng-spa
Statmt-news_commentary-18-eng-spa… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.es-en.quickmt-train.id-en
quickmt id-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-news_commentary-14-eng-ind
Statmt-news_commentary-15-eng-ind
Statmt-news_commentary-16-eng-ind
Statmt-news_commentary-17-eng-ind
Statmt-news_commentary-18-eng-ind
Statmt-news_commentary-18.1-eng-ind
Statmt-ccaligned-1-eng-ind_ID
Facebook-wikimatrix-1-eng-ind
Neulab-tedtalks_train-1-eng-ind
Neulab-tedtalks_dev-1-eng-ind… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.id-en.quickmt-train.tr-en
quickmt tr-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-newsdev_tren-2016-tur-eng
Statmt-newsdev_entr-2016-eng-tur
Statmt-newstest_tren-2016-tur-eng
Statmt-newstest_entr-2016-eng-tur
Statmt-newstest_entr-2017-eng-tur
Statmt-newstest_tren-2017-tur-eng
Statmt-newstest_entr-2018-eng-tur
Statmt-newstest_tren-2018-tur-eng
Statmt-ccaligned-1-eng-tur_TR
Tilde-worldbank-1-eng-tur… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.tr-en.quickmt-train.ro-en
quickmt ro-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-europarl-7-ron-eng
Statmt-newsdev_enro-2016-eng-ron
Statmt-newsdev_roen-2016-ron-eng
Statmt-newstest_enro-2016-eng-ron
Statmt-newstest_roen-2016-ron-eng
Statmt-europarl-10-ron-eng
Statmt-ccaligned-1-eng-ron_RO
ParaCrawl-paracrawl-6-eng-ron
ParaCrawl-paracrawl-7.1-eng-ron
ParaCrawl-paracrawl-8-eng-ron
ParaCrawl-paracrawl-9-eng-ron… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.ro-en.quickmt-train.da-en
quickmt da-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-europarl-7-dan-eng
Statmt-europarl-10-dan-eng
Statmt-ccaligned-1-dan_DK-eng
ParaCrawl-paracrawl-6-eng-dan
ParaCrawl-paracrawl-7.1-eng-dan
ParaCrawl-paracrawl-8-eng-dan
ParaCrawl-paracrawl-9-eng-dan
Tilde-eesc-2017-dan-eng
Tilde-ema-2016-dan-eng
Tilde-ecb-2017-dan-eng
Tilde-rapid-2016-dan-eng
Facebook-wikimatrix-1-dan-eng… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.da-en.finetranslations-sample-ar-enSample of https://huggingface.co/datasets/HuggingFaceFW/finetranslations filtered and split into sentences by this script intended to be used for training sentence-level machine translation models.
quickmt-train.cs-en
quickmt cs-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-commoncrawl_wmt13-1-ces-eng
Statmt-europarl_wmt13-7-ces-eng
Statmt-news_commentary_wmt18-13-ces-eng
Statmt-europarl-9-ces-eng
Statmt-europarl-7-ces-eng
Statmt-news_commentary-14-ces-eng
Statmt-news_commentary-15-ces-eng
Statmt-news_commentary-16-ces-eng
Statmt-news_commentary-17-ces-eng
Statmt-news_commentary-18-ces-eng… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.cs-en.quickmt-train.vi-en
quickmt vi-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
OPUS-ccaligned-v1-eng-vie
Facebook-wikimatrix-1-eng-vie
OPUS-ccmatrix-v1-eng-vie
Neulab-tedtalks_train-1-eng-vie
Neulab-tedtalks_test-1-eng-vie
Neulab-tedtalks_dev-1-eng-vie
ELRC-hrw_dataset_v1-1-eng-vie
OPUS-elrc_3086_wikipedia_health-v1-eng-vie
OPUS-elrc_wikipedia_health-v1-eng-vie
OPUS-elrc_2922-v1-eng-vie
OPUS-gnome-v1-eng-vie… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.vi-en.quickmt-train.pl-en
quickmt pl-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Google-wmt24pp-1-eng-pol_PL
Statmt-newsdev_plen-2020-pol-eng
Statmt-newsdev_enpl-2020-eng-pol
Statmt-europarl-10-pol-eng
Statmt-ccaligned-1-eng-pol_PL
Tilde-eesc-2017-eng-pol
Tilde-ema-2016-eng-pol
Tilde-czechtourism-1-eng-pol
Tilde-ecb-2017-eng-pol
Tilde-rapid-2019-eng-pol
Tilde-worldbank-1-eng-pol
Facebook-wikimatrix-1-eng-pol… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.pl-en.quickmt-train.el-en
quickmt el-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-europarl-10-ell-eng
Facebook-wikimatrix-1-ell-eng
Neulab-tedtalks_train-1-eng-ell
Neulab-tedtalks_test-1-eng-ell
Neulab-tedtalks_dev-1-eng-ell
OPUS-books-v1-ell-eng
OPUS-dgt-v2019-ell-eng
OPUS-dgt-v4-ell-eng
OPUS-ecb-v1-ell-eng
OPUS-ecdc-v20160316-ell-eng
OPUS-elitr_eca-v1-ell-eng
OPUS-elra_w0164-v1-ell-eng
OPUS-elra_w0196-v1-ell-eng… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.el-en.quickmt-train.ja-en
quickmt ja-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-generaltest-2022_refA-eng-jpn
Statmt-generaltest-2022_refA-jpn-eng
Statmt-newstest_enja-2020-eng-jpn
Statmt-newstest_jaen-2020-jpn-eng
Statmt-newstest_enja-2021-eng-jpn
Statmt-newstest_jaen-2021-jpn-eng
Statmt-news_commentary-14-eng-jpn
Statmt-news_commentary-15-eng-jpn
Statmt-news_commentary-16-eng-jpn
Statmt-news_commentary-17-eng-jpn… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.ja-en.quickmt-train.zh-en
quickmt zh-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-news_commentary_wmt18-13-zho-eng
Statmt-news_commentary-14-eng-zho
Statmt-news_commentary-15-eng-zho
Statmt-news_commentary-16-eng-zho
Statmt-news_commentary-17-eng-zho
Statmt-news_commentary-18-eng-zho
Statmt-news_commentary-18.1-eng-zho
Statmt-wiki_titles-1-zho-eng
Statmt-wiki_titles-2-zho-eng
Statmt-wikititles-3-zho-eng… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.zh-en.newscrawl2024-en-backtranslated-fr
NewsCrawl 2023 en Translated into fr
This dataset is a subset of NewsCrawl-en-2024 translated from en into fr by the quickmt/quickmt-en-fr model (beam size 4) intended to be used for training translation models from fr into en.
References
Findings of the WMT24 General Machine Translation Shared Task: The LLM Era Is Here but MT Is Not Solved Yet (Kocmi et al., WMT 2024)
P3-Latvian-QuickMTThis is an automatically translated version of P3 (Public Pool of Prompts) using quickmt-en-lv.
Languages
The data in P3-Latvian-Full are in Latvian (BCP-47 lv).
Dataset Structure
Data Instances
An example of "train" looks as follows:
{
'answer_choices': ['mobilais tālrunis', 'televīzija', 'ledusskapis', 'lidmašīna'],
'inputs_pretokenized': 'Kura tehnoloģija tika izstrādāta pavisam nesen? Iespējas: - mobilais tālrunis - televizors - ledusskapis - lidmašīna'… See the full description on the dataset page: https://huggingface.co/datasets/matiss/P3-Latvian-QuickMT.newscrawl2024-en-backtranslated-zh
NewsCrawl 2023 en Translated into zh
This dataset is a subset of NewsCrawl-en-2024 translated from en into zh by the quickmt/quickmt-en-zh model (beam size 4) intended to be used for training translation models from zh into en.
References
Findings of the WMT24 General Machine Translation Shared Task: The LLM Era Is Here but MT Is Not Solved Yet (Kocmi et al., WMT 2024)
quickmt-train.fr-en
quickmt fr-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication, and basic filtering with quickmt:
Statmt-commoncrawl_wmt13-1-fra-eng
Statmt-europarl_wmt13-7-fra-eng
Statmt-europarl-7-fra-eng
Statmt-news_commentary-14-eng-fra
Statmt-news_commentary-15-eng-fra
Statmt-news_commentary-16-eng-fra
Statmt-news_commentary-17-eng-fra
Statmt-news_commentary-18-eng-fra
Statmt-news_commentary-18.1-eng-fra
Statmt-newstest_fren-2014-fra-eng… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.fr-en.newscrawl2024-en-backtranslated-is
NewsCrawl 2023 en Translated into is
This dataset is a subset of NewsCrawl-en-2024 translated from en into is by the quickmt/quickmt-en-is model (beam size 4) intended to be used for training translation models from is into en.
References
Findings of the WMT24 General Machine Translation Shared Task: The LLM Era Is Here but MT Is Not Solved Yet (Kocmi et al., WMT 2024)
quickmt-train.th-en
quickmt th-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-ccaligned-1-eng-tha_TH
Neulab-tedtalks_train-1-eng-tha
Neulab-tedtalks_test-1-eng-tha
Neulab-tedtalks_dev-1-eng-tha
ELRC-wikipedia_health-1-eng-tha
ELRC-hrw_dataset_v1-1-eng-tha
OPUS-ccaligned-v1-eng-tha
OPUS-elrc_3048_wikipedia_health-v1-eng-tha
OPUS-elrc_wikipedia_health-v1-eng-tha
OPUS-elrc_2922-v1-eng-tha
OPUS-gnome-v1-eng-tha… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.th-en.quickmt-train.lv-en
quickmt pl-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-europarl-7-lav-eng
Statmt-newsdev_lven-2017-lav-eng
Statmt-newsdev_enlv-2017-eng-lav
Statmt-newstest_lven-2017-lav-eng
Statmt-europarl-10-lav-eng
Statmt-ccaligned-1-eng-lav_LV
ParaCrawl-paracrawl-6-eng-lav
ParaCrawl-paracrawl-7.1-eng-lav
ParaCrawl-paracrawl-8-eng-lav
ParaCrawl-paracrawl-9-eng-lav
Tilde-eesc-2017-eng-lav… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.lv-en.quickmt-train.ko-en
quickmt ko-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
OPUS-ccaligned-v1-eng-kor
OPUS-ccmatrix-v1-eng-kor
OPUS-elrc_3070_wikipedia_health-v1-eng-kor
OPUS-elrc_wikipedia_health-v1-eng-kor
OPUS-elrc_2922-v1-eng-kor
OPUS-gnome-v1-eng-kor
OPUS-globalvoices-v2017q3-eng-kor
OPUS-globalvoices-v2018q4-eng-kor
OPUS-kde4-v2-eng-kor
OPUS-linguatools_wikititles-v2014-eng-kor
OPUS-mdn_web_docs-v20230925-eng-kor… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.ko-en.quickmt-train.sv-en
quickmt sv-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-europarl-7-swe-eng
Statmt-dcep_wmt17-1-swe-eng
Statmt-books_wmt17-1-swe-eng
Statmt-europarl-10-swe-eng
Statmt-ccaligned-1-eng-swe_SE
ParaCrawl-paracrawl-6-eng-swe
ParaCrawl-paracrawl-8-eng-swe
ParaCrawl-paracrawl-9-eng-swe
Tilde-eesc-2017-eng-swe
Tilde-ema-2016-eng-swe
Tilde-ecb-2017-eng-swe
Tilde-rapid-2016-eng-swe… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.sv-en.quickmt-train.ur-en
quickmt ur-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-pmindia-1-eng-urd
Statmt-ccaligned-1-eng-urd_PK
JoshuaDec-indian_training-1-urd-eng
JoshuaDec-indian_dev-1-urd-eng
JoshuaDec-indian_test-1-urd-eng
JoshuaDec-indian_devtest-1-urd-eng
JoshuaDec-indian_dict-1-urd-eng
Neulab-tedtalks_train-1-eng-urd
Neulab-tedtalks_test-1-eng-urd
Neulab-tedtalks_dev-1-eng-urd
ELRC-hrw_dataset_v1-1-eng-urd… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.ur-en.
