datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quickstart-3d
Dataset Card for quickstart-3d
This is a FiftyOne dataset with 200 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = fouh.load_from_hub("Voxel51/quickstart-3d")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/quickstart-3d.LLaVA-OneVision-1.5-Mid-Training-Webdataset-Quick-Start-3Mquickerlifelongwm-lifelongQuickdrawHDquickmt-train.de-en
quickmt de-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-commoncrawl_wmt13-1-deu-eng
Statmt-europarl_wmt13-7-deu-eng
Statmt-news_commentary_wmt18-13-deu-eng
Statmt-europarl-9-deu-eng
Statmt-europarl-7-deu-eng
Statmt-news_commentary-14-deu-eng
Statmt-news_commentary-15-deu-eng
Statmt-news_commentary-16-deu-eng
Statmt-news_commentary-17-deu-eng
Statmt-news_commentary-18-deu-eng… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.de-en.quickmt-train.it-en
quickmt it-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-europarl-7-ita-eng
Statmt-news_commentary-14-eng-ita
Statmt-news_commentary-15-eng-ita
Statmt-news_commentary-16-eng-ita
Statmt-news_commentary-17-eng-ita
Statmt-news_commentary-18-eng-ita
Statmt-news_commentary-18.1-eng-ita
Statmt-europarl-10-ita-eng
Tilde-eesc-2017-eng-ita
Tilde-ema-2016-eng-ita
Tilde-czechtourism-1-eng-ita… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.it-en.quick-canvas-benchmarkquickmt-train.hi-en
quickmt hi-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
IITB-hien_dev-1.5-hin-eng
Neulab-tedtalks_test-1-eng-hin
Google-wmt24pp-1-eng-hin_IN
IITB-hien_test-1.5-hin-eng
Statmt-news_commentary-14-eng-hin
Statmt-news_commentary-15-eng-hin
Statmt-news_commentary-16-eng-hin
Statmt-news_commentary-17-eng-hin
Statmt-news_commentary-18-eng-hin
Statmt-news_commentary-18.1-eng-hin
Statmt-pmindia-1-eng-hin… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.hi-en.split-text-quickmt-train.zh-enspilit text (sentence)
quickdraw
Dataset Card for Quick, Draw!
This is a processed version of Google's Quick, Draw dataset to be compatible with the latest versions of 🤗 Datasets that support .parquet files. NOTE: this dataset only contains the "preprocessed_bitmaps" subset of the original dataset.
madlad400-en-backtranslated-ar
madlad400 en Sample Translated into ar
This dataset is a subset of MADLAD-400 translated from en into ar by the quickmt/quickmt-en-ar model (beam size 4) intended to be used for training translation models from ar into en.
quickmt-train.bn-en
quickmt bn-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
OPUS-ccaligned-v1-ben-eng
OPUS-ccmatrix-v1-ben-eng
OPUS-nllb-v1-ben-eng
OPUS-wikimatrix-v1-ben-eng
Statmt-pmindia-1-eng-ben
JoshuaDec-indian_training-1-ben-eng
JoshuaDec-indian_dev-1-ben-eng
JoshuaDec-indian_test-1-ben-eng
JoshuaDec-indian_devtest-1-ben-eng
JoshuaDec-indian_dict-1-ben-eng
Neulab-tedtalks_train-1-eng-ben… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.bn-en.quickmt-train.pt-en
quickmt pt-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-europarl-7-por-eng
Statmt-news_commentary-14-eng-por
Statmt-news_commentary-15-eng-por
Statmt-news_commentary-16-eng-por
Statmt-news_commentary-17-eng-por
Statmt-news_commentary-18-eng-por
Statmt-news_commentary-18.1-eng-por
Statmt-europarl-10-por-eng
Tilde-eesc-2017-eng-por
Tilde-ema-2016-eng-por
Tilde-czechtourism-1-eng-por… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.pt-en.quickmt-train.es-en
quickmt es-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-newstest-2009-eng-spa
Statmt-newstest-2010-eng-spa
Statmt-newstest-2011-eng-spa
Statmt-europarl_wmt13-7-spa-eng
Statmt-europarl-7-spa-eng
Statmt-news_commentary-14-eng-spa
Statmt-news_commentary-15-eng-spa
Statmt-news_commentary-16-eng-spa
Statmt-news_commentary-17-eng-spa
Statmt-news_commentary-18-eng-spa… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.es-en.quickmt-train.id-en
quickmt id-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-news_commentary-14-eng-ind
Statmt-news_commentary-15-eng-ind
Statmt-news_commentary-16-eng-ind
Statmt-news_commentary-17-eng-ind
Statmt-news_commentary-18-eng-ind
Statmt-news_commentary-18.1-eng-ind
Statmt-ccaligned-1-eng-ind_ID
Facebook-wikimatrix-1-eng-ind
Neulab-tedtalks_train-1-eng-ind
Neulab-tedtalks_dev-1-eng-ind… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.id-en.quickmt-train.tr-en
quickmt tr-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-newsdev_tren-2016-tur-eng
Statmt-newsdev_entr-2016-eng-tur
Statmt-newstest_tren-2016-tur-eng
Statmt-newstest_entr-2016-eng-tur
Statmt-newstest_entr-2017-eng-tur
Statmt-newstest_tren-2017-tur-eng
Statmt-newstest_entr-2018-eng-tur
Statmt-newstest_tren-2018-tur-eng
Statmt-ccaligned-1-eng-tur_TR
Tilde-worldbank-1-eng-tur… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.tr-en.quickmt-train.ro-en
quickmt ro-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-europarl-7-ron-eng
Statmt-newsdev_enro-2016-eng-ron
Statmt-newsdev_roen-2016-ron-eng
Statmt-newstest_enro-2016-eng-ron
Statmt-newstest_roen-2016-ron-eng
Statmt-europarl-10-ron-eng
Statmt-ccaligned-1-eng-ron_RO
ParaCrawl-paracrawl-6-eng-ron
ParaCrawl-paracrawl-7.1-eng-ron
ParaCrawl-paracrawl-8-eng-ron
ParaCrawl-paracrawl-9-eng-ron… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.ro-en.quickmt-train.da-en
quickmt da-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-europarl-7-dan-eng
Statmt-europarl-10-dan-eng
Statmt-ccaligned-1-dan_DK-eng
ParaCrawl-paracrawl-6-eng-dan
ParaCrawl-paracrawl-7.1-eng-dan
ParaCrawl-paracrawl-8-eng-dan
ParaCrawl-paracrawl-9-eng-dan
Tilde-eesc-2017-dan-eng
Tilde-ema-2016-dan-eng
Tilde-ecb-2017-dan-eng
Tilde-rapid-2016-dan-eng
Facebook-wikimatrix-1-dan-eng… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.da-en.finetranslations-sample-ar-enSample of https://huggingface.co/datasets/HuggingFaceFW/finetranslations filtered and split into sentences by this script intended to be used for training sentence-level machine translation models.
quickdraw-small
Dataset Card for "quickdraw-small"
More Information needed
quickmt-train.cs-en
quickmt cs-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-commoncrawl_wmt13-1-ces-eng
Statmt-europarl_wmt13-7-ces-eng
Statmt-news_commentary_wmt18-13-ces-eng
Statmt-europarl-9-ces-eng
Statmt-europarl-7-ces-eng
Statmt-news_commentary-14-ces-eng
Statmt-news_commentary-15-ces-eng
Statmt-news_commentary-16-ces-eng
Statmt-news_commentary-17-ces-eng
Statmt-news_commentary-18-ces-eng… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.cs-en.quickmt-train.vi-en
quickmt vi-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
OPUS-ccaligned-v1-eng-vie
Facebook-wikimatrix-1-eng-vie
OPUS-ccmatrix-v1-eng-vie
Neulab-tedtalks_train-1-eng-vie
Neulab-tedtalks_test-1-eng-vie
Neulab-tedtalks_dev-1-eng-vie
ELRC-hrw_dataset_v1-1-eng-vie
OPUS-elrc_3086_wikipedia_health-v1-eng-vie
OPUS-elrc_wikipedia_health-v1-eng-vie
OPUS-elrc_2922-v1-eng-vie
OPUS-gnome-v1-eng-vie… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.vi-en.ov2_quickstart
OV2 Quickstart
Quickstart bundle for LLaVA-OneVision-2 (OV2). Contains everything needed to reproduce SFT training and run inference: packed SFT data, ready-to-use HF inference model, Megatron-Core checkpoint, and a Megatron training environment snapshot.
Total size: ~374 GB across 329 files.
Contents
1. packed_mixed_sft_cap_v30s/ — 308 GB
Packed mixed SFT (image + video + caption) dataset, sharded for distributed training via Megatron-Energon.
Format:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/ov2_quickstart.quickmt-train.pl-en
quickmt pl-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Google-wmt24pp-1-eng-pol_PL
Statmt-newsdev_plen-2020-pol-eng
Statmt-newsdev_enpl-2020-eng-pol
Statmt-europarl-10-pol-eng
Statmt-ccaligned-1-eng-pol_PL
Tilde-eesc-2017-eng-pol
Tilde-ema-2016-eng-pol
Tilde-czechtourism-1-eng-pol
Tilde-ecb-2017-eng-pol
Tilde-rapid-2019-eng-pol
Tilde-worldbank-1-eng-pol
Facebook-wikimatrix-1-eng-pol… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.pl-en.quickmt-train.el-en
quickmt el-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-europarl-10-ell-eng
Facebook-wikimatrix-1-ell-eng
Neulab-tedtalks_train-1-eng-ell
Neulab-tedtalks_test-1-eng-ell
Neulab-tedtalks_dev-1-eng-ell
OPUS-books-v1-ell-eng
OPUS-dgt-v2019-ell-eng
OPUS-dgt-v4-ell-eng
OPUS-ecb-v1-ell-eng
OPUS-ecdc-v20160316-ell-eng
OPUS-elitr_eca-v1-ell-eng
OPUS-elra_w0164-v1-ell-eng
OPUS-elra_w0196-v1-ell-eng… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.el-en.quickdrawquickmt-train.ja-en
quickmt ja-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-generaltest-2022_refA-eng-jpn
Statmt-generaltest-2022_refA-jpn-eng
Statmt-newstest_enja-2020-eng-jpn
Statmt-newstest_jaen-2020-jpn-eng
Statmt-newstest_enja-2021-eng-jpn
Statmt-newstest_jaen-2021-jpn-eng
Statmt-news_commentary-14-eng-jpn
Statmt-news_commentary-15-eng-jpn
Statmt-news_commentary-16-eng-jpn
Statmt-news_commentary-17-eng-jpn… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.ja-en.quickmt-train.zh-en
quickmt zh-en Training Corpus
Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:
Statmt-news_commentary_wmt18-13-zho-eng
Statmt-news_commentary-14-eng-zho
Statmt-news_commentary-15-eng-zho
Statmt-news_commentary-16-eng-zho
Statmt-news_commentary-17-eng-zho
Statmt-news_commentary-18-eng-zho
Statmt-news_commentary-18.1-eng-zho
Statmt-wiki_titles-1-zho-eng
Statmt-wiki_titles-2-zho-eng
Statmt-wikititles-3-zho-eng… See the full description on the dataset page: https://huggingface.co/datasets/quickmt/quickmt-train.zh-en.newscrawl2024-en-backtranslated-fr
NewsCrawl 2023 en Translated into fr
This dataset is a subset of NewsCrawl-en-2024 translated from en into fr by the quickmt/quickmt-en-fr model (beam size 4) intended to be used for training translation models from fr into en.
References
Findings of the WMT24 General Machine Translation Shared Task: The LLM Era Is Here but MT Is Not Solved Yet (Kocmi et al., WMT 2024)
