CoolFace
14 results

wmt16

wmt /wmt16 Dataset Card for "wmt16" Dataset Summary Warning: There are issues with the Common Crawl corpus data (training-parallel-commoncrawl.tgz): Non-English files contain many English sentences. Their "parallel" sentences in English are not aligned: they are uncorrelated with their counterpart. We have contacted the WMT organizers, and in response, they have indicated that they do not have plans to update the Common Crawl corpus data. Their rationale pertains… See the full description on the dataset page: https://huggingface.co/datasets/wmt/wmt16.texttranslation1M<n<10M27 likes5.5k downloads2y agoHugging Faceqanastek /WMT-16-PubMedWMT'16 Biomedical Translation Task - PubMed parallel datasets http://www.statmt.org/wmt16/biomedical-translation-task.htmltexttranslation100K<n<1M6 likes262 downloads4y agoHugging FaceExpertFlowPredictor /wmt16_Qwen3-30B-A3B_moe_patternstext10K<n<100K0 likes181 downloads11mo agoHugging Faceenimai /MuST-C-and-WMT16-de-entext1M<n<10M0 likes116 downloads4y agoHugging FaceLvxue /wmt16text1M<n<10M0 likes49 downloads4y agoHugging Facestas /wmt16-en-ro-pre-processed WMT16 English-Romanian Translation Data w/ further preprocessing The original instructions are here. This pre-processed dataset was created by running: git clone https://github.com/rsennrich/wmt16-scripts cd wmt16-scripts cd sample ./download_files.sh ./preprocess.sh It was originally used by transformers finetune_trainer.py The data itself resides at https://cdn-datasets.huggingface.co/translation/wmt_en_ro.tar.gz If you would like to convert it to jsonlines I've included a small… See the full description on the dataset page: https://huggingface.co/datasets/stas/wmt16-en-ro-pre-processed.text100K<n<1M0 likes44 downloads6y agoHugging Face