CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /wmt24pp WMT24++ This repository contains the human translation and post-edit data for the 55 en->xx language pairs released in the publication WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects. If you are interested in the MT/LLM system outputs and automatic metric scores, please see MTME. If you are interested in the images of the source URLs for each document, please see here. Schema Each language pair is stored in its own jsonl file. Each row… See the full description on the dataset page: https://huggingface.co/datasets/google/wmt24pp.texttranslation10K<n<100K95 likes16k downloads2mo agoHugging Face02Lulu19971017 /wmt24pp WMT24++ This repository contains the human translation and post-edit data for the 55 en->xx language pairs released in the publication WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects. If you are interested in the MT/LLM system outputs and automatic metric scores, please see MTME. If you are interested in the images of the source URLs for each document, please see here. Schema Each language pair is stored in its own jsonl file. Each row is… See the full description on the dataset page: https://huggingface.co/datasets/Lulu19971017/wmt24pp.texttranslation10K<n<100K0 likes407 downloads9mo agoHugging Face03synquid /wmt24pp WMT24++ This repository contains the human translation and post-edit data for the 55 en->xx language pairs released in the publication WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects. If you are interested in the MT/LLM system outputs and automatic metric scores, please see MTME. If you are interested in the images of the source URLs for each document, please see here. Schema Each language pair is stored in its own jsonl file. Each row is… See the full description on the dataset page: https://huggingface.co/datasets/synquid/wmt24pp.texttranslation10K<n<100K0 likes360 downloads8mo agoHugging Face04alwaysgood /wmt24pp-kr-reversed WMT24++ Parallel Mix (Direction-Flipped) Source dataset dir: /Users/inertia/Desktop/preprocessing/processed/wmt24pp_parallel_80_20 Transformation: direction flip per row (source/target swap) Pair+sample_key overlap with source: 0 Primary(stats): {'en->ko': 450, 'ko->en': 449} Auxiliary(flipped): ja->en, en->ja, zh->en, en->zh Target primary ratio: 0.8000 Achieved primary ratio: 0.7998 Tag template: <{tgt_upper}> Text template: {src_tag} {source} {tgt_tag} {target}… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/wmt24pp-kr-reversed.texttranslation1K<n<10K0 likes48 downloads5mo agoHugging Face05amalia-llm /wmt24pp_xx_to_pt_10k WMT24PP XX to PT 10k This is a dataset of translations from various source languages into Portuguese. The pairs are derived from google/wmt24pp, where the same 998 English sentences are each translated into many target languages. For every sentence, we take its translation in another language as the source and its Portuguese translation as the target, yielding {source language}→Portuguese pairs across many languages. Note that both sides are translations of the… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wmt24pp_xx_to_pt_10k.texttranslation10K<n<100K1 likes25 downloads3mo agoHugging Face06alwaysgood /wmt24pp-kr WMT24++ Parallel Mix Primary(all): en->ko, ko->en Primary mode: disjoint_halves Auxiliary(sampled, no oversampling): en->ja, ja->en, en->zh, zh->en Target primary ratio: 0.8000 Achieved primary ratio: 0.7998 Tag template: <{tgt_upper}> Text template: {src_tag} {source} {tgt_tag} {target} Normalize doubled quotes: True Strip control chars: True Drop bad source: True Drop canary: True Drop @user handles: True Drop one-word sentences: True Columns id, dataset, pair_config… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/wmt24pp-kr.texttranslation1K<n<10K0 likes23 downloads5mo agoHugging Face07NM-development /wmt24pp-ce WMT24++ Reference Translations for Chechen Description WMT24++ benchmark in Chechen Original WMT24++ benchmark (55 languages): https://huggingface.co/datasets/google/wmt24pp The reference translation have been created by human translator based on the Russian version of WMT24++ The dataset uses both cases of Cyrillic Palochka Letter where it is grammatically correct. For preparation of the Chechen version of the dataset we hired a professional native speaker… See the full description on the dataset page: https://huggingface.co/datasets/NM-development/wmt24pp-ce.texttranslationn<1K0 likes22 downloads2mo agoHugging Face08Rashik24 /wmt24pp-en-bn WMT24++ English-Bengali Filtered Subset This dataset is a filtered subset of google/wmt24pp using en-bn_IN.jsonl as the source file. Filtering rules: keep rows where is_bad_source is false keep rows with non-empty source keep rows with non-empty target Result: input rows: 998 kept rows: 960 removed rows: 38 Columns are preserved from the source JSONL. texttranslationn<1K0 likes18 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.