CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zouhar /last-translation-benchmark Last Translation Benchmark Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. Standard benchmarks for machine translation evaluation are often either trivial (having few authentic mistakes) or unrealistic (overly synthetically contrived). Furthermore, automatic translation metrics become less reliable and reward-hacked as models get stronger, and their outputs are… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/last-translation-benchmark.texttranslation1K<n<10K60 likes966 downloads21d agoHugging Face02ltg /norsumm-nob-nno-translation Nynorsk-Bokmål translation pairs A multi-sentence parallel corpus of manual Nynorsk-Bokmål translations. These translations were extracted from the SamiaT/NorSumm dataset. You can read more about how the original dataset was created (including details about the manual translation process) in Benchmarking Abstractive Summarisation: A Dataset of Human-authored Summaries of Norwegian News Articles by Samia Touileb et al.. Contact David Samuel (davisamu@ifi.uio.no)… See the full description on the dataset page: https://huggingface.co/datasets/ltg/norsumm-nob-nno-translation.texttranslationn<1K1 likes858 downloads8mo agoHugging Face03McGill-NLP /speech-translation-and-summarization English-Centric Multilingual Audio Dataset This dataset contains generated article and summary audio for English-centric multilingual directions. Each direction folder contains metadata JSONL files and corresponding audio files for few_shot and test splits. Included directions amharic_english / english_amharic arabic_english / english_arabic bengali_english / english_bengali chinese_simplified_english / english_chinese_simplified english_english french_english /… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/speech-translation-and-summarization.audioautomatic-speech-recognition10K<n<100K6 likes760 downloads1mo agoHugging Face04recursal /Europarl-Translation-Instruct Dataset Card for Europarl-Translation-Instruct Waifu to catch your attention. Dataset Details Dataset Description europarl-translation-instruct is a translation instruct dataset built from europarl data. Curated by: M8than Funded by: Recursal.ai Shared by: M8than Language(s) (NLP): English instruct (but various languages in) License: cc-by-sa-4.0 Dataset Sources Source Data: https://www.statmt.org/europarl/ (Transcript source) Processing… See the full description on the dataset page: https://huggingface.co/datasets/recursal/Europarl-Translation-Instruct.texttext-generation10M<n<100M4 likes495 downloads2y agoHugging Face05LLaMAX /BenchMAX_General_Translation Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_General_Translation is a dataset of BenchMAX, which evaluates the translation capability on the general domain. We collect parallel test data from Flore-200, TED-talk, and WMT24. Usage Run the following commands to generate… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_General_Translation.texttranslation100K<n<1M0 likes435 downloads1y agoHugging Face06LLaMAX /BenchMAX_Domain_Translation Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Domain_Translation is a dataset of BenchMAX, which evaluates the translation capability on specific domains. We collect the domain multi-way parallel data from other tasks in BenchMAX, such as math data, code data, etc. Each sample contains one… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Domain_Translation.texttranslation10K<n<100K0 likes319 downloads2y agoHugging Face07talmp /en-vi-translation To join all training set files together run python join_dataset.py file, final result will be join_dataset.json file texttranslation1M<n<10M11 likes272 downloads3y agoHugging Face08botisan-ai /cantonese-mandarin-translations Dataset Card for cantonese-mandarin-translations Dataset Summary This is a machine-translated parallel corpus between Cantonese (a Chinese dialect that is mainly spoken by Guangdong (province of China), Hong Kong, Macau and part of Malaysia) and Chinese (written form, in Simplified Chinese). Supported Tasks and Leaderboards N/A Languages Cantonese (yue) Simplified Chinese (zh-CN) Dataset Structure JSON lines with yue field and zh field… See the full description on the dataset page: https://huggingface.co/datasets/botisan-ai/cantonese-mandarin-translations.texttranslation10K<n<100K31 likes124 downloads3y agoHugging Face09mesolitica /standard-malay-translation-instructionstext1M<n<10M0 likes115 downloads3y agoHugging Face10salmankhanpm /translation-checkpointstext1K<n<10K0 likes111 downloads7mo agoHugging Face11david9dragon9 /shp_translationsThis dataset contains translations of three splits (askscience, explainlikeimfive, legaladvice) of the Stanford Human Preference (SHP) dataset, used for training domain-invariant reward models. The translation was conducted using the No Language Left Behind (NLLB) 3.3 B 200 model. References: Stanford Human Preference Dataset: https://huggingface.co/datasets/stanfordnlp/SHP NLLB: https://huggingface.co/facebook/nllb-200-3.3B textquestion-answering100K<n<1M0 likes107 downloads2y agoHugging Face12joonkeene /Patent_Translation_datasetPatent Dataset .. text1M<n<10M0 likes99 downloads1y agoHugging Face13agentlans /translation-quality Multilingual Translation Quality Dataset This dataset provides multilingual text chunks translated into English, accompanied by automated quality evaluations generated by multiple large language models. Dataset Details Source Data: agentlans/HuggingFaceFW-finetranslations-100-languages-sample Target Language: English Content: Multilingual chunks mapped to their English translations alongside automated judge scores. Evaluation Methodology The… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/translation-quality.tabulartranslation100K<n<1M0 likes98 downloads20d agoHugging Face14Emulated-Inc /parallel-translation-training-pool Parallel translation training pool Sentences in eleven languages beside their translations, from five public parallel corpora read at the pinned revisions named below and laid out twice. Ten languages are paired with English in both directions, twenty directions in all. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 4975238 rows, one JSON object per line, with these fields. Field What it holds id a row identifier… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/parallel-translation-training-pool.texttranslation1M<n<10M0 likes95 downloads13d agoHugging Face15CodeTranslatorLLM /Code-Translationtext10K<n<100K1 likes90 downloads3y agoHugging Face16ic-org /wellbeing-in-translation Wellbeing in Translation Raw outputs and translated materials for Does AI Wellbeing Survive Translation? We test whether the unchanged CAIS 1-7 self-report battery measures the same positive-minus-negative gap after translation. Paper · Code · Source instrument Headline result Language sensitivity is specific to the model-battery pair. Model Gap spread across 7 languages English rank English stimulus / local battery Local stimulus / English battery… See the full description on the dataset page: https://huggingface.co/datasets/ic-org/wellbeing-in-translation.image10K<n<100K0 likes85 downloads1mo agoHugging Face17ParsBench /parsinlu-machine-translation-en-fa-alpaca-style ParsiNLU Machine Translation En-Fa in Alpaca Style This dataset is an Alpaca-style and instruction-included version of the ParsiNLU original dataset. texttranslation1M<n<10M7 likes84 downloads2y agoHugging Face18Eloquent /HalluciGen-Translation Task 2: HalluciGen - Tranlsation This dataset contains the trial and test splits per language pair for the Translation scenario of the HalluciGen task, which is part of the 2024 ELOQUENT lab. NOTE: A gold-labeled version of the dataset will be released in a new repository. Dataset schema id: unique identifier of the example langpair: the source and target language pair of the example source: original model input for translation hyp1: first alternative translation of the… See the full description on the dataset page: https://huggingface.co/datasets/Eloquent/HalluciGen-Translation.text1K<n<10K0 likes83 downloads2y agoHugging Face19RoxanneWsyw /ESFT-translationtext10K<n<100K0 likes80 downloads1y agoHugging Face20agentlans /wikidata-entity-translationstexttranslation10M<n<100M1 likes77 downloads3mo agoHugging Face21NNEngine /English-Hindi_Translation 📘 README.md 👉 Copy everything below into your repository README.md English–Hindi Massive Synthetic Translation Dataset 🧠 Overview This dataset is a large-scale synthetic parallel corpus for English → Hindi machine translation, designed to stress-test modern sequence-to-sequence models, tokenizers, and large-scale training pipelines. The corpus contains 10 million aligned sentence pairs generated using a high-entropy template engine with: 100+ subjects 100+… See the full description on the dataset page: https://huggingface.co/datasets/NNEngine/English-Hindi_Translation.texttranslation10M<n<100M0 likes76 downloads8mo agoHugging Face22squarelike /sharegpt_deepl_ko_translationhttps://github.com/jwj7140/Gugugo sharegpt_deepl_ko를 한-영 번역데이터로 변환한 데이터입니다. translation_data_sharegpt.json: 최대 약 1300자 분량의 번역 데이터 모음 translation_data_sharegpt_long.json: 1300자~7000자 분량의 번역 데이터 모음 translation_data_sharegpt_long_newlineClean.json: translation_data_sharegpt_long.json에서 개행이 번역되지 않은 항목를 제거한 데이터 translation_data_sharegpt_long_newlineClean.json: translation_data_sharegpt.json에서 개행이 번역되지 않은 항목를 제거한 데이터 sharegpt_deepl_ko에서 몇 가지의 데이터 전처리를 진행했습니다. text100K<n<1M17 likes74 downloads3y agoHugging Face23zouhar /optimal-reference-translationsThis is the dataset for two papers: Quality and Quantity of Machine Translation References for Automated Metrics [paper] - effect of reference quality and quantity on automatic metric performance, and Evaluating Optimal Reference Translations [paper] - creation of the data and human aspects of annotation and translation. Please see the original repository for more information and the raw data or contact the authors with any questions. Please make sure that you have the latest datasets… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/optimal-reference-translations.texttranslationn<1K1 likes74 downloads3y agoHugging Face24kunishou /databricks-dolly-69k-ja-en-translationThis dataset was created by automatically translating "databricks-dolly-15k" into Japanese.This dataset contains 69K ja-en-translation task data and is licensed under CC BY SA 3.0. Last Update : 2023-04-18 databricks-dolly-15k-jahttps://github.com/kunishou/databricks-dolly-15k-jadatabricks-dolly-15khttps://github.com/databrickslabs/dolly/tree/master/data text10K<n<100K15 likes64 downloads3y agoHugging Face25talmp /en-vi-translation-testtext1K<n<10K0 likes62 downloads3y agoHugging Face26ezosa /Dolci-Think-SFT-7B-translationstext1K<n<10K0 likes62 downloads6mo agoHugging Face27melanieyes /adaption-piguard-translation-handoff This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. adaption-piguard_translation_handoff This dataset contains 100 English prompts curated for translation guardrail research, comprising a balanced mix of 50 benign instructions and 50 prompt injection attempts. The samples include diverse content such as jailbreak personas, requests for harmful actions like spyware installation, and complex instruction overrides. Each entry is… See the full description on the dataset page: https://huggingface.co/datasets/melanieyes/adaption-piguard-translation-handoff.tabular1K<n<10K1 likes62 downloads10d agoHugging Face28AndreasThinks /welsh-translation-instructionThis is a set of Alpaca formatted Welsh-English translation instructions, obtained from the Welsh Government website. texttranslation10K<n<100K1 likes61 downloads2y agoHugging Face29agentlans /translation-contrastive-tripletstext100K<n<1M0 likes61 downloads17d agoHugging Face30CGICAI /cherokee-english-translation Cherokee–English Parallel Corpus (Archivist Project) A curated Cherokee (ᏣᎳᎩ / Tsalagi) ↔ English parallel corpus for machine translation, assembled from public sources, deduplicated, benchmark-decontaminated, and conflict-cleaned. Built to train and evaluate English→Cherokee translation models for one of the most endangered languages in North America. Files File Rows Purpose train_en2chr_v2.jsonl 138,307 Flagship training set. English→Cherokee SFT… See the full description on the dataset page: https://huggingface.co/datasets/CGICAI/cherokee-english-translation.tabulartranslation100K<n<1M0 likes57 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.