CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01natgillin /translations-raw natgillin/translations-raw Frozen, canonical raw bitext consolidated from upstream alvations/mtdata-raw* snapshots (since deleted). This is the read-only source-of-truth for downstream quality-filtering pipelines. 31,663 parquet files (1566.8 GB) 49 language pairs under data/<src-tgt>/ Schema: 5 columns — see below Read-only for downstream pipelines. Do not delete or modify. Schema Each parquet has 5 columns: column type description source string… See the full description on the dataset page: https://huggingface.co/datasets/natgillin/translations-raw.text1M<n<10M5 likes15k downloads4mo agoHugging Face02MaLA-LM /mala-bilingual-translation-corpus MaLA Corpus: Massive Language Adaptation Corpus This MaLA-LM/mala-bilingual-translation-corpus is the MaLA bilingual translation corpus, collected and processed from various sources. As a part of MaLA Corpus that aims to enhance massive language adaptation in many languages, it contains bilingual translation data (aka, parallel data and bitexts) in 2,500+ language pairs (500+ languages). Key statistics of all language pairs available at… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-bilingual-translation-corpus.texttranslation10B<n<100B8 likes6.5k downloads2mo agoHugging Face03aisingapore /NLG-Machine-Translationgated SEA Machine Translation SEA Machine Translation evaluates a model's ability to translate a document from a source language into a target language coherently and fluently. It is sampled from FLORES 200 for Burmese, Chinese, English, Indonesian, Khmer, Malay, Tamil, Thai, and Vietnamese, and NusaX for Indonesian, Javanese, and Sundanese. Supported Tasks and Leaderboards SEA Machine Translation is designed for evaluating chat or instruction-tuned large language models… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Machine-Translation.texttext-generation10K<n<100K2 likes2.8k downloads9mo agoHugging Face04ayymen /Weblate-Translations Dataset Card for Weblate Translations A dataset containing strings from projects hosted on Weblate and their translations into other languages. Please consider donating or contributing to Weblate if you find this dataset useful. To avoid rows with values like "None" and "N/A" being interpreted as missing values, pass the keep_default_na parameter like this: from datasets import load_dataset dataset = load_dataset("ayymen/Weblate-Translations", keep_default_na=False)… See the full description on the dataset page: https://huggingface.co/datasets/ayymen/Weblate-Translations.texttranslation10M<n<100M21 likes2.8k downloads2y agoHugging Face05premio-ai /OpenSubtitles_Translations_Datasettext100M<n<1B0 likes1.2k downloads2y agoHugging Face06ayymen /Pontoon-Translations Dataset Card for Pontoon Translations This is a dataset containing strings from various Mozilla projects on Mozilla's Pontoon localization platform and their translations into more than 200 languages. Source strings are in English. To avoid rows with values like "None" and "N/A" being interpreted as missing values, pass the keep_default_na parameter like this: from datasets import load_dataset dataset = load_dataset("ayymen/Pontoon-Translations", keep_default_na=False)… See the full description on the dataset page: https://huggingface.co/datasets/ayymen/Pontoon-Translations.texttranslation1M<n<10M19 likes1.1k downloads3y agoHugging Face07Tamazight-NLP /Weblate-Translations Dataset Card for Weblate Translations A dataset containing strings from projects hosted on Weblate and their translations into other languages. Please consider donating or contributing to Weblate if you find this dataset useful. Dataset Details Dataset Description Curated by: Mohamed Aymane Farhi Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): Check the README YAML metadata… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/Weblate-Translations.texttranslation10M<n<100M4 likes1.1k downloads3mo agoHugging Face08yipyany /ted-translation-decisions-en-zh TED Translation Decision Dataset (EN–ZH 英-简中) 🎁🎁 DATASET UPDATED REGULARLY! COME BACK FOR NEW ENTRIES! 🎁🎁 🧩 Searchable Keywords translation, EN-ZH, bilingual, rationale, subtitle, human decisions,TED Talks, translation choices, linguistic annotation, cross-lingual, semantic nuance, translation rationale dataset, Chinese translation, English translation dataset, word-level translation, interpretability, translation pedagogy, translation teaching… See the full description on the dataset page: https://huggingface.co/datasets/yipyany/ted-translation-decisions-en-zh.tabulartranslationn<1K1 likes1.1k downloads10h agoHugging Face09soynade-research /Bambara-Speech-Translation-Data AfVoices-Translated (Bambara-English) This is a Bambara speech translation dataset, which is built on the African Next Voices (AfVoices) Bambara ASR corpus. It provides English translations for the human-corrected subset of the original collection, creating a parallel corpus for Bambara-English machine translation and speech-to-text tasks. Methodology We machine-translated the human-validated transcriptions from AfVoices using the Oolel-translator repository. Inference… See the full description on the dataset page: https://huggingface.co/datasets/soynade-research/Bambara-Speech-Translation-Data.audioautomatic-speech-recognition100K<n<1M1 likes951 downloads7mo agoHugging Face10zouhar /last-translation-benchmark Last Translation Benchmark Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. Standard benchmarks for machine translation evaluation are often either trivial (having few authentic mistakes) or unrealistic (overly synthetically contrived). Furthermore, automatic translation metrics become less reliable and reward-hacked as models get stronger, and their outputs are… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/last-translation-benchmark.texttranslation1K<n<10K59 likes933 downloads19d agoHugging Face11ltg /norsumm-nob-nno-translation Nynorsk-Bokmål translation pairs A multi-sentence parallel corpus of manual Nynorsk-Bokmål translations. These translations were extracted from the SamiaT/NorSumm dataset. You can read more about how the original dataset was created (including details about the manual translation process) in Benchmarking Abstractive Summarisation: A Dataset of Human-authored Summaries of Norwegian News Articles by Samia Touileb et al.. Contact David Samuel (davisamu@ifi.uio.no)… See the full description on the dataset page: https://huggingface.co/datasets/ltg/norsumm-nob-nno-translation.texttranslationn<1K1 likes875 downloads8mo agoHugging Face12McGill-NLP /speech-translation-and-summarization English-Centric Multilingual Audio Dataset This dataset contains generated article and summary audio for English-centric multilingual directions. Each direction folder contains metadata JSONL files and corresponding audio files for few_shot and test splits. Included directions amharic_english / english_amharic arabic_english / english_arabic bengali_english / english_bengali chinese_simplified_english / english_chinese_simplified english_english french_english /… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/speech-translation-and-summarization.audioautomatic-speech-recognition10K<n<100K6 likes758 downloads1mo agoHugging Face13theblackcat102 /instruction_translationsTranslation of Instruction datasettexttext-generation100K<n<1M5 likes716 downloads4y agoHugging Face14ilsp /ancient-modern_greek_translations Dataset Card for Ancient-Modern Greek translations The Ancient-Modern Greek translations dataset includes 100 sentences of Ancient Greek texts manually translated into Modern Greek. Original texts and translations have been extracted from the web sources cited below. Δημοσθένους, Ὑπὲρ τῆς Ῥοδίων ἐλευθερίας (ell: Δημοσθένους, Υπέρ της Ελευθερίας των Ροδίων; eng: Demosthenes, On the Liberty of the Rhodians). Μτφρ. Β.Η. Τσακατίκας. χ.χ. Λόγοι του Δημοσθένη. Γ' Ολυνθιακός, Υπέρ της… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/ancient-modern_greek_translations.texttranslationn<1K2 likes688 downloads2y agoHugging Face15tech1984 /swe_smith_back_translationBack translate the swe-smith data to get the problem statment following the R2E sylte prompt. More details in https://github.com/SWE-bench/SWE-smith/issues/127 text10K<n<100K2 likes661 downloads1y agoHugging Face16SultanR /AraMix-Translation-Scores AraMix-Translation-Scores AdaMLLab/AraMix (minhash_deduped subset, 178,883,241 rows) with a machine-translation-detection score added to every document. All original columns are preserved. Columns column type description id string unchanged from AraMix source string unchanged from AraMix text string unchanged from AraMix mmbert_quality_score float64 AraMix's original mmbert_score, renamed mmbert_translated_score float64 new —… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/AraMix-Translation-Scores.tabular100M<n<1B0 likes594 downloads2mo agoHugging Face17recursal /Europarl-Translation-Instruct Dataset Card for Europarl-Translation-Instruct Waifu to catch your attention. Dataset Details Dataset Description europarl-translation-instruct is a translation instruct dataset built from europarl data. Curated by: M8than Funded by: Recursal.ai Shared by: M8than Language(s) (NLP): English instruct (but various languages in) License: cc-by-sa-4.0 Dataset Sources Source Data: https://www.statmt.org/europarl/ (Transcript source) Processing… See the full description on the dataset page: https://huggingface.co/datasets/recursal/Europarl-Translation-Instruct.texttext-generation10M<n<100M4 likes543 downloads2y agoHugging Face18Aletheia-ng /african_languages_translationtext1M<n<10M1 likes516 downloads1y agoHugging Face19hpprc /enwiki-translations英語WikipediaをLLMを用いて英日翻訳したデータセットです。 collectionサブセットは翻訳元となった英語Wikipediaの文章、datasetサブセットは英日翻訳されたテキストペアと翻訳元にした事例のcollectionサブセットにおけるidを収載したものです。 なお、出力の利用に際しては、翻訳に使用した各モデルの出力に関するライセンス規約に従ってください。 Phi3.5 MoEを利用して作成されたデータについては、Wikipediaのライセンスにしたがい、CC-BY-SA 4.0で利用可能であるものとします。 text10M<n<100M0 likes490 downloads2y agoHugging Face20NuBerea /translation-analysisgated NuBerea Translation Verse Texts Verse-level texts of historical Bible translations (Clementine Vulgate, Luther Bible 1545, Matthew's Bible 1537). Part of the NuBerea curated corpus estate of biblical and historical texts. Attribution Upstream Data Sources Source License Clementine Vulgate, NOCR Public Domain Luther Bible 1545, NOCR Public Domain Matthew's Bible 1537, Textus Receptus Bibles Public Domain NuBerea project. Licensed… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/translation-analysis.tabularfeature-extraction100K<n<1M0 likes456 downloads10d agoHugging Face21semeru /code-code-translation-java-csharp Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/code-to-code-trans in Semeru CodeXGLUE -- Code2Code Translation Task Definition Code translation aims to migrate legacy software from one programming language in a platform toanother. In CodeXGLUE, given a piece of Java (C#) code, the task is to translate the code into C#… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-translation-java-csharp.text10K<n<100K2 likes452 downloads3y agoHugging Face22LLaMAX /BenchMAX_General_Translation Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_General_Translation is a dataset of BenchMAX, which evaluates the translation capability on the general domain. We collect parallel test data from Flore-200, TED-talk, and WMT24. Usage Run the following commands to generate… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_General_Translation.texttranslation100K<n<1M0 likes417 downloads1y agoHugging Face23bot-yaya /undl_ru2en_translation Dataset Card for "undl_ru2en_translation" More Information needed text100K<n<1M0 likes396 downloads3y agoHugging Face24llm-jp /relaion2B-en-research-safe-japanese-translation relaion2B-en-research-safe-japanese-translation This dataset is the Japanese translation of the English subset of ReLAION-5B (laion/relaion2B-en-research-safe), translated by gemma-2-9b-it. We used text2dataset for translating with open-weight LLMs. By leveraging the fast LLM inference library vLLM, this tool enables the rapid translation of large English datasets into Japanese. Prompt The following is the prompt used for translation with Gemma. You are an… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/relaion2B-en-research-safe-japanese-translation.image1B<n<10B4 likes365 downloads1y agoHugging Face25Yehor /en-uk-translation-conversationstext10M<n<100M0 likes341 downloads1y agoHugging Face26DigitalUmuganda /monolingual_machine_translation_datatext100K<n<1M0 likes337 downloads3y agoHugging Face27LLaMAX /BenchMAX_Domain_Translation Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Domain_Translation is a dataset of BenchMAX, which evaluates the translation capability on specific domains. We collect the domain multi-way parallel data from other tasks in BenchMAX, such as math data, code data, etc. Each sample contains one… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Domain_Translation.texttranslation10K<n<100K0 likes321 downloads2y agoHugging Face28syeda-raisa /idiom_translation_finaltext10K<n<100K0 likes313 downloads3y agoHugging Face29bot-yaya /undl_zh2en_translation 联合国语料中文的机翻英文 rt,翻译工具用的argostranslate,因为对齐之前需要做一次机翻,所以这里上传了一份,以便后续pipeline使用,其它语言也一样。注意不要直接拿这份去练机翻模型,因为它们不是人翻的。 在机翻之前,已经用脚本洗掉了一部分制表噪声和分隔符,所用函数如下: def clean_paragraph(paragraph): lines = paragraph.split('\n') para = '' table = [] for line in lines: line = line.strip() # 表格线或其他分割线 if re.match(r'^\+[-=+]+\+|-+|=+|_+$', line): if not para.endswith('\n'): para += '\n' if len(table) > 0: para… See the full description on the dataset page: https://huggingface.co/datasets/bot-yaya/undl_zh2en_translation.text100K<n<1M0 likes295 downloads2y agoHugging Face30mfmezger /sandboxai_german_to_english_translations_seperatedtext1M<n<10M2 likes293 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.