CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01natgillin /translations-raw natgillin/translations-raw Frozen, canonical raw bitext consolidated from upstream alvations/mtdata-raw* snapshots (since deleted). This is the read-only source-of-truth for downstream quality-filtering pipelines. 31,663 parquet files (1566.8 GB) 49 language pairs under data/<src-tgt>/ Schema: 5 columns — see below Read-only for downstream pipelines. Do not delete or modify. Schema Each parquet has 5 columns: column type description source string… See the full description on the dataset page: https://huggingface.co/datasets/natgillin/translations-raw.text1M<n<10M5 likes16k downloads3mo agoHugging Face02zaibihassan /Quranic-Translation-Audio-Data Overview Quranic Translation Audio Data is a highly curated, standardized, and streaming-optimized multilingual audio dataset containing the complete recitation of translation audios and commentaries of the Holy Quran across 51 different translation directories. Every audio track has been meticulously converted from heavy .mp3 source files into the modern, high-fidelity Opus (.opus) format at a streaming-optimized bitrate of 32kbps. Alongside… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Translation-Audio-Data.audioaudio-to-audio1K<n<10K1 likes4.5k downloads3mo agoHugging Face03MaLA-LM /mala-bilingual-translation-corpus MaLA Corpus: Massive Language Adaptation Corpus This MaLA-LM/mala-bilingual-translation-corpus is the MaLA bilingual translation corpus, collected and processed from various sources. As a part of MaLA Corpus that aims to enhance massive language adaptation in many languages, it contains bilingual translation data (aka, parallel data and bitexts) in 2,500+ language pairs (500+ languages). Key statistics of all language pairs available at… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-bilingual-translation-corpus.texttranslation10B<n<100B8 likes3.6k downloads2mo agoHugging Face04ayymen /Weblate-Translations Dataset Card for Weblate Translations A dataset containing strings from projects hosted on Weblate and their translations into other languages. Please consider donating or contributing to Weblate if you find this dataset useful. To avoid rows with values like "None" and "N/A" being interpreted as missing values, pass the keep_default_na parameter like this: from datasets import load_dataset dataset = load_dataset("ayymen/Weblate-Translations", keep_default_na=False)… See the full description on the dataset page: https://huggingface.co/datasets/ayymen/Weblate-Translations.texttranslation10M<n<100M21 likes2.8k downloads2y agoHugging Face05aisingapore /NLG-Machine-Translationgated SEA Machine Translation SEA Machine Translation evaluates a model's ability to translate a document from a source language into a target language coherently and fluently. It is sampled from FLORES 200 for Burmese, Chinese, English, Indonesian, Khmer, Malay, Tamil, Thai, and Vietnamese, and NusaX for Indonesian, Javanese, and Sundanese. Supported Tasks and Leaderboards SEA Machine Translation is designed for evaluating chat or instruction-tuned large language models… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Machine-Translation.texttext-generation10K<n<100K2 likes2.8k downloads9mo agoHugging Face06HuggingFace-CN-community /translationimagen<1K61 likes1.9k downloads3y agoHugging Face07premio-ai /OpenSubtitles_Translations_Datasettext100M<n<1B0 likes1.2k downloads2y agoHugging Face08ayymen /Pontoon-Translations Dataset Card for Pontoon Translations This is a dataset containing strings from various Mozilla projects on Mozilla's Pontoon localization platform and their translations into more than 200 languages. Source strings are in English. To avoid rows with values like "None" and "N/A" being interpreted as missing values, pass the keep_default_na parameter like this: from datasets import load_dataset dataset = load_dataset("ayymen/Pontoon-Translations", keep_default_na=False)… See the full description on the dataset page: https://huggingface.co/datasets/ayymen/Pontoon-Translations.texttranslation1M<n<10M19 likes1.2k downloads3y agoHugging Face09Tamazight-NLP /Weblate-Translations Dataset Card for Weblate Translations A dataset containing strings from projects hosted on Weblate and their translations into other languages. Please consider donating or contributing to Weblate if you find this dataset useful. Dataset Details Dataset Description Curated by: Mohamed Aymane Farhi Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): Check the README YAML metadata… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/Weblate-Translations.texttranslation10M<n<100M4 likes1.1k downloads3mo agoHugging Face10yipyany /ted-translation-decisions-en-zh TED Translation Decision Dataset (EN–ZH 英-简中) 🎁🎁 DATASET UPDATED REGULARLY! COME BACK FOR NEW ENTRIES! 🎁🎁 🧩 Searchable Keywords translation, EN-ZH, bilingual, rationale, subtitle, human decisions,TED Talks, translation choices, linguistic annotation, cross-lingual, semantic nuance, translation rationale dataset, Chinese translation, English translation dataset, word-level translation, interpretability, translation pedagogy, translation teaching… See the full description on the dataset page: https://huggingface.co/datasets/yipyany/ted-translation-decisions-en-zh.tabulartranslationn<1K1 likes1k downloads8h agoHugging Face11soynade-research /Bambara-Speech-Translation-Data AfVoices-Translated (Bambara-English) This is a Bambara speech translation dataset, which is built on the African Next Voices (AfVoices) Bambara ASR corpus. It provides English translations for the human-corrected subset of the original collection, creating a parallel corpus for Bambara-English machine translation and speech-to-text tasks. Methodology We machine-translated the human-validated transcriptions from AfVoices using the Oolel-translator repository. Inference… See the full description on the dataset page: https://huggingface.co/datasets/soynade-research/Bambara-Speech-Translation-Data.audioautomatic-speech-recognition100K<n<1M1 likes1k downloads7mo agoHugging Face12zouhar /last-translation-benchmark Last Translation Benchmark Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. Standard benchmarks for machine translation evaluation are often either trivial (having few authentic mistakes) or unrealistic (overly synthetically contrived). Furthermore, automatic translation metrics become less reliable and reward-hacked as models get stronger, and their outputs are… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/last-translation-benchmark.texttranslation1K<n<10K59 likes920 downloads18d agoHugging Face13FrancophonIA /Translations_Hungarian_public_websites [!NOTE] Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/18982 Description A webcrawl of 14 different websites covering parallel corpora of Hungarian with Polish, Czech, Swedish, Finnish, French, German, Italian, English and Slovenian Citation Translations of Hungarian from public websites (2022). Version 1.0. [Dataset (Text corpus)]. Source: European Language Grid. https://live.european-language-grid.eu/catalogue/corpus/18982 translation0 likes878 downloads1y agoHugging Face14McGill-NLP /speech-translation-and-summarization English-Centric Multilingual Audio Dataset This dataset contains generated article and summary audio for English-centric multilingual directions. Each direction folder contains metadata JSONL files and corresponding audio files for few_shot and test splits. Included directions amharic_english / english_amharic arabic_english / english_arabic bengali_english / english_bengali chinese_simplified_english / english_chinese_simplified english_english french_english /… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/speech-translation-and-summarization.audioautomatic-speech-recognition10K<n<100K6 likes769 downloads1mo agoHugging Face15ltg /norsumm-nob-nno-translation Nynorsk-Bokmål translation pairs A multi-sentence parallel corpus of manual Nynorsk-Bokmål translations. These translations were extracted from the SamiaT/NorSumm dataset. You can read more about how the original dataset was created (including details about the manual translation process) in Benchmarking Abstractive Summarisation: A Dataset of Human-authored Summaries of Norwegian News Articles by Samia Touileb et al.. Contact David Samuel (davisamu@ifi.uio.no)… See the full description on the dataset page: https://huggingface.co/datasets/ltg/norsumm-nob-nno-translation.texttranslationn<1K1 likes750 downloads8mo agoHugging Face16theblackcat102 /instruction_translationsTranslation of Instruction datasettexttext-generation100K<n<1M5 likes717 downloads4y agoHugging Face17ilsp /ancient-modern_greek_translations Dataset Card for Ancient-Modern Greek translations The Ancient-Modern Greek translations dataset includes 100 sentences of Ancient Greek texts manually translated into Modern Greek. Original texts and translations have been extracted from the web sources cited below. Δημοσθένους, Ὑπὲρ τῆς Ῥοδίων ἐλευθερίας (ell: Δημοσθένους, Υπέρ της Ελευθερίας των Ροδίων; eng: Demosthenes, On the Liberty of the Rhodians). Μτφρ. Β.Η. Τσακατίκας. χ.χ. Λόγοι του Δημοσθένη. Γ' Ολυνθιακός, Υπέρ της… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/ancient-modern_greek_translations.texttranslationn<1K2 likes688 downloads2y agoHugging Face18tech1984 /swe_smith_back_translationBack translate the swe-smith data to get the problem statment following the R2E sylte prompt. More details in https://github.com/SWE-bench/SWE-smith/issues/127 text10K<n<100K2 likes659 downloads1y agoHugging Face19SultanR /AraMix-Translation-Scores AraMix-Translation-Scores AdaMLLab/AraMix (minhash_deduped subset, 178,883,241 rows) with a machine-translation-detection score added to every document. All original columns are preserved. Columns column type description id string unchanged from AraMix source string unchanged from AraMix text string unchanged from AraMix mmbert_quality_score float64 AraMix's original mmbert_score, renamed mmbert_translated_score float64 new —… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/AraMix-Translation-Scores.tabular100M<n<1B0 likes593 downloads2mo agoHugging Face20recursal /Europarl-Translation-Instruct Dataset Card for Europarl-Translation-Instruct Waifu to catch your attention. Dataset Details Dataset Description europarl-translation-instruct is a translation instruct dataset built from europarl data. Curated by: M8than Funded by: Recursal.ai Shared by: M8than Language(s) (NLP): English instruct (but various languages in) License: cc-by-sa-4.0 Dataset Sources Source Data: https://www.statmt.org/europarl/ (Transcript source) Processing… See the full description on the dataset page: https://huggingface.co/datasets/recursal/Europarl-Translation-Instruct.texttext-generation10M<n<100M4 likes548 downloads2y agoHugging Face21hpprc /enwiki-translations英語WikipediaをLLMを用いて英日翻訳したデータセットです。 collectionサブセットは翻訳元となった英語Wikipediaの文章、datasetサブセットは英日翻訳されたテキストペアと翻訳元にした事例のcollectionサブセットにおけるidを収載したものです。 なお、出力の利用に際しては、翻訳に使用した各モデルの出力に関するライセンス規約に従ってください。 Phi3.5 MoEを利用して作成されたデータについては、Wikipediaのライセンスにしたがい、CC-BY-SA 4.0で利用可能であるものとします。 text10M<n<100M0 likes489 downloads2y agoHugging Face22NuBerea /translation-analysisgated NuBerea Translation Verse Texts Verse-level texts of historical Bible translations (Clementine Vulgate, Luther Bible 1545, Matthew's Bible 1537). Part of the NuBerea curated corpus estate of biblical and historical texts. Attribution Upstream Data Sources Source License Clementine Vulgate, NOCR Public Domain Luther Bible 1545, NOCR Public Domain Matthew's Bible 1537, Textus Receptus Bibles Public Domain NuBerea project. Licensed… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/translation-analysis.tabularfeature-extraction100K<n<1M0 likes458 downloads9d agoHugging Face23semeru /code-code-translation-java-csharp Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/code-to-code-trans in Semeru CodeXGLUE -- Code2Code Translation Task Definition Code translation aims to migrate legacy software from one programming language in a platform toanother. In CodeXGLUE, given a piece of Java (C#) code, the task is to translate the code into C#… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-translation-java-csharp.text10K<n<100K2 likes451 downloads3y agoHugging Face24LLaMAX /BenchMAX_General_Translation Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_General_Translation is a dataset of BenchMAX, which evaluates the translation capability on the general domain. We collect parallel test data from Flore-200, TED-talk, and WMT24. Usage Run the following commands to generate… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_General_Translation.texttranslation100K<n<1M0 likes429 downloads1y agoHugging Face25kajuma /translation_source0 likes404 downloads10mo agoHugging Face26ackermannj /lior-translationgated2 likes399 downloads8d agoHugging Face27bot-yaya /undl_ru2en_translation Dataset Card for "undl_ru2en_translation" More Information needed text100K<n<1M0 likes395 downloads3y agoHugging Face28Aletheia-ng /african_languages_translationtext1M<n<10M1 likes384 downloads1y agoHugging Face29llm-jp /relaion2B-en-research-safe-japanese-translation relaion2B-en-research-safe-japanese-translation This dataset is the Japanese translation of the English subset of ReLAION-5B (laion/relaion2B-en-research-safe), translated by gemma-2-9b-it. We used text2dataset for translating with open-weight LLMs. By leveraging the fast LLM inference library vLLM, this tool enables the rapid translation of large English datasets into Japanese. Prompt The following is the prompt used for translation with Gemma. You are an… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/relaion2B-en-research-safe-japanese-translation.image1B<n<10B4 likes365 downloads1y agoHugging Face30Yehor /en-uk-translation-conversationstext10M<n<100M0 likes363 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.