CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agentlans /translation-quality Multilingual Translation Quality Dataset This dataset provides multilingual text chunks translated into English, accompanied by automated quality evaluations generated by multiple large language models. Dataset Details Source Data: agentlans/HuggingFaceFW-finetranslations-100-languages-sample Target Language: English Content: Multilingual chunks mapped to their English translations alongside automated judge scores. Evaluation Methodology The… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/translation-quality.tabulartranslation100K<n<1M0 likes98 downloads20d agoHugging Face02ic-org /wellbeing-in-translation Wellbeing in Translation Raw outputs and translated materials for Does AI Wellbeing Survive Translation? We test whether the unchanged CAIS 1-7 self-report battery measures the same positive-minus-negative gap after translation. Paper · Code · Source instrument Headline result Language sensitivity is specific to the model-battery pair. Model Gap spread across 7 languages English rank English stimulus / local battery Local stimulus / English battery… See the full description on the dataset page: https://huggingface.co/datasets/ic-org/wellbeing-in-translation.image10K<n<100K0 likes85 downloads1mo agoHugging Face03melanieyes /adaption-piguard-translation-handoff This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. adaption-piguard_translation_handoff This dataset contains 100 English prompts curated for translation guardrail research, comprising a balanced mix of 50 benign instructions and 50 prompt injection attempts. The samples include diverse content such as jailbreak personas, requests for harmful actions like spyware installation, and complex instruction overrides. Each entry is… See the full description on the dataset page: https://huggingface.co/datasets/melanieyes/adaption-piguard-translation-handoff.tabular1K<n<10K1 likes62 downloads9d agoHugging Face04CGICAI /cherokee-english-translation Cherokee–English Parallel Corpus (Archivist Project) A curated Cherokee (ᏣᎳᎩ / Tsalagi) ↔ English parallel corpus for machine translation, assembled from public sources, deduplicated, benchmark-decontaminated, and conflict-cleaned. Built to train and evaluate English→Cherokee translation models for one of the most endangered languages in North America. Files File Rows Purpose train_en2chr_v2.jsonl 138,307 Flagship training set. English→Cherokee SFT… See the full description on the dataset page: https://huggingface.co/datasets/CGICAI/cherokee-english-translation.tabulartranslation100K<n<1M0 likes57 downloads2mo agoHugging Face05freococo /tipitaka_myanmar_translation_books Myanmar Tipitaka Translation (60 Books) This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga. The texts have been converted into a clean, structured JSONL format, suitable for: Natural Language Processing (NLP) LLM Training & Fine-tuning Digital Humanities Research Dhamma Study Applications 📊 Dataset Statistics Total Books: 60 Total Content Lines: 194… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_myanmar_translation_books.tabulartext-generation100K<n<1M0 likes45 downloads8mo agoHugging Face06Voider22 /bhagavad-gita-verses-sanskrit-translations Bhagavad Gita – Sanskrit, Transliteration & Multi-Commentary Dataset A complete dataset of all 700 verses of the Bhagavad Gita, sourced directly from the open-source VedicScriptures API (MIT-licensed).This dataset includes: 📜 Original Sanskrit slokas 🔡 IAST transliteration 🌐 Multiple English & Hindi translations 🧠 Traditional commentaries from many teachers 🔢 Structured metadata (chapter, verse, IDs, authors) This dataset is ideal for NLP, LLM fine-tuning, translation… See the full description on the dataset page: https://huggingface.co/datasets/Voider22/bhagavad-gita-verses-sanskrit-translations.tabularn<1K2 likes44 downloads10mo agoHugging Face07nizarun /English_Arabic_Translation_Pairs English · العربية English→Arabic Technical & Reasoning Translation Dataset High-quality English → Modern Standard Arabic translation pairs focused on native-English educational, scientific, and reasoning content. English source text is drawn from real corpora (FineWeb-Edu and three NVIDIA reasoning datasets); Arabic translations are produced by DeepSeek-v4-flash under a strict translation-only prompt that preserves notation, numbers, formulas, code, and citations. Pairs… See the full description on the dataset page: https://huggingface.co/datasets/nizarun/English_Arabic_Translation_Pairs.tabulartranslation100K<n<1M0 likes38 downloads2mo agoHugging Face08OpenFormosa /zh_translation_benchmark zh_translation_benchmark zh_translation_benchmark is a 2,000-example synthetic benchmark for evaluating whether a translation or rewriting model can produce natural Taiwan Traditional Chinese (zh-TW) from English, Mainland Chinese, Hong Kong Traditional Chinese, Cantonese-style written Chinese, or code-mixed English/Chinese documents. The dataset now exposes a single Hugging Face subset/config: full. It is not split into dev and test files. The full 2,000 rows are loaded… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/zh_translation_benchmark.tabulartranslation1K<n<10K0 likes24 downloads4mo agoHugging Face09agentlans /en-translations Multilingual Parallel Sentences with Semantic Similarity Scores and Quality Metrics This dataset is a diverse collection of parallel sentences in English and various other languages, sourced from multiple high-quality datasets. Each sentence pair includes a semantic similarity score calculated using the Language-agnostic BERT Sentence Embedding (LaBSE) model, along with additional quality metrics. Supported Tasks This dataset supports: Machine Translation… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-translations.tabulartranslation100K<n<1M0 likes20 downloads2y agoHugging Face10Porsera /lmsys_persian_translationgatedtabular10K<n<100K0 likes15 downloads16d agoHugging Face11c123ian /dpo_irish_eng_translationsThis is a test for my DPO dataset for Irish ENglish trasnlslations, raw data origin : https://www.gaois.ie/en/corpora/parallel?Query=Apple&Language=en&SearchMode=exact&PerPage=50, used COMETXL refrernce free maodel Unbabel/wmt23-cometkiwi-da-xl (which has been trained to asses Irish) to score accepted/rejected. Used GPT4 to generate translations to compare with human stranslations of Irish legislation (which has to have a Irisng/English copy by law) tabularn<1K0 likes13 downloads2y agoHugging Face12Arabic-Clip-Archive /ImageCaptions-7M-Translations-Arabic-subset-150000image100K<n<1M0 likes11 downloads3y agoHugging Face13c123ian /Irish_English_Translation Irish English graded Translations Data collected in order to fine-tune a Irish-English LLM, see blog post for more details. See here for the next phase, preference dataset formated (for DPO) here Data Sources: translated_gaois_graded.jsonl Parallel English-Irish corpus of legislation collected by Gaois. This corpus contains high-quality, human-translated paragraph pairs, making it a valuable resource. translated_tatoeba_graded.jsonl One draw-back is uses a lot of… See the full description on the dataset page: https://huggingface.co/datasets/c123ian/Irish_English_Translation.tabular10K<n<100K0 likes8 downloads2y agoHugging Face14welyjesch /bombo_hil_eng_raw_translationstabular10K<n<100K0 likes8 downloads4mo agoHugging Face15NorthernTribe-Research /maasai-translation-corpusgated Maasai-English Translation Corpus Parallel English↔Maasai translation pairs for low-resource MT, language preservation, and culturally grounded tooling. Overview Total pairs: 9,910 Splits: 8,434 train / 738 valid / 738 test Directions: 4,955 en→mas and 4,955 mas→en Quality tiers: 8,444 gold and 1,466 silver Main sources: 8,444 Bible-derived pairs, 680 cultural manual pairs, 70 knowledge-driven cultural pairs, 132 public-domain Hollis proverb pairs, 504 public-domain… See the full description on the dataset page: https://huggingface.co/datasets/NorthernTribe-Research/maasai-translation-corpus.tabulartranslation10K<n<100K0 likes3 downloads6mo agoHugging Face16MohammadKhodadad /chem-machine-translation-samplestabularn<1K0 likes2 downloads4mo agoHugging Face17qducnguyen /ultrafeedback_translation_refinedgatedtabular10K<n<100K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.