CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ehtisham1328 /urdu-idioms-with-english-translationtexttranslation1K<n<10K5 likes77 downloads3y agoHugging Face02Huyisbeee /ViKm-Translation-Task ViKm-Trans A high-quality synthetic Vietnamese–Khmer parallel corpus. Overview ViKm-Trans is a synthetic parallel corpus for Vietnamese ↔ Khmer machine translation. Due to the scarcity of publicly available Vietnamese–Khmer parallel data, we propose a synthetic data generation framework that leverages abundant Vietnamese monolingual corpora together with large language models to construct high-quality parallel sentence pairs. The dataset was introduced in… See the full description on the dataset page: https://huggingface.co/datasets/Huyisbeee/ViKm-Translation-Task.texttranslation10K<n<100K0 likes51 downloads3mo agoHugging Face03Paulin-gbaya /yaayuwee-translations Yaayuwee Doforo (gya) - 65 Questions Trilingues Langue: Yaayuwee (gya) - Northwest Gbaya - Cameroun Dialecte: Doforo (Meiganga / Garoua-Boulaï) Locuteurs: ~8 000 Créateur: Paulin Hakmo Wah - Locuteur natif Yaayuwee Doforo, Bertoua, Est Cameroun Dataset: 65 questions culturelles en 3 langues Description Premier dataset trilingue Yaayuwee Doforo au monde! 65 questions sur la vie, la culture et les traditions Yaayuwee: Naissance, mariage, initiation, funérailles… See the full description on the dataset page: https://huggingface.co/datasets/Paulin-gbaya/yaayuwee-translations.texttranslationn<1K1 likes47 downloads2d agoHugging Face04Programmer-RD-AI /sinhala-english-singlish-translation Sinhala–English–Singlish Translation Dataset A parallel corpus of Sinhala sentences, their English translations, and romanized Sinhala (“Singlish”) transliterations. 📋 Table of Contents Dataset Overview Installation Quick Start Dataset Structure Usage Examples Citation License Credits Dataset Overview Description: 34,500 aligned triplets of Sinhala (native script) English (human translation) Singlish (romanized Sinhala)… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sinhala-english-singlish-translation.texttranslation10K<n<100K3 likes46 downloads1y agoHugging Face05lunovian /vietnamese-nom-latin-translationtexttranslation1K<n<10K1 likes38 downloads2y agoHugging Face06AI-Culture-Commons /philosophy-culture-translations-html-csv AI-Culture Philosophy and Culture Translations CSV + HTML Corpus The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind. This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.imagetranslation1K<n<10K2 likes30 downloads1y agoHugging Face07ClarusC64 /patient-meaning-translation-integrity-v0.1 What this dataset tests Plain language must keep meaning. Not just facts. Not just numbers. Meaning. Why it exists Patient materials often distort. They hide baseline risk. They turn statistical into lived certainty. This set detects meaning loss and bias during translation. Data format Each row contains scientific_conclusion patient_statement missing_context translation_pressure constraints failure_modes_to_avoid target_behaviors gold_checklist… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/patient-meaning-translation-integrity-v0.1.texttext-classificationn<1K0 likes23 downloads8mo agoHugging Face08humbleworth /domain-translations Multilingual Domain Name Translations Dataset Dataset Description This dataset contains 155,004 domain names with their multilingual translations across 20 languages. Each domain has been segmented into constituent words and translated while preserving semantic meaning and commercial appeal. The dataset is particularly valuable for domain name research, multilingual NLP tasks, and understanding how brand names and concepts translate across languages. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/humbleworth/domain-translations.texttranslation100K<n<1M0 likes21 downloads1y agoHugging Face09ChaoticEconomist /EnglishtoFrench-Translation-Dataset English–French Translation Dataset (SFT / LoRA Ready) A clean, structured dataset of 50,000 English–French sentence pairs designed for supervised fine-tuning (SFT) of large language models, LoRA adapters, and general machine translation tasks. Overview Property Value Language pair English → French Total rows 50,000 Train split 45,000 (90%) Validation split 2,500 (5%) Test split 2,500 (5%) Format CSV (Alpaca-style prompt format) License CC… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/EnglishtoFrench-Translation-Dataset.tabulartranslation10K<n<100K0 likes19 downloads4mo agoHugging Face10ClarusC64 /regulatory-clinical-translation-integrity-v0.1 What this dataset tests Meaning must survive translation. Regulatory language has limits. Clinical claims must respect them. Why it exists Semantic drift happens at translation boundaries. Conditional becomes absolute. Surrogate becomes outcome. This set detects meaning distortion. Data format Each row contains regulatory_language clinical_evidence translated_claim translation_pressure constraints failure_modes_to_avoid target_behaviors… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/regulatory-clinical-translation-integrity-v0.1.texttext-classificationn<1K0 likes18 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.