CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yipyany /ted-translation-decisions-en-zh TED Translation Decision Dataset (EN–ZH 英-简中) 🎁🎁 DATASET UPDATED REGULARLY! COME BACK FOR NEW ENTRIES! 🎁🎁 🧩 Searchable Keywords translation, EN-ZH, bilingual, rationale, subtitle, human decisions,TED Talks, translation choices, linguistic annotation, cross-lingual, semantic nuance, translation rationale dataset, Chinese translation, English translation dataset, word-level translation, interpretability, translation pedagogy, translation teaching… See the full description on the dataset page: https://huggingface.co/datasets/yipyany/ted-translation-decisions-en-zh.tabulartranslationn<1K1 likes1.1k downloads15h agoHugging Face02Marcolini /cross-species-translational-alignment Cross-Species Translational Alignment — TG-GATEs + DrugMatrix × Tox21 Goal: build a training substrate for detecting subtle / pre-histopathological toxicity signatures in animal transcriptome data, with mechanism-of-toxicity labels attached. This directory contains the compound-level linkage layer: every compound that has rat in-vivo perturbation data cross-referenced to Tox21 mechanism assays via standardized chemical identifiers. Background — the hackathon Built… See the full description on the dataset page: https://huggingface.co/datasets/Marcolini/cross-species-translational-alignment.tabulartabular-classificationn<1K0 likes205 downloads3mo agoHugging Face03agentlans /tatoeba-english-translations Tatoeba English Translation Dataset Dataset Summary This dataset is derived from the Tatoeba database, focusing on English sentences and their translations. It includes assessments of English sentences using text quality, sentiment, and readability models. The dataset is designed for tasks related to multilingual text quality, readability, and sentiment analysis. Supported Tasks and Leaderboards Quality Assessment Readability Prediction Sentiment Analysis… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/tatoeba-english-translations.tabulartext-classification1M<n<10M2 likes88 downloads2y agoHugging Face04abdelhaqueidali /Amazigh-Quran-Translation-Jouhadi Dataset Card: Tamazight (Tifinagh) Quran Translation - Lahoucine Jouhadi This dataset provides a digitized, partial translation of the meanings of the Holy Quran into Amazigh (Tachelhit) using the Neo-Tifinagh script. The content is based on the full translation work of Lahoucine Jouhadi (Lhocine Jouhadi Baamrani) based on Warsh recitation used in Morocco. Original Sources & References Author's Website - Down currently: Jouhadi Lahoussine Publications… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Amazigh-Quran-Translation-Jouhadi.tabular1K<n<10K0 likes54 downloads28d agoHugging Face05zhangtaolab /cross_species_leaf_absolute_translationtabular10K<n<100K0 likes39 downloads3mo agoHugging Face06zhangtaolab /cross_species_leaf_on_off_translationtabular10K<n<100K0 likes27 downloads3mo agoHugging Face07ChaoticEconomist /EnglishtoFrench-Translation-Dataset English–French Translation Dataset (SFT / LoRA Ready) A clean, structured dataset of 50,000 English–French sentence pairs designed for supervised fine-tuning (SFT) of large language models, LoRA adapters, and general machine translation tasks. Overview Property Value Language pair English → French Total rows 50,000 Train split 45,000 (90%) Validation split 2,500 (5%) Test split 2,500 (5%) Format CSV (Alpaca-style prompt format) License CC… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/EnglishtoFrench-Translation-Dataset.tabulartranslation10K<n<100K0 likes19 downloads4mo agoHugging Face08Calibration-Translation /Calibration-translation-human-eval Translation Evaluation Dataset: Tower vs Calibration This dataset compares translations generated by two models ("Tower-system" and "Calibration") along with human ratings. tabularn<1K0 likes17 downloads1y agoHugging Face09NAMAA-Space /ASCAT-Arabic-Scientific-Translation ASCAT: Arabic Scientific Corpus for Advanced Translation ASCAT (Arabic Scientific Corpus for Advanced Translation) is a high-quality English–Arabic parallel corpus of full scientific abstracts designed for rigorous evaluation and training of domain-specific machine translation (MT) systems. Unlike existing Arabic–English corpora that rely on short sentences or narrow domains, ASCAT targets long-form scientific abstracts validated through a multi-engine translation and expert… See the full description on the dataset page: https://huggingface.co/datasets/NAMAA-Space/ASCAT-Arabic-Scientific-Translation.tabulartranslationn<1K1 likes16 downloads6mo agoHugging Face10ychen /english-darija-arabizi-translationtabular1K<n<10K0 likes14 downloads2y agoHugging Face11sltAI /crowdsourced-text-to-sign-language-rule-based-translation-corpus Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/sltAI/crowdsourced-text-to-sign-language-rule-based-translation-corpus.tabular1K<n<10K0 likes13 downloads6mo agoHugging Face12misclassified /meps_speeches_with_translation.csvtabular10K<n<100K0 likes12 downloads3y agoHugging Face13Junhoee /Jeju-Standard-Translationtabular100K<n<1M0 likes10 downloads2y agoHugging Face14JerrySweeney /Irish_Translations Irish_Focloir This dataset contains English phrases along with their Irish translations. Dataset Structure The dataset contains the following fields: english: The English phrase. irish: The Irish translation of the phrase. Example Here is an example of the data structure: english,irish "Hello","Dia dhuit" "Goodbye","Slán" "Thank you","Go raibh maith agat" tabulartranslationn<1K0 likes6 downloads2y agoHugging Face15DebasishDhal99 /wat2025-translation-collectiontabular1K<n<10K0 likes6 downloads11mo agoHugging Face16bokatiq /referenceless_machine_translation_evaluationBengali is a low resource language in natural language processing (NLP), with dialects like Sylheti, Chittagong, and Barisal being even more underrepresented. To address this, ONUBAD introduced a parallel corpus translating these dialects into Standard Bangla and English using expert translators, providing 1,540 words, 130 clauses, and 980 sentences per dialect. We focused on the Sylheti-English pair and adapted the dataset for LLM-based machine translation (MT) evaluation. We extracted the… See the full description on the dataset page: https://huggingface.co/datasets/bokatiq/referenceless_machine_translation_evaluation.tabulartranslation1K<n<10K0 likes6 downloads10mo agoHugging Face17ClarusC64 /legal-scientific-evidence-translation-coherence-v0.1What this dataset is You receive scientific finding court translation uncertainty bounds method limits overstatement signals You decide Does the translation preserve the limits of the science Answer coherent or incoherent Why this matters When translation drifts juries misread certainty appeals rise convictions or verdicts destabilise tabulartext-classificationn<1K0 likes4 downloads7mo agoHugging Face18vidula123 /translation_2tabularn<1K0 likes3 downloads2y agoHugging Face19OdiaGenAIdata /wat24_text_to_text_translationtabular100K<n<1M0 likes3 downloads2y agoHugging Face20Anonymous-Account /Calibration-translation-human-eval Translation Evaluation Dataset: Tower vs Calibration This dataset compares translations generated by two models ("Tower-system" and "Calibration") along with human ratings. tabularn<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.