CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ayymen /Pontoon-Translations Dataset Card for Pontoon Translations This is a dataset containing strings from various Mozilla projects on Mozilla's Pontoon localization platform and their translations into more than 200 languages. Source strings are in English. To avoid rows with values like "None" and "N/A" being interpreted as missing values, pass the keep_default_na parameter like this: from datasets import load_dataset dataset = load_dataset("ayymen/Pontoon-Translations", keep_default_na=False)… See the full description on the dataset page: https://huggingface.co/datasets/ayymen/Pontoon-Translations.texttranslation1M<n<10M19 likes1.2k downloads3y agoHugging Face02yipyany /ted-translation-decisions-en-zh TED Translation Decision Dataset (EN–ZH 英-简中) 🎁🎁 DATASET UPDATED REGULARLY! COME BACK FOR NEW ENTRIES! 🎁🎁 🧩 Searchable Keywords translation, EN-ZH, bilingual, rationale, subtitle, human decisions,TED Talks, translation choices, linguistic annotation, cross-lingual, semantic nuance, translation rationale dataset, Chinese translation, English translation dataset, word-level translation, interpretability, translation pedagogy, translation teaching… See the full description on the dataset page: https://huggingface.co/datasets/yipyany/ted-translation-decisions-en-zh.tabulartranslationn<1K1 likes1k downloads9h agoHugging Face03syeda-raisa /idiom_translation_finaltext10K<n<100K0 likes341 downloads3y agoHugging Face04Gsk068 /JP-TH_Literary_Translation_URL_Alignment_Index JP–TH Literary Translation URL Alignment Index This release provides a copyright-conscious metadata index and reproducibility package for a Japanese–Thai literary translation dataset associated with the study Context-Aware Prompting for Japanese–Thai Literary Translation in a Low-Resource Setting. Overview The release is designed to support reproducible academic research on Japanese–Thai literary machine translation, context-aware prompting, prompt engineering… See the full description on the dataset page: https://huggingface.co/datasets/Gsk068/JP-TH_Literary_Translation_URL_Alignment_Index.textn<1K0 likes234 downloads2mo agoHugging Face05abidlabs /test-translation-datasettextn<1K0 likes216 downloads5y agoHugging Face06Marcolini /cross-species-translational-alignment Cross-Species Translational Alignment — TG-GATEs + DrugMatrix × Tox21 Goal: build a training substrate for detecting subtle / pre-histopathological toxicity signatures in animal transcriptome data, with mechanism-of-toxicity labels attached. This directory contains the compound-level linkage layer: every compound that has rat in-vivo perturbation data cross-referenced to Tox21 mechanism assays via standardized chemical identifiers. Background — the hackathon Built… See the full description on the dataset page: https://huggingface.co/datasets/Marcolini/cross-species-translational-alignment.tabulartabular-classificationn<1K0 likes208 downloads3mo agoHugging Face07gplsi /ES-VA_translation_test Subtask (ES-VA_translation) of Phrases adaptability task This dataset was built from 200,000 sentences extracted from the Common Voice tool, an open resource that collects text contributions in various languages. These sentences were subjected to a rigorous filtering process, selecting only those with the greatest linguistic richness to ensure their usefulness in applications requiring language diversity and complexity. Subsequently, the selected sentences were translated from… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/ES-VA_translation_test.texttranslation1K<n<10K0 likes143 downloads9mo agoHugging Face08strongminsu /ko-en-structured-translations Korean–English Multistyle Parallel Corpus 한국어 사용자에게 익숙한 표현 기반의 다도메인·다문체 한–영 병렬 코퍼스 소개(Introduction) 저는 머신러닝, 인공지능 수업을 진행하는 강사입니다.Seq2Seq, Attention, Transformer 등 자연어처리(NLP) 수업을 진행하며한국 학습자에게 자연스럽고 익숙한 한–영 번역 데이터셋의 부족을 경험했습니다. 기존 공개 데이터셋은 도메인 다양성이 부족하거나 문체가 한국 사용자에게 자연스럽지 않거나 전반적으로 문장의 퀄리티가 매우 부족하여 학습한 번역 모델의 실제 성능이 기대만큼 나오지 않는 문제가 있었습니다. 이 문제를 해결하기 위해, 딥러닝 강사로서 langchain을 사용하여 직접 고품질 병렬 데이터를 자동으로 생성·정제하여 구성한 데이터셋입니다. 한국어 사용자에게 익숙한 표현을 중심으로 다양한 문체, 문장 구조를… See the full description on the dataset page: https://huggingface.co/datasets/strongminsu/ko-en-structured-translations.texttranslation1K<n<10K17 likes136 downloads10mo agoHugging Face09Tamazight-NLP /Pontoon-Translations Pontoon Translations Amazigh subset of Pontoon Translations. texttranslation10K<n<100K2 likes126 downloads3y agoHugging Face10traintogpb /aihub-koen-translation-integrated-large-10m AI Hub Ko-En Translation Dataset (Integrated) AI Hub의 한-영 번역 관련 데이터셋 8개를 병합한 자료입니다. 병합 시 총 데이터 개수는 10,416,509개 이며, train / validation / test는 8:1:1 비율로 분할되었습니다. base-10m: 병합 데이터 100% 사용, 총 10,416,509개 mini-1m: 병합 데이터 10% 사용 (base-10m의 각 세트 내에서 10% 임의 선택), 총 1,041,651개 tiny-100k: 병합 데이터 1% 사용 (base-10m의 각 세트 내에서 1% 임의 선택), 총 104,165개 Subsets 활용한 데이터셋 목록은 다음과 같으며, 데이터셋 이름 옆 번호는 aihubshell에서의 datasetkey입니다. 전문분야 한영 말뭉치 (111) 총 개수: 1,350,000 중복 제거 후 개수: 1,350,000 사용 칼럼:… See the full description on the dataset page: https://huggingface.co/datasets/traintogpb/aihub-koen-translation-integrated-large-10m.texttranslation10M<n<100M18 likes93 downloads3y agoHugging Face11rhyliieee /tagalog-filipino-english-translationThis dataset is a Tagalog-English translation data. It is a compiled comma-separated values dataset from different existing HuggingFace and External dataset. Here are the collected and compiled data: saillab/alpaca_tamil_taco DIBT/MPEP_FILIPINO Nag, S., Ma, S., Ntalli, A., & Dulay, K. M. (2024, June 10). TalkTogether. https://doi.org/10.17605/OSF.IO/3ZDFN texttranslation100K<n<1M6 likes90 downloads2y agoHugging Face12jaio98 /basque_dialect_machine_translationtext100K<n<1M0 likes86 downloads3mo agoHugging Face13liechticonsulting /swiss-legal-translation Swiss Legal Translation Dataset A large-scale parallel corpus of Swiss legal texts in German, French, and Italian, extracted from official government sources. Dataset Overview Metric Value Total articles 62,594 Languages German, French, Italian All 3 languages 28,000 articles (44.7%) At least 2 languages 62,594 articles (100%) Source laws 743+ unique laws File size ~88 MB Language Coverage by Source Source Articles DE FR IT… See the full description on the dataset page: https://huggingface.co/datasets/liechticonsulting/swiss-legal-translation.text10K<n<100K2 likes83 downloads4mo agoHugging Face14Ehtisham1328 /urdu-idioms-with-english-translationtexttranslation1K<n<10K5 likes79 downloads3y agoHugging Face15MaroneAI /Wolof-to-French_Translation-Dataset Dataset Wolof ↔ Français 🧩 Présentation Ce dataset contient plus de 30 000 paires phrase Wolof – phrase Française.Chaque ligne est structurée comme suit : Wolof (input) Français (target) Phrase en Wolof Phrase correspondante en Français Il a été conçu pour la traduction automatique et les tâches de NLP impliquant le Wolof et le Français. 📚 Provenance et nettoyage Le dataset a été créé en compilant différentes sources accessibles… See the full description on the dataset page: https://huggingface.co/datasets/MaroneAI/Wolof-to-French_Translation-Dataset.texttranslation10K<n<100K3 likes74 downloads1y agoHugging Face16billingsmoore /tibetan-to-spanish-translation-datasetThis dataset consists of three columns, the first of which is a sentence or phrase in Tibetan, the second is the phonetic transliteration of the Tibetan, and the third is the Spanish translation of the Tibetan. The dataset was scraped from Lotsawa House and is released under the same license as the texts from which it is sourced. The dataset is part of the larger MLotsawa project, the code repo for which can be found here. texttranslation10K<n<100K1 likes65 downloads2y agoHugging Face17allandclive /Luganda_Sci-Math-Bio_Translations Luganda Sci-Math-Bio Translations This dataset contains Luganda and English translations of biologicial, mathematical and scientific terms texttranslation1K<n<10K3 likes64 downloads3y agoHugging Face18lordjia /Cantonese_English_Translation Cantonese_English_Translation Overview | 總括 This dataset provides parallel text translations between Cantonese and English, suitable for research and development in natural language processing and machine translation. | 呢個資料庫提供廣東話同英文嘅對應翻譯,啱晒用嚟做語言處理同機器翻譯嘅研究同開發。 Dataset Structure | 資料組織 english_cantonese_translation.csv: Contains two fields: "english" and "cantonese". | 有兩個位: "english" 同 "cantonese"。 Usage Example | 用法例子 import pandas as pd # Load… See the full description on the dataset page: https://huggingface.co/datasets/lordjia/Cantonese_English_Translation.texttranslation100K<n<1M8 likes62 downloads2y agoHugging Face19Amani27 /massive_translation_dataset Dataset Card for Massive Dataset for Translation Dataset Summary This dataset is derived from AmazonScience/MASSIVE dataset for translation task purpose. Supported Tasks and Leaderboards Translation Languages English (en_US) German (de_DE) Hindi (hi_IN) Spanish (es_ES) French (fr_FR) Italian (it_IT) Arabic (ar_SA) Dutch (nl_NL) Japanese (ja_JP) Portugese (pt_PT) texttranslation10K<n<100K10 likes59 downloads3y agoHugging Face20abdelhaqueidali /Amazigh-Quran-Translation-Jouhadi Dataset Card: Tamazight (Tifinagh) Quran Translation - Lahoucine Jouhadi This dataset provides a digitized, partial translation of the meanings of the Holy Quran into Amazigh (Tachelhit) using the Neo-Tifinagh script. The content is based on the full translation work of Lahoucine Jouhadi (Lhocine Jouhadi Baamrani) based on Warsh recitation used in Morocco. Original Sources & References Author's Website - Down currently: Jouhadi Lahoussine Publications… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Amazigh-Quran-Translation-Jouhadi.tabular1K<n<10K0 likes54 downloads26d agoHugging Face21Charif-Ayfarah /English-to-Afar-language-translation Author Created by Charif Ayfarah. Contact: afbarit@gmail.com License Licensed under CC BY 4.0. You are free to use, modify, and distribute this dataset, including for commercial purposes, as long as you give appropriate credit. textn<1K2 likes52 downloads6mo agoHugging Face22pythainlp /thai-local-language-translation-dataset Thai Local Language Translation Dataset Thai Local Language Translation Dataset is a translation dataset for translate Thai Local Language to Thai Central Language. We create the dataset from Thai Dialect Corpus (Thai dialects ASR corpus). We select train set only from Thai Dialect Corpus. The dataset support Khummuang, Korat, and Pattani. Reference Suwanbandit, A., Naowarat, B., Sangpetch, O., Chuangsuwanich, E. (2023) Thai Dialect Corpus and Transfer-based Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-local-language-translation-dataset.texttranslation10K<n<100K4 likes51 downloads2y agoHugging Face23traintogpb /aihub-koen-translation-integrated-small-100k AI Hub Ko-En Translation Dataset (Integrated) AI Hub의 한-영 번역 관련 데이터셋 8개를 병합한 자료입니다. 병합 시 총 데이터 개수는 10,416,509개 이며, train / validation / test는 8:1:1 비율로 분할되었습니다. base-10m: 병합 데이터 100% 사용, 총 10,416,509개 mini-1m: 병합 데이터 10% 사용 (base-10m의 각 세트 내에서 10% 임의 선택), 총 1,041,651개 tiny-100k: 병합 데이터 1% 사용 (base-10m의 각 세트 내에서 1% 임의 선택), 총 104,165개 Subsets 활용한 데이터셋 목록은 다음과 같으며, 데이터셋 이름 옆 번호는 aihubshell에서의 datasetkey입니다. 전문분야 한영 말뭉치 (111) 총 개수: 1,350,000 중복 제거 후 개수: 1,350,000 사용 칼럼:… See the full description on the dataset page: https://huggingface.co/datasets/traintogpb/aihub-koen-translation-integrated-small-100k.texttranslation100K<n<1M3 likes48 downloads3y agoHugging Face24traintogpb /aihub-koen-translation-integrated-base-1m AI Hub Ko-En Translation Dataset (Integrated) AI Hub의 한-영 번역 관련 데이터셋 8개를 병합한 자료입니다. 병합 시 총 데이터 개수는 10,416,509개 이며, train / validation / test는 8:1:1 비율로 분할되었습니다. base-10m: 병합 데이터 100% 사용, 총 10,416,509개 mini-1m: 병합 데이터 10% 사용 (base-10m의 각 세트 내에서 10% 임의 선택), 총 1,041,651개 tiny-100k: 병합 데이터 1% 사용 (base-10m의 각 세트 내에서 1% 임의 선택), 총 104,165개 Subsets 활용한 데이터셋 목록은 다음과 같으며, 데이터셋 이름 옆 번호는 aihubshell에서의 datasetkey입니다. 전문분야 한영 말뭉치 (111) 총 개수: 1,350,000 중복 제거 후 개수: 1,350,000 사용 칼럼:… See the full description on the dataset page: https://huggingface.co/datasets/traintogpb/aihub-koen-translation-integrated-base-1m.texttranslation1M<n<10M3 likes45 downloads3y agoHugging Face25Huyisbeee /ViKm-Translation-Task ViKm-Trans A high-quality synthetic Vietnamese–Khmer parallel corpus. Overview ViKm-Trans is a synthetic parallel corpus for Vietnamese ↔ Khmer machine translation. Due to the scarcity of publicly available Vietnamese–Khmer parallel data, we propose a synthetic data generation framework that leverages abundant Vietnamese monolingual corpora together with large language models to construct high-quality parallel sentence pairs. The dataset was introduced in… See the full description on the dataset page: https://huggingface.co/datasets/Huyisbeee/ViKm-Translation-Task.texttranslation10K<n<100K0 likes44 downloads3mo agoHugging Face26traintogpb /aihub-kozh-translation-integrated-large-5.9m AI Hub Ko-Zh Translation Dataset (Integrated) AI Hub의 한-중 번역 관련 데이터셋 10개를 병합한 자료입니다. 병합 시 총 데이터 개수는 5,934,596개이며, 이중 10,000개의 validation set와 2,000개의 test set가 분리되어 모든 데이터 사이즈(large-5.9m, base-1m, small-100k)에서 동일하게 사용됩니다. large-5.9m (train): 병합 데이터 100% 사용; 총 5,922,596개 base-1m (train): 병합 데이터 중 1M개 사용; 총 1,000,000개 small-100k (train): 병합 데이터 중 100K개 사용; 총 100,000개 Subsets Name Total Size Chinese Size (Utilized Only) URL Datasetkey (AIHub) 한국어-중국어 번역… See the full description on the dataset page: https://huggingface.co/datasets/traintogpb/aihub-kozh-translation-integrated-large-5.9m.texttranslation1M<n<10M1 likes39 downloads2y agoHugging Face27lunovian /vietnamese-nom-poetry-translationtexttranslationn<1K1 likes37 downloads2y agoHugging Face28Programmer-RD-AI /sinhala-english-singlish-translation Sinhala–English–Singlish Translation Dataset A parallel corpus of Sinhala sentences, their English translations, and romanized Sinhala (“Singlish”) transliterations. 📋 Table of Contents Dataset Overview Installation Quick Start Dataset Structure Usage Examples Citation License Credits Dataset Overview Description: 34,500 aligned triplets of Sinhala (native script) English (human translation) Singlish (romanized Sinhala)… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sinhala-english-singlish-translation.texttranslation10K<n<100K3 likes36 downloads1y agoHugging Face29zhangtaolab /cross_species_leaf_absolute_translationtabular10K<n<100K0 likes36 downloads3mo agoHugging Face30sellersew /carrot-engine-normalization-translation-v2text10M<n<100M1 likes35 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.