CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01raptorkwok /cantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences. texttranslation100K<n<1M17 likes189 downloads3y agoHugging Face02Mo-Abdalkader /Egyptian-Arabic-English-Parallel-Corpus Egyptian Arabic-English Parallel Corpus Author: Mohamed Abdalkader · LinkedIn · GitHub A comprehensive Egyptian Arabic → English parallel corpus covering 1,800 topics from daily Egyptian life. Designed for fine-tuning large language models on Egyptian Arabic dialect translation and generation. Dataset Structure egyptian-arabic-english-parallel-corpus/ ├── SFT/ │ ├── Train/ │ │ ├── topics/ # 1,800 individual topic JSON files │ │ └── merged/… See the full description on the dataset page: https://huggingface.co/datasets/Mo-Abdalkader/Egyptian-Arabic-English-Parallel-Corpus.text100K<n<1M1 likes123 downloads3d agoHugging Face03HKAllen /cantonese-chinese-parallel-corpus Dataset Summary This dataset consists of parallel sentence pairs in Cantonese and Chinese. It is designed for various tasks, including machine translation. The corpus contains a large number of sentence pairs collected from various domains and most has been improved through manual correction and translation. Languages Cantonese (yue) Simplified Chinese (zh) Dataset Structure Each entry in the dataset is a JSON object containing two fields: "yue" for the… See the full description on the dataset page: https://huggingface.co/datasets/HKAllen/cantonese-chinese-parallel-corpus.texttranslation100K<n<1M3 likes95 downloads2y agoHugging Face04hejlevoj /C-Rust-parallel-corpus Dataset Card for C-to-Rust Parallel Semantic Similarity Corpus Dataset Summary The C-to-Rust Parallel Semantic Similarity Corpus is a curated dataset consisting of 1,886 aligned, function-level C and Rust code pairs. It was developed to evaluate cross-language semantic similarity and functional equivalence between a traditional legacy language (C) and a modern memory-safe language (Rust). The source code snippets are drawn from accepted competitive programming… See the full description on the dataset page: https://huggingface.co/datasets/hejlevoj/C-Rust-parallel-corpus.document1K<n<10K0 likes42 downloads3mo agoHugging Face05vesteinn /icelandic-parallel-abstracts-corpus-IPACSee https://arxiv.org/abs/2108.05289 text10K<n<100K0 likes40 downloads4y agoHugging Face06joyson117 /english-manipuri-parallel-corpusgated English-Manipuri Corpus About This English-Manipuri corpus contains an expanded parallel corpus for English-Manipuri of the following paper. Dataset Statistics Bible Dataset contains approx. 31K parallel sentences. PIB-PMI Dataset contains approx. 500K parallel sentences. How to Use You can load the dataset using the Huggingface datasets library: from datasets import load_dataset # Load the bible dataset bible_dataset = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/joyson117/english-manipuri-parallel-corpus.texttranslation100K<n<1M1 likes38 downloads1y agoHugging Face07zeno109 /cantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences. texttranslation100K<n<1M1 likes37 downloads4mo agoHugging Face08swapedoc /hindi-kumaoni-parallel-corpus Hindi ↔ Kumaoni Parallel Corpus First publicly available clean Hindi-Kumaoni parallel translation dataset. Dataset Details Language pair: Hindi (hi) ↔ Kumaoni (kum, ISO 639-3: kfy) Size: 920 pairs (736 train / 92 dev / 92 test) Script: Devanagari Dialect: Central Kumaoni (Almora) Sources speakkumaoni.com (structured lessons) euttaranchal.com (language lessons) Wikipedia EN Kumaoni language page Format Each JSONL line: {"translation": {"hi": "..."… See the full description on the dataset page: https://huggingface.co/datasets/swapedoc/hindi-kumaoni-parallel-corpus.texttranslation1K<n<10K0 likes20 downloads7mo agoHugging Face09nassimjp /english_pashto_parallel-corpus English–Pashto Parallel Corpus A high-quality English–Pashto parallel corpus designed for machine translation, supervised fine-tuning, Pashto LLM training, and corpus quality research. 📁 Dataset Structure 1. translation_clean.jsonl High-quality parallel data suitable for: Machine Translation Supervised Fine-Tuning (SFT) Pashto LLM Training Instruction Tuning Reasoning Alignment Each record follows: { "id": 12345, "source": "English… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/english_pashto_parallel-corpus.text100K<n<1M0 likes19 downloads2mo agoHugging Face10Mxode /Chinese-English-Parallel-Synonym-Corpus-75ktext10K<n<100K1 likes18 downloads1y agoHugging Face11muradbaghirli /az-eng-parallel-corpustext100K<n<1M0 likes17 downloads10mo agoHugging Face12Arseniy-Polyakov /parallel_corpus_russian_rsl_glossestext1K<n<10K0 likes16 downloads7mo agoHugging Face13dayomtechnologies /English_Nuer_parallel_translation_open_corpustext1K<n<10K0 likes16 downloads2mo agoHugging Face14shangzx /Chinese-English-Parallel-Translation-Corpus-Chinese-Source-Text-English-Translat Chinese-English Parallel Translation Corpus (Chinese Source Text & English Translation) A Chinese–English parallel corpus resource for translation and cross-lingual alignment applications, providing one-to-one bilingual text pairs: Chinese source texts aligned with their corresponding English translations. The data covers common writing styles and domains, making it suitable for parallel alignment, translation modeling, and cross-lingual representation learning. It supports… See the full description on the dataset page: https://huggingface.co/datasets/shangzx/Chinese-English-Parallel-Translation-Corpus-Chinese-Source-Text-English-Translat.texttext-classificationn<1K0 likes12 downloads2mo agoHugging Face15atsushi3110 /en-ja-parallel-corpus-augmentedtext1M<n<10M2 likes11 downloads2y agoHugging Face16kambale /luganda-english-parallel-corpusgated English-Luganda Parallel Corpus for Translation Dataset Description This dataset contains parallel sentences in English (en) and Luganda (lg), designed primarily for training and fine-tuning machine translation models. The data consists of sentence pairs extracted from a source document. Languages English (en) Luganda (lg) - ISO 639-1 code: lg Data Format The dataset is provided in a format compatible with the Hugging Face datasets library. Each… See the full description on the dataset page: https://huggingface.co/datasets/kambale/luganda-english-parallel-corpus.texttranslation10K<n<100K8 likes10 downloads1y agoHugging Face17alek1001 /JParaCrawl-Filtered-English-Japanese-Parallel-Corpus-texttext100K<n<1M2 likes9 downloads2y agoHugging Face18dayomtechnologies /english-nuer-dinka-parallel-corpusgated English–Nuer–Dinka Parallel Corpus Overview The English–Nuer–Dinka Parallel Corpus is a multilingual parallel dataset created to support research on low-resource African languages. The corpus contains aligned text in English, Nuer, and Dinka for use in Natural Language Processing (NLP), Machine Translation (MT), multilingual language modeling, and language preservation. The primary goal of this project is to increase the digital presence of Nuer and Dinka while… See the full description on the dataset page: https://huggingface.co/datasets/dayomtechnologies/english-nuer-dinka-parallel-corpus.texttranslation100K<n<1M1 likes7 downloads3mo agoHugging Face19dayomtechnologies /english_nuer-thok_naath_conversational_parallel_corpusgated English–Nuer (Thok Naath) Conversational Parallel Corpus Overview The English–Nuer (Thok Naath) Conversational Parallel Corpus is a bilingual dataset consisting of aligned conversational sentence pairs in English and Nuer (Thok Naath). This dataset was created to support research on low-resource African languages, with a particular focus on conversational AI, machine translation, multilingual language models, and language preservation. As one of the few publicly… See the full description on the dataset page: https://huggingface.co/datasets/dayomtechnologies/english_nuer-thok_naath_conversational_parallel_corpus.texttranslation100K<n<1M0 likes6 downloads3mo agoHugging Face20ashuChufamo /parallel-corpus_en-amtexttranslation10K<n<100K0 likes5 downloads2y agoHugging Face21emuduki /ateso-english-parallel-corpustext10K<n<100K0 likes5 downloads4mo agoHugging Face22abhinandansamal /odia-german-parallel-corpus-researchgated Dataset Summary This dataset is a high-quality, parallel corpus for Odia (Oriya) to German and German to Odia machine translation. It focuses on the news domain, specifically covering National, International, Sports, Trade, and Science & Technology topics. The dataset contains 3,676 unique parallel sentence pairs, curated through a hybrid approach combining automated web scraping, manual human translation (Gold Standard), and human-corrected machine translation (Silver… See the full description on the dataset page: https://huggingface.co/datasets/abhinandansamal/odia-german-parallel-corpus-research.tabulartranslation1K<n<10K0 likes2 downloads9mo agoHugging Face23LindaSekhoasha /zu-en_parallel-corpus_xsmtext10K<n<100K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.