CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01adeshkin /khakas-russian-parallel-corpus Khakas-Russian Parallel Corpus The creation of this dataset is aimed at supporting the development of natural language processing (NLP) tools and machine translation for the Khakas language, which is classified as a "Definitely Endangered" language. By providing high-quality parallel data, this project helps preserve the linguistic heritage of the Khakas people. Dataset Overlap: The Khakas sentences in this corpus do not overlap with those in the Khakas… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-parallel-corpus.texttranslation100K<n<1M2 likes277 downloads12d agoHugging Face02tunis-ai /tunisian-msa-parallel-corpus Dataset Description This is an ambitious project to create a high-quality, reproducible parallel corpus for Modern Standard Arabic (MSA) and Tunisian Arabic (aeb) through a sophisticated synthetic data generation pipeline. The dataset is being developed by the Tunisia.AI community to address the scarcity of high-quality dialectal data for training and evaluating language models. The primary goal is to provide a rich, well-documented resource for the research and development of:… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus.tabulartranslation1K<n<10K0 likes98 downloads1y agoHugging Face03abdelhaqueidali /Kabyle-Latin-to-Tifinagh-Parallel-Corpus Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.texttranslation1M<n<10M0 likes57 downloads3mo agoHugging Face04sarjukesumo /quran-parallel-corpus Quran Parallel Corpus Verse-aligned Quran parallel corpus — Arabic (Uthmani), English (Sahih International), and Indonesian (Ministry of Religious Affairs). Stats Total verses: 6236 Languages: Arabic, English, Indonesian Translation pairs: Arabic↔English, Arabic↔Indonesian, English↔Indonesian Formats: JSONL, CSV, Parquet Structure Each verse record contains: Field Description surah_number Chapter (1–114) surah_name_arabic Arabic surah… See the full description on the dataset page: https://huggingface.co/datasets/sarjukesumo/quran-parallel-corpus.tabulartranslation100K<n<1M0 likes57 downloads1mo agoHugging Face05DatarrX /Myanmar-Written-Spoken-Parallel-Corpus Myanmar Written-Spoken Parallel Corpus (MWSPC) Dataset Description Myanmar Written-Spoken Parallel Corpus (MWSPC) is a high-quality open-source dataset designed to bridge the gap between formal written Burmese and daily spoken Burmese. This dataset is crucial for building natural-sounding AI models that understand the linguistic nuances of the Myanmar language. Curated by: Khant Sint Heinn (Kalix Louis) Organization: DatarrX | ဒေတာ-အက်စ် Language: Burmese… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Myanmar-Written-Spoken-Parallel-Corpus.texttext-generation1K<n<10K6 likes47 downloads4mo agoHugging Face06nickoo004 /kaa-parallel-corpus Kaa Karakalpak-English Parallel Corpus (FineTranslations) 📌 Overview This repository contains a high-quality, curated parallel corpus for the Karakalpak (kaa) language, paired with English (en). Karakalpak is a low-resource Turkic language spoken primarily in the Republic of Karakalpakstan. This dataset is a specialized subset extracted from the massive HuggingFaceFW/finetranslations project. The goal of this repo is to provide a dedicated and easy-to-access resource… See the full description on the dataset page: https://huggingface.co/datasets/nickoo004/kaa-parallel-corpus.tabulartranslation10K<n<100K0 likes41 downloads5mo agoHugging Face07Lots-of-LoRAs /task441_eng_guj_parallel_corpus_gu-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task441_eng_guj_parallel_corpus_gu-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task441_eng_guj_parallel_corpus_gu-en_language_identification.texttext-generation1K<n<10K0 likes34 downloads2y agoHugging Face08tunis-ai /tunisian-msa-parallel-corpus-evaluated Dataset Description This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb). It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP. The primary goals are to support: Machine translation between Tunisian Arabic and MSA. Research in dialectal-aware text generation and evaluation. Cross-dialect representation learning in… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus-evaluated.tabulartranslation1K<n<10K2 likes34 downloads1y agoHugging Face09Lots-of-LoRAs /task440_eng_guj_parallel_corpus_gu-en_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task440_eng_guj_parallel_corpus_gu-en_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task440_eng_guj_parallel_corpus_gu-en_classification.texttext-generation1K<n<10K0 likes30 downloads2y agoHugging Face10kingkaung /islamqainfo_parallel_corpus Dataset Card for IslamQA Info Parallel Corpus Dataset Description The IslamQA Info Parallel Corpus is a multilingual dataset derived from the IslamQA repository. It contains curated question-and-answer pairs across 17 languages, making it a valuable resource for multilingual and cross-lingual natural language processing (NLP) tasks. The dataset has been created over nearly three decades (since 1997) by Sheikhul Islam Muhammad Saalih al-Munajjid and his team. Key… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/islamqainfo_parallel_corpus.texttable-question-answering10K<n<100K2 likes30 downloads2y agoHugging Face11Lots-of-LoRAs /task439_eng_guj_parallel_corpus_gu_en_translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task439_eng_guj_parallel_corpus_gu_en_translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task439_eng_guj_parallel_corpus_gu_en_translation.texttext-generation1K<n<10K0 likes27 downloads2y agoHugging Face12Lots-of-LoRAs /task438_eng_guj_parallel_corpus_en_gu_translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task438_eng_guj_parallel_corpus_en_gu_translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task438_eng_guj_parallel_corpus_en_gu_translation.texttext-generation1K<n<10K0 likes23 downloads2y agoHugging Face13tejasta /Kannada-English-Parallel-Corpus Dataset Card for Kannada-English Parallel Corpus Dataset Description This dataset provides a high-quality, curated collection of parallel English-Kannada sentence pairs. It is designed to address the critical data scarcity in low-resource language modeling. This corpus is intended to facilitate advancements in Neural Machine Translation (NMT), cross-lingual transfer learning, and instruction-tuning for Large Language Models (LLMs) to better serve Kannada speakers.… See the full description on the dataset page: https://huggingface.co/datasets/tejasta/Kannada-English-Parallel-Corpus.texttranslation1K<n<10K7 likes13 downloads3mo agoHugging Face14bekan /english_karakalpak_parallel_corpus_v1 English-Karakalpak Parallel Corpus (en-kaa) Dataset Description English-Karakalpak Parallel Corpus is a high-quality dataset containing 10,441 aligned sentence pairs in English and Karakalpak (kaa). This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. The corpus utilizes the official… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v1.texttranslation10K<n<100K2 likes12 downloads10mo agoHugging Face15hbenayed /tunisian-msa-parallel-corpus-evaluated Dataset Description This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb). It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP. The primary goals are to support: Machine translation between Tunisian Arabic and MSA. Research in dialectal-aware text generation and evaluation. Cross-dialect representation learning in Arabic… See the full description on the dataset page: https://huggingface.co/datasets/hbenayed/tunisian-msa-parallel-corpus-evaluated.tabulartranslation1K<n<10K0 likes12 downloads6mo agoHugging Face16dayomtechnologies /english-nuer-dinka-parallel-corpusgated English–Nuer–Dinka Parallel Corpus Overview The English–Nuer–Dinka Parallel Corpus is a multilingual parallel dataset created to support research on low-resource African languages. The corpus contains aligned text in English, Nuer, and Dinka for use in Natural Language Processing (NLP), Machine Translation (MT), multilingual language modeling, and language preservation. The primary goal of this project is to increase the digital presence of Nuer and Dinka while… See the full description on the dataset page: https://huggingface.co/datasets/dayomtechnologies/english-nuer-dinka-parallel-corpus.texttranslation100K<n<1M1 likes7 downloads3mo agoHugging Face17ashuChufamo /parallel-corpus_en-amtexttranslation10K<n<100K0 likes5 downloads2y agoHugging Face18bekan /english_karakalpak_pairs_parallel_corpus_v2_8907 English-Karakalpak Parallel Corpus v2 (8.9K) Dataset Description English-Karakalpak Parallel Corpus v2 is a high-quality dataset containing 8,906 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. This resource… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_pairs_parallel_corpus_v2_8907.texttranslation1K<n<10K1 likes4 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.