CoolFace
8 results

parallel-corpus

cloverx-id /xone-repository-parallel-en-id-corpusWe are currently developing new version of LMSE translation scoring model and processing additional data sources. We estimate the dataset will expand, with significantly improved quality.(Delayed..) A score of 55% and above indicates high-quality translation pairs, even if the first version of the model we developed gave them such a score. We will try to release a newer model in the future with better quality and consistently fast scoring speeds, and release it to the public once we decide… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/xone-repository-parallel-en-id-corpus.tabulartranslation10M<n<100M1 likes1.3k downloads4d agoHugging Facebrowndw /human-ai-parallel-corpus Human-AI Parallel English Corpus (HAP-E) 🙃 Purpose The HAP-E corpus is designed for comparisions of the writing produced by humans and the writing produced by large language models (LLMs). The corpus was created by seeding an LLM with an approximately 500-word chunk of human-authored text and then prompting the model to produce an additional 500 words. Thus, a second 500-word chunk of human-authored text (what actually comes next in the original text) can be compared to… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus.texttext-classification10K<n<100K3 likes596 downloads2y agoHugging FaceFrancophonIA /UFAL_Parallel_Corpus_of_North_Levantine_1.0 [!NOTE] Dataset origin: https://zenodo.org/records/4012218 UFAL Parallel Corpus of North Levantine 1.0 March 10, 2023 Authors Shadi Saleh <saleh@ufal.mff.cuni.cz> Hashem Sellat <sellat@ufal.mff.cuni.cz> Mateusz Krubiński <krubinski@ufal.mff.cuni.cz> Adam Posppíšil <adam.pospisil@ff.cuni.cz> Petr Zemánek <petr.zemanek@ff.cuni.cz> Pavel Pecina <pecina@ufal.mff.cuni.cz> Overview This is the first release of the UFAL Parallel Corpus of North Levantine… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/UFAL_Parallel_Corpus_of_North_Levantine_1.0.0 likes310 downloads1y agoHugging Faceadeshkin /khakas-russian-parallel-corpus Khakas-Russian Parallel Corpus The creation of this dataset is aimed at supporting the development of natural language processing (NLP) tools and machine translation for the Khakas language, which is classified as a "Definitely Endangered" language. By providing high-quality parallel data, this project helps preserve the linguistic heritage of the Khakas people. Dataset Overlap: The Khakas sentences in this corpus do not overlap with those in the Khakas… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-parallel-corpus.texttranslation100K<n<1M2 likes277 downloads12d agoHugging Facemteb /english-danish-parallel-corpus DanishMedicinesAgencyBitextMining An MTEB dataset Massive Text Embedding Benchmark A Bilingual English-Danish parallel corpus from The Danish Medicines Agency. Task category t2t Domains Medical, Written Reference https://sprogteknologi.dk/dataset/bilingual-english-danish-parallel-corpus-from-the-danish-medicines-agency How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/english-danish-parallel-corpus.texttranslation10K<n<100K0 likes258 downloads1y agoHugging Facebekan /english_karakalpak_parallel_corpus_v5 English-Karakalpak Parallel Corpus This dataset contains parallel sentences in English and Karakalpak language. It is created to support AI development for the Karakalpak language. Dataset Description English-Karakalpak Parallel Corpus is a high-quality, dynamic dataset containing carefully aligned sentence pairs in English (en) and Karakalpak (kaa). Note: This dataset is updated frequently. New sentence pairs are added on a regular basis to continuously increase… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v5.texttranslation10K<n<100K4 likes227 downloads8d agoHugging Face