CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sentence-transformers /parallel-sentences-ccmatrix Dataset Card for Parallel Sentences - CCMatrix This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. The texts originate from the CCMatrix dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices parallel-sentences-muse parallel-sentences-jw300 parallel-sentences-news-commentary… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-ccmatrix.textfeature-extraction1B<n<10B15 likes6.8k downloads2y agoHugging Face02sentence-transformers /parallel-sentences-talks Dataset Card for Parallel Sentences - Talks This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the Talks dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-talks.textfeature-extraction10M<n<100M12 likes6.1k downloads2y agoHugging Face03Bingsu /st-parallel-sentences Dataset Card for "st-parallel-sentences" More Information needed text100M<n<1B1 likes4.9k downloads3y agoHugging Face04sentence-transformers /parallel-sentences-opensubtitles Dataset Card for Parallel Sentences - OpenSubtitles This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the OpenSubtitles dataset. Warning! The quality of this dataset is not great; many of the english and non-english texts don't match well, or are fully empty. Related Datasets The following… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-opensubtitles.textfeature-extraction100M<n<1B4 likes4.7k downloads2y agoHugging Face05oss-codes /NCERT-Parallel-Dataset-Indictexttranslation100K<n<1M2 likes4.5k downloads1y agoHugging Face06mrlbenchmarks /global-piqa-parallel Global PIQA Parallel Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems. In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.imagequestion-answering10K<n<100K10 likes4k downloads4mo agoHugging Face07sentence-transformers /parallel-sentences-tatoeba Dataset Card for Parallel Sentences - Tatoeba This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the Tatoeba dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-tatoeba.textfeature-extraction1M<n<10M0 likes3.8k downloads2y agoHugging Face08strombergnlp /bornholmsk_parallelThis dataset is parallel text for Bornholmsk and Danish. For more details, see the paper [Bornholmsk Natural Language Processing: Resources and Tools](https://aclanthology.org/W19-6138/).texttranslation1K<n<10K2 likes3.6k downloads4y agoHugging Face09sentence-transformers /parallel-sentences-wikimatrix Dataset Card for Parallel Sentences - WikiMatrix This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the WikiMatrix dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-wikimatrix.textfeature-extraction10M<n<100M8 likes3.5k downloads2y agoHugging Face10ghanaopenai /kasem-speech-text-parallel This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Kasem Speech-Text Parallel Dataset Dataset Description This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kasem-speech-text-parallel.audioautomatic-speech-recognition10K<n<100K0 likes2.7k downloads3mo agoHugging Face11sentence-transformers /parallel-sentences-jw300 Dataset Card for Parallel Sentences - JW300 This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the JW300 dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-jw300.textfeature-extraction10M<n<100M10 likes2.1k downloads2y agoHugging Face12sentence-transformers /parallel-sentences-opus-100 Dataset Card for Parallel Sentences - OPUS-100 This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. The sentences originate from the OPUS-100 website. In particular, this dataset is a reformatting of the OPUS-100 dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-opus-100.textfeature-extraction10M<n<100M4 likes1.8k downloads2y agoHugging Face13sentence-transformers /parallel-sentences-europarl Dataset Card for Parallel Sentences - Europarl This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the Europarl dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-europarl.textfeature-extraction10M<n<100M1 likes1.8k downloads2y agoHugging Face14aiana94 /polynews-parallel Dataset Card for PolyNewsParallel Dataset Summary PolyNewsParallel is a multilingual paralllel dataset containing news titles for 833 language pairs. It covers 64 languages and 17 scripts. Uses This dataset can be used for machine translation or text retrieval. Languages There are 64 languages avaiable: Code Language Script amh_Ethi Amharic Ethiopic arb_Arab Modern Standard Arabic Arabic ayr_Latn Central Aymara Latin bam_Latn Bambara… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/polynews-parallel.imagetranslation1M<n<10M15 likes1.7k downloads2y agoHugging Face15ghanaopenai /ga-speech-text-parallel-90k This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github Ga Speech-Text Parallel Dataset Dataset Description This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-speech-text-parallel-90k.audioautomatic-speech-recognition10K<n<100K0 likes1.3k downloads3mo agoHugging Face16cloverx-id /xone-repository-parallel-en-id-corpusWe are currently developing new version of LMSE translation scoring model and processing additional data sources. We estimate the dataset will expand, with significantly improved quality.(Delayed..) A score of 55% and above indicates high-quality translation pairs, even if the first version of the model we developed gave them such a score. We will try to release a newer model in the future with better quality and consistently fast scoring speeds, and release it to the public once we decide… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/xone-repository-parallel-en-id-corpus.tabulartranslation10M<n<100M1 likes1.3k downloads4d agoHugging Face17Mathoctopus /GSM8KInstruct_Paralleltextquestion-answering10K<n<100K11 likes1.3k downloads3y agoHugging Face18sentence-transformers /parallel-sentences-global-voices Dataset Card for Parallel Sentences - Global Voices This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the Global Voices dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-global-voices.textfeature-extraction1M<n<10M1 likes1.2k downloads2y agoHugging Face19ghanaopenai /twi-trigrams-speech-text-parallel This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Trigrams Speech-Text Parallel Dataset Dataset Description This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-trigrams-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M0 likes1.2k downloads3mo agoHugging Face20oss-codes /Cyber-Parallel-Dataset-Indictext1K<n<10K0 likes1k downloads1y agoHugging Face21oss-codes /Finance-Parallel-Dataset-Indictext100K<n<1M0 likes822 downloads1y agoHugging Face22Lego-MT /Parallel_Dataset Dataset Sources Paper: LegoMT2: Selective Asynchronous Sharded Data Parallel Training for Massive Neural Machine Translation Link: https://aclanthology.org/2025.findings-acl.1200.pdf Repository: https://github.com/CONE-MT/CONE texttranslation1M<n<10M0 likes755 downloads1y agoHugging Face23haoranxu /X-ALMA-Parallel-Data This is the translation parallel dataset used by X-ALMA. @misc{xu2024xalmaplugplay, title={X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at Scale}, author={Haoran Xu and Kenton Murray and Philipp Koehn and Hieu Hoang and Akiko Eriguchi and Huda Khayrallah}, year={2024}, eprint={2410.03115}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2410.03115}, } text100K<n<1M8 likes657 downloads2y agoHugging Face24oss-codes /Law-Parallel-Dataset-Indictext100K<n<1M0 likes625 downloads1y agoHugging Face25browndw /human-ai-parallel-corpus Human-AI Parallel English Corpus (HAP-E) 🙃 Purpose The HAP-E corpus is designed for comparisions of the writing produced by humans and the writing produced by large language models (LLMs). The corpus was created by seeding an LLM with an approximately 500-word chunk of human-authored text and then prompting the model to produce an additional 500 words. Thus, a second 500-word chunk of human-authored text (what actually comes next in the original text) can be compared to… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus.texttext-classification10K<n<100K3 likes596 downloads2y agoHugging Face26akberto /ParallelCacheFlow_77_12textn<1K0 likes488 downloads11mo agoHugging Face27Sudehsna /Romansh_German_Parallel_Data Romansh–German Parallel Dataset (FineWeb-Based) This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction. Description This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.tabular10K<n<100K2 likes472 downloads1y agoHugging Face28tiny-aya-translate /tr-hi-parallel-speech-v2 TR↔HI Parallel Speech (v2) — synthetic TTS corpus The raw speech corpus behind TinyAya Stage 2: ~911 hours of synthetic Turkish⇄Hindi parallel speech, 53,506 rows, generated with OmniVoice across 14 voice designs. This is the pre-encoding source. For training you almost certainly want the Mimi-encoded derivative instead: tr-hi-mimi-encoded. Layout path contents data/train-*.parquet the loadable table (schema in the YAML header above) audio/*.wav ~9… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-parallel-speech-v2.audioaudio-to-audio100K<n<1M1 likes468 downloads2mo agoHugging Face29parallel-reasoner /Step-analysistext100K<n<1M0 likes465 downloads2mo agoHugging Face30haoranxu /ALMA-Human-Parallel Dataset Card for "ALMA-Human-Parallel" This is human-written parallel dataset used by ALMA translation models. @misc{xu2023paradigm, title={A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models}, author={Haoran Xu and Young Jin Kim and Amr Sharaf and Hany Hassan Awadalla}, year={2023}, eprint={2309.11674}, archivePrefix={arXiv}, primaryClass={cs.CL} } @misc{xu2024contrastive, title={Contrastive… See the full description on the dataset page: https://huggingface.co/datasets/haoranxu/ALMA-Human-Parallel.text10K<n<100K8 likes455 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.