CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
015CD-AI /Vietnamese-THUIR-T2Ranking-gg-translated 📚 5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated 📝 Overview Vietnamese-THUIR-T2Ranking-gg-translated is a large-scale dataset for passage ranking in Vietnamese.It is translated from the original THUIR/T2Ranking [1] using Google Translate, inspired by the approach of mMARCO [2].The dataset aims to provide a large-scale dataset for research and applications in Information Retrieval (IR) in Vietnamese. In IR, passage ranking is an essential and challenging task… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated.tabulartext-retrieval100M<n<1B22 likes558 downloads1y agoHugging Face02Sara237 /gsm8k-translatedtext10K<n<100K0 likes222 downloads2y agoHugging Face03math-across-languages /gsm8k-translated Multilingual GSM8K Translations This dataset contains machine-translated versions of GSM8K in these languages: French (fr) German (de) Hindi (hi) Dataset Structure For each language, we provide the original GSM8K train and test splits: train: 7,473 samples test: 1,319 samples Each sample consists of a question and an answer. The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.textquestion-answering10K<n<100K0 likes62 downloads3mo agoHugging Face04h2oai /h2o-translated-chinese-med-prompts Translated Chinese Medical Prompts This repository contains medical prompts translated originally from Chinese, which can be used as training data for natural language processing (NLP) tasks related to the medical domain in English language. Dataset Description The dataset consists of a collection of medical prompts originally in Chinese, which have been translated into English. These prompts cover various medical topics, including symptoms, diagnoses, treatments, medications, and… See the full description on the dataset page: https://huggingface.co/datasets/h2oai/h2o-translated-chinese-med-prompts.text10K<n<100K0 likes59 downloads3y agoHugging Face05srinivasbilla /semeval-2016-absa-reviews-english-translated-stanford-alpaca Dataset Card for Dataset Name Derived from eastwind/semeval-2016-absa-reviews-arabic using Helsinki-NLP/opus-mt-tc-big-ar-en texttext-classification10K<n<100K3 likes45 downloads3y agoHugging Face06aashay96 /translated-dataset-synthetic-retrieval-taskstext10K<n<100K2 likes27 downloads3y agoHugging Face07kornwtp /xnli-translated-khm-pairclassificationtext1K<n<10K0 likes23 downloads2y agoHugging Face08kornwtp /xnli-translated-zsm-pairclassificationtext1K<n<10K0 likes23 downloads2y agoHugging Face09kornwtp /xnli-translated-lao-pairclassificationtext1K<n<10K0 likes22 downloads2y agoHugging Face10DrAbdulmalek /Translated_Books Translated Books ⚠️ Disclaimer: This is a personal translation project and is NOT part of the OmniMedical Suite ecosystem. It is unrelated to medical OCR, handwriting recognition, or any of the author's medical AI work. Dataset Description A personal collection of English-to-Arabic book translations compiled as a parallel corpus. This dataset is maintained separately from the author's professional medical AI projects. Files File Format… See the full description on the dataset page: https://huggingface.co/datasets/DrAbdulmalek/Translated_Books.texttranslationn<1K0 likes21 downloads3mo agoHugging Face11srinivasbilla /semeval-2016-absa-reviews-english-translated-resampled Dataset Card for Hotel Review ABSA (SemEval 2016 Translated from Arabic) Dataset Description Derived from eastwind/semeval-2016-absa-reviews-english-translated-stanford-alpaca, by upsampling the neutral class and then resampling 3k examples from each class text10K<n<100K5 likes20 downloads3y agoHugging Face12seongs /dell-qa-en-to-ko-translated-by-ke-t5-base Dell QA English to Korean Translation Dataset Dataset Description This dataset, dell-qa-en-to-ko-translated-by-ke-t5-base, is a Korean translation of the original English Dell QA dataset. Source The original dataset, dell_qa, is designed for question-answering tasks and contains questions and answers related to Dell technologies. This translated version extends the utility to Korean language tasks. Dataset Structure Data Fields input… See the full description on the dataset page: https://huggingface.co/datasets/seongs/dell-qa-en-to-ko-translated-by-ke-t5-base.textquestion-answering10K<n<100K1 likes20 downloads1y agoHugging Face13utkarsharora100 /google_go_emotions_hindi_translatedtabular100K<n<1M1 likes20 downloads2y agoHugging Face14Zynab /sts-arabic-translated-modifiedtabular1K<n<10K0 likes15 downloads3y agoHugging Face15luiseduardobrito /ptbr-quora-translated Dataset Summary The Quora dataset is composed of question pairs, and the task is to determine if the questions are paraphrases of each other (have the same meaning). The dataset was translated to Portuguese using the model seamless-m4t-medium. Languages Portuguese tabulartext-classification100K<n<1M3 likes14 downloads3y agoHugging Face16Arabic-Clip /ccs_synthetic_translated_arabicThe columns inside the dataset as follows: index url caption_en caption_ar The dataset size is 12556500 rows × 4 columns image10M<n<100M0 likes13 downloads2y agoHugging Face17Arabic-Clip /ccs_synthetic_translated_arabic_processedimage10M<n<100M1 likes13 downloads2y agoHugging Face18wikd /translated_datatext1K<n<10K0 likes12 downloads2y agoHugging Face19Xondamir /translated_datasettextsummarizationn<1K1 likes12 downloads2y agoHugging Face20prasannad28 /translated_facts_test_settext1K<n<10K0 likes12 downloads2y agoHugging Face21tomTAMAPSM /semeval-2016-absa-reviews-english-translated-stanford-alpaca Dataset Card for Dataset Name Derived from eastwind/semeval-2016-absa-reviews-arabic using Helsinki-NLP/opus-mt-tc-big-ar-en texttext-classification10K<n<100K0 likes11 downloads5mo agoHugging Face22Mabeck /translated_da_entext10K<n<100K1 likes10 downloads2y agoHugging Face23marcelohaps /translated_sqltext10K<n<100K3 likes10 downloads2y agoHugging Face24atharvanighot /ignmilton-translated-dataset-v1.0text100K<n<1M1 likes10 downloads2y agoHugging Face25DiegoAlysson /Translated_Expanded_CC3M-Brazilian_Portuguese-Hindi-Xhosa CC3M Multilingual & Augmented Variants This repository provides four multilingual, augmented, and similarity-enhanced variants of the Conceptual Captions 3M (CC3M) dataset.The goal is to support research in vision–language modeling, multimodal alignment, data augmentation, and low-resource language evaluation. All versions include translations generated with Google Translate and MarianMT, and caption augmentations produced with BLIP2, generating five additional captions per… See the full description on the dataset page: https://huggingface.co/datasets/DiegoAlysson/Translated_Expanded_CC3M-Brazilian_Portuguese-Hindi-Xhosa.text1M<n<10M0 likes9 downloads10mo agoHugging Face26Cognitive-Lab /hh_dpo_kannada_translatedtext10K<n<100K0 likes8 downloads3y agoHugging Face27pasan-SK /Davidson_back_translated_alltext10K<n<100K0 likes8 downloads2y agoHugging Face28prasannad28 /translated_factsThis dataset contains LLM-based translations (via Aya-Expanse) of the full fact-space for SemEval Task 7 - Multilingual and Crosslingual Fact-Checked Claim Retrieval. However, it was not used in the final pipeline due to a lack of performance gains. Further improvements via translation refinement of the fact-space were not pursued due to high computational costs, and no definitive conclusions were drawn about the feasibility of this direction. text100K<n<1M0 likes8 downloads2y agoHugging Face29petrovortex /data_problems_translatedtext1K<n<10K0 likes8 downloads11mo agoHugging Face30Turbs /translated-datasettext1K<n<10K0 likes7 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.