datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vietnamese-THUIR-T2Ranking-gg-translated
📚 5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated
📝 Overview
Vietnamese-THUIR-T2Ranking-gg-translated is a large-scale dataset for passage ranking in Vietnamese.It is translated from the original THUIR/T2Ranking [1] using Google Translate, inspired by the approach of mMARCO [2].The dataset aims to provide a large-scale dataset for research and applications in Information Retrieval (IR) in Vietnamese.
In IR, passage ranking is an essential and challenging task… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated.gsm8k-translatedgsm8k-translated
Multilingual GSM8K Translations
This dataset contains machine-translated versions of GSM8K in these languages:
French (fr)
German (de)
Hindi (hi)
Dataset Structure
For each language, we provide the original GSM8K train and test splits:
train: 7,473 samples
test: 1,319 samples
Each sample consists of a question and an answer.
The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.h2o-translated-chinese-med-prompts
Translated Chinese Medical Prompts
This repository contains medical prompts translated originally from Chinese, which can be used as training data for natural language processing (NLP) tasks related to the medical domain in English language.
Dataset Description
The dataset consists of a collection of medical prompts originally in Chinese, which have been translated into English. These prompts cover various medical topics, including symptoms, diagnoses, treatments, medications, and… See the full description on the dataset page: https://huggingface.co/datasets/h2oai/h2o-translated-chinese-med-prompts.semeval-2016-absa-reviews-english-translated-stanford-alpaca
Dataset Card for Dataset Name
Derived from eastwind/semeval-2016-absa-reviews-arabic using Helsinki-NLP/opus-mt-tc-big-ar-en
translated-dataset-synthetic-retrieval-tasksxnli-translated-khm-pairclassificationxnli-translated-zsm-pairclassificationxnli-translated-lao-pairclassificationTranslated_Books
Translated Books
⚠️ Disclaimer: This is a personal translation project and is NOT part of the OmniMedical Suite ecosystem.
It is unrelated to medical OCR, handwriting recognition, or any of the author's medical AI work.
Dataset Description
A personal collection of English-to-Arabic book translations compiled as a parallel corpus.
This dataset is maintained separately from the author's professional medical AI projects.
Files
File
Format… See the full description on the dataset page: https://huggingface.co/datasets/DrAbdulmalek/Translated_Books.semeval-2016-absa-reviews-english-translated-resampled
Dataset Card for Hotel Review ABSA (SemEval 2016 Translated from Arabic)
Dataset Description
Derived from eastwind/semeval-2016-absa-reviews-english-translated-stanford-alpaca, by upsampling the neutral class and then resampling 3k examples from each class
dell-qa-en-to-ko-translated-by-ke-t5-base
Dell QA English to Korean Translation Dataset
Dataset Description
This dataset, dell-qa-en-to-ko-translated-by-ke-t5-base, is a Korean translation of the original English Dell QA dataset.
Source
The original dataset, dell_qa, is designed for question-answering tasks and contains questions and answers related to Dell technologies. This translated version extends the utility to Korean language tasks.
Dataset Structure
Data Fields
input… See the full description on the dataset page: https://huggingface.co/datasets/seongs/dell-qa-en-to-ko-translated-by-ke-t5-base.google_go_emotions_hindi_translatedsts-arabic-translated-modifiedptbr-quora-translated
Dataset Summary
The Quora dataset is composed of question pairs, and the task is to determine if the questions are paraphrases of each
other (have the same meaning). The dataset was translated to Portuguese using the model seamless-m4t-medium.
Languages
Portuguese
ccs_synthetic_translated_arabicThe columns inside the dataset as follows:
index
url
caption_en
caption_ar
The dataset size is 12556500 rows × 4 columns
ccs_synthetic_translated_arabic_processedtranslated_datatranslated_datasettranslated_facts_test_setsemeval-2016-absa-reviews-english-translated-stanford-alpaca
Dataset Card for Dataset Name
Derived from eastwind/semeval-2016-absa-reviews-arabic using Helsinki-NLP/opus-mt-tc-big-ar-en
translated_da_entranslated_sqlignmilton-translated-dataset-v1.0Translated_Expanded_CC3M-Brazilian_Portuguese-Hindi-Xhosa
CC3M Multilingual & Augmented Variants
This repository provides four multilingual, augmented, and similarity-enhanced variants of the Conceptual Captions 3M (CC3M) dataset.The goal is to support research in vision–language modeling, multimodal alignment, data augmentation, and low-resource language evaluation.
All versions include translations generated with Google Translate and MarianMT, and caption augmentations produced with BLIP2, generating five additional captions per… See the full description on the dataset page: https://huggingface.co/datasets/DiegoAlysson/Translated_Expanded_CC3M-Brazilian_Portuguese-Hindi-Xhosa.hh_dpo_kannada_translatedDavidson_back_translated_alltranslated_factsThis dataset contains LLM-based translations (via Aya-Expanse) of the full fact-space for SemEval Task 7 - Multilingual and Crosslingual Fact-Checked Claim Retrieval. However, it was not used in the final pipeline due to a lack of performance gains. Further improvements via translation refinement of the fact-space were not pursued due to high computational costs, and no definitive conclusions were drawn about the feasibility of this direction.
data_problems_translatedtranslated-dataset
