datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
translationrelaion2B-en-research-safe-japanese-translation
relaion2B-en-research-safe-japanese-translation
This dataset is the Japanese translation of the English subset of ReLAION-5B (laion/relaion2B-en-research-safe), translated by gemma-2-9b-it.
We used text2dataset for translating with open-weight LLMs.
By leveraging the fast LLM inference library vLLM, this tool enables the rapid translation of large English datasets into Japanese.
Prompt
The following is the prompt used for translation with Gemma.
You are an… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/relaion2B-en-research-safe-japanese-translation.wellbeing-in-translation
Wellbeing in Translation
Raw outputs and translated materials for Does AI Wellbeing Survive Translation? We test whether the unchanged CAIS 1-7 self-report battery measures the same positive-minus-negative gap after translation.
Paper · Code · Source instrument
Headline result
Language sensitivity is specific to the model-battery pair.
Model
Gap spread across 7 languages
English rank
English stimulus / local battery
Local stimulus / English battery… See the full description on the dataset page: https://huggingface.co/datasets/ic-org/wellbeing-in-translation.machine-translation-for-vision
Machine Translation for Vision (MTV)
Given the rise of multimedia content, human translators increasingly focus on culturally adapting not only words but also other modalities such as images to convey the same meaning. While several applications stand to benefit from this, machine translation systems remain confined to dealing with language in speech
and text. In this work, we introduce a new task of translating images to make them culturally relevant (image transcreation). For… See the full description on the dataset page: https://huggingface.co/datasets/cmu-lti/machine-translation-for-vision.ImageCaptions-7M-Translations-Arabicpixmo-captions-arabic-translations_100K
PixMo CAP - Arabic Translations
Dataset Description
This dataset contains 100,000 randomly selected samples from the allenai/pixmo-cap dataset with English captions translated to Modern Standard Arabic.
Translation Details
Translator Model: quickmt/quickmt-en-ar
Beam Size: 5
Language Pair: English → Arabic
Total Samples: 10,000
Dataset Structure
Data Fields
image_link: URL link to the original image
original_caption:… See the full description on the dataset page: https://huggingface.co/datasets/JadwalAlmaa/pixmo-captions-arabic-translations_100K.ImgMT-Dataset-For-Evaluating-Text-Image-Translationcauldron_translations_nu
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/zakcination/cauldron_translations_nu.ImageCaptions-7M-Translationsversion https://git-lfs.github.com/spec/v1
oid sha256:835f3f7d88a86e05a882c6a6b6333da6ab874776385f85473798769d767c2fca
size 27
filtered-llamaindex-with-translationsphilosophy-culture-translations-html-csv
AI-Culture Philosophy and Culture Translations CSV + HTML Corpus
The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind.
This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.multilingual-image-text-translation
Multilingual Image-Text Translation Dataset (MMT)
Overview
This dataset contains Multilingual Multimodal Translation (MMT) pairs from FLORES-200, featuring image-text combinations across 11 languages. The dataset is designed for training and evaluating multimodal translation models that can translate text while considering visual context.
MMT (Multilingual Multimodal Translation)
Multilingual: 11 languages (en, id, ja, kk, ko, ru, ur, uz, vi, zh-cn, zh-tw)… See the full description on the dataset page: https://huggingface.co/datasets/rileykim/multilingual-image-text-translation.graph_translation_eng_kh
Dataset Card for hf_10k_charts
Features: image_eng:image, image_kh:image, text_eng:string, text_kh:string
Total examples: 10000 (train: 8000, test: 2000)
Total file size: 470MB
graph_translation_eng_kh_syncgraph_translation_eng_kh_200k
Dataset Card for hf_10k_charts
Features: image_eng:image, image_kh:image, text_eng:string, text_kh:string
Total examples: 200000 (train: 160000, test: 40000)
Total file size: 8GB
vlm-project-with-images-distribution-q2-translation-all-language-officialpixmo-captions-arabic-translations
PixMo CAP - Arabic Translations
Dataset Description
This dataset contains 10,000 randomly selected samples from the allenai/pixmo-cap dataset with English captions translated to Modern Standard Arabic.
Translation Details
Translator Model: quickmt/quickmt-en-ar
Beam Size: 5
Language Pair: English → Arabic
Total Samples: 10,000
Dataset Structure
Data Fields
image_link: URL link to the original image
original_caption:… See the full description on the dataset page: https://huggingface.co/datasets/JadwalAlmaa/pixmo-captions-arabic-translations.Ainu-Japan_translation_model
Dataset Card for [Dataset Name]
Table of Contents
Dataset Description
Dataset Summary
Supported Tasks and Leaderboards
Languages
Dataset Structure
Data Instances
Data Fields
Data Splits
Dataset Creation
Curation Rationale
Source Data
Annotations
Personal and Sensitive Information
Considerations for Using the Data
Social Impact of Dataset
Discussion of Biases
Other Known Limitations
Additional Information
Dataset Curators
Licensing Information
Citation Information… See the full description on the dataset page: https://huggingface.co/datasets/Sotaro0124/Ainu-Japan_translation_model.translation_laqwen3_dataset_translationtranslation_vi_2kImageCaptions-7M-Translations-Arabic-subset-150000filtered-colpali-with-translationsvlm-project-with-images-distribution-q2-translation-all-languagetranslation_viparagraph_translation_eng_kh
Machine Translation Dataset: Khmer to English
A machine translation dataset project that pairs Khmer document images with their corresponding English translations. This dataset is designed for training machine translation models to convert Khmer text (via OCR from images) to English text.
Project Overview
This project processes bilingual documents from BOP (Bank of PNG) Bulletins, creating a Khmer-English machine translation dataset. Each sample pairs a Khmer image… See the full description on the dataset page: https://huggingface.co/datasets/vichetkao/paragraph_translation_eng_kh.table_translation_eng_kh_100k
Multilingual Chart Dataset
Dataset Overview
Total examples: 100000 (train: 90000, test: 10000)
Total size: 8GB
Languages: English (en), Khmer (km)
Chart types: 22+ types (bar, line, pie, gauge, heatmap, candlestick, etc.)
Features
image_eng: PNG image of chart in English
image_kh: PNG image of chart in Khmer
text_eng: JSON metadata in English (chart_type, title, axis labels, data)
text_kh: JSON metadata in Khmer (chart_type, title, axis labels, data)
import… See the full description on the dataset page: https://huggingface.co/datasets/vichetkao/table_translation_eng_kh_100k.vlm-project-with-images-distribution-q2-translation-inlcude-a2vlm-project-with-images-distribution-q2-translation-inlcude-a2-q3-q4vlm-project-with-images-distribution-q2-translation-all-language-v2
