datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-edu-translated
Helsinki-NLP/fineweb-edu-translated
fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu.
Translations are based on OPUS-MT and HPLT-MT models.
The data in v1.0 covers 36,704,000 documents with over 28 billion space-searated tokens of English data translated into 36 languages.
The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages.
In the v1.1 release, additional translations… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/fineweb-edu-translated.nemotron-cc-translated
Helsinki-NLP/nemotron-cc-translated
nemotron-cc-tanslated is a collection of automatically translated documents from nemotron-cc taken out of the high-quality subset.
Translations are based on OPUS-MT and HPLT-MT models.
The data in v1.0 covers 156,431,999 documents with over 70 billion space-searated tokens of English data translated into 36 languages.
The total v1.0 data set includes over 2.4 trillion tokens and the translated documents are aligned across all languages.
v1.1… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/nemotron-cc-translated.Dolci-Instruct-SFT-translatedDolci-Think-SFT-translated
Dolci-Think-SFT-translated
Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations.
Columns
Each row is a translated conversation plus the result of a post-translation quality filter:
id — source record id.
messages — the translated conversation (list of {content, role}).
filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.nemotron-cc-10K-sample-translated
Translated Nemotron-cc-hq samples
This dataset contains translated samples from https://huggingface.co/datasets/spyysalo/nemotron-cc-10K-sample
Currently, the following are available, we will add other models and languages:
Model
Languages
Gemma-3-4b-it
["Bulgarian", "Czech", "Danish", "German", "Estonian", "Finnish", "French", "Croatian", "Dutch"]
EuroLLM-9B-Instruct
["Bulgarian", "Czech", "Danish", "German", "Greek", "Estonia", "Finnish", "French", "Irish"… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/nemotron-cc-10K-sample-translated.bhasha-wiki-translated
Bhasha Wikipedia Translated
Translated wikipedia articles
Dataset Details
Dataset is being updated
Dataset Description
We have translated 6.185 million English wikipedia articles into 6 Indic languages. The translations were done using IndicTrans2 model.
Curated by: Soket AI labs
Language(s) (NLP): Hindi, Bengali, Gujarati, Tamil, Kannada, Urdu
License: cc-by-sa-4.0
Uses
For pretraining or Fine tuning for Indic language models
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/soketlabs/bhasha-wiki-translated.Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedVietnamese-nampdn-ai-tiny-webtext-gg-translatedyeji-bazi-translated-ko
██████╗ █████╗ ███████╗██╗ ████████╗██████╗ █████╗ ███╗ ██╗███████╗
██╔══██╗██╔══██╗╚══███╔╝██║ ╚══██╔══╝██╔══██╗██╔══██╗████╗ ██║██╔════╝
██████╔╝███████║ ███╔╝ ██║ ██║ ██████╔╝███████║██╔██╗ ██║███████╗
██╔══██╗██╔══██║ ███╔╝ ██║ ██║ ██╔══██╗██╔══██║██║╚██╗██║╚════██║
██████╔╝██║ ██║███████╗██║ ██║ ██║ ██║██║ ██║██║ ╚████║███████║
╚═════╝ ╚═╝ ╚═╝╚══════╝╚═╝ ╚═╝ ╚═╝ ╚═╝╚═╝ ╚═╝╚═╝ ╚═══╝╚══════╝
⚡ MASSIVE TRANSLATION CORPUS… See the full description on the dataset page: https://huggingface.co/datasets/tellang/yeji-bazi-translated-ko.marathi-alpaca-cleaned-translated
Marathi Alpaca Cleaned Translated
A Marathi translation of the 51,760-row Alpaca-Cleaned instruction-tuning dataset — Unsloth's hosted fork of yahma/alpaca-cleaned, which fixes hallucinations, empty outputs, and formatting errors found in the original Stanford Alpaca-52k dataset.
Translated using Meta's facebook/nllb-200-distilled-600M model. Built to reproduce and evaluate the Marathi instruction-tuning experiment from Khade et al., CHiPSAL 2025. The original paper translated… See the full description on the dataset page: https://huggingface.co/datasets/lubzo/marathi-alpaca-cleaned-translated.Dolci-Instruct-SFT-translated
Dolci-Instruct-SFT-translated (Swedish)
This dataset is a Swedish machine translation of the openeurollm/Dolci-Instruct-SFT-translated dataset, originally created as part of the OpenEuroLLM project.
Dataset details
Examples: 494,841 multi-turn conversations
Language: Swedish (sv-SE)
Format: Chat/messages format (id, messages)
License: Apache 2.0
Translation
All English source texts were machine-translated to Swedish using Google Gemma 3 27B-IT (w8a8_fp8… See the full description on the dataset page: https://huggingface.co/datasets/AI-Sweden-Models/Dolci-Instruct-SFT-translated.gsm8k-translated
Multilingual GSM8K Translations
This dataset contains machine-translated versions of GSM8K in these languages:
French (fr)
German (de)
Hindi (hi)
Dataset Structure
For each language, we provide the original GSM8K train and test splits:
train: 7,473 samples
test: 1,319 samples
Each sample consists of a question and an answer.
The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.FairytaleQA-translated-spanish
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Spanish machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-spanish.bluemoon-fandom-1-1-rp-jp-translated
bluemoon-fandom-1-1-rp-jp-translated
A subset of Squish42/bluemoon-fandom-1-1-rp-cleaned translated to Japanese using command-r-08-2024.
Misc. info
I used openrouter's api for inference with command-r-08-2024. Doing so is roughly 4x quicker than running the model locally, doesn't use up 95% of my vram, and doesn't make my 3090 as loud as my neighbours.
I decided to use command-r-08-2024 because it is completely uncensored for nsfw translation and provides translation… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/bluemoon-fandom-1-1-rp-jp-translated.translated-babylm-telugu
Translated BabyLM — Telugu (translated-babylm-telugu)
Dataset Description
This dataset is a Telugu translation of the English BabyLM 2026 corpus, produced using IndicTrans2, a state-of-the-art neural machine translation model developed by AI4Bharat for Indic languages. The dataset is intended for training and evaluating language models on Telugu, following the BabyLM challenge setup.
Translated by: IndicTrans2 (ai4bharat/indictrans2-en-indic-1B)
Source language:… See the full description on the dataset page: https://huggingface.co/datasets/pulipakav-1/translated-babylm-telugu.Open_o1_sft_Pro_translated_jp
概要
このデータセットはOpen_o1_sft_ProデータセットをQwen社のQwen2.5-14B-Instructを用いて日本語に翻訳したものになります。
テンプレート
テンプレートは以下です。
{"conversations": [{"role": "user", "content": "入力"}, {"role": "assistant", "thought": "思考",
"content": "出力"}, ...],
"id": id(整数),
"dataset": "元データセットの名前"}
ライセンス
ライセンスは元データセットに準じます。
謝辞
データセットの製作者様,Qwenの開発者様,計算資源を貸してくださったVolt mindの皆様に感謝します。
Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedFairytaleQA-translated-ptBR
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Brazilian Portuguese (pt-BR) machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-ptBR.hermes-3-dataset-ru-translated-prompts
Переведенные промты из hermes-3-dataset
Модель-переводчик Gemma-3-27b-it.
Переведены все промты.
Multi-turn промты переведены с учетом контекста англоязычного ответа.
Будет полезно для создания крупных русскоязычных инструктивных датасетов или Online RL.
Translated prompts from hermes-3-dataset
Translator model: Gemma-3-27b-it.
All prompts have been translated.
Multi-turn prompts were translated considering the context of the English response.
This will be useful… See the full description on the dataset page: https://huggingface.co/datasets/kristaller486/hermes-3-dataset-ru-translated-prompts.translated-babylm-hindi
Translated BabyLM — Hindi (translated-babylm-hindi)
Dataset Description
This dataset is a Hindi translation of the English BabyLM 2026 corpus, produced using IndicTrans2, a state-of-the-art neural machine translation model developed by AI4Bharat for Indic languages. The dataset is intended for training and evaluating language models on Hindi, following the BabyLM challenge setup.
Translated by: IndicTrans2 (ai4bharat/indictrans2-en-indic-1B)
Source language:… See the full description on the dataset page: https://huggingface.co/datasets/pulipakav-1/translated-babylm-hindi.s1K-1.1-Translated
s1K-1.1-Translated
This dataset contains translated versions of the s1K-1.1 dataset across multiple languages.
Languages
The dataset contains the following language subsets: Zh, Fr, Ja, Af, Th, Lv, Mr, Te, Sw, En
Translation Method
This dataset was created using Gemini 2.0 Flash for automatic translation.
Dataset Structure
Each language subset contains conversational data in the following format:
{
'conversations': [
{'from': 'human'… See the full description on the dataset page: https://huggingface.co/datasets/joshbarua/s1K-1.1-Translated.Vietnamese-nvidia-OpenMathInstruct-1-50k-gg-translatedFairytaleQA-translated-ptPT
Dataset Card for FairytaleQA-translated-ptPT
Dataset Summary
This repository contains the European Portuguese (pt-PT) machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-ptPT.multi-doc-qa-zh-translated
中文多文档QA数据集
从togethercomputer/Long-Data-Collections中的多文档QA任务,使用谷歌翻译机翻成中文得到。
任务:给定多个参考文档和一个问题,只有一个文档包含有用信息,模型需要根据参考文档回答问题,并指出哪个文档包含有用信息。
对于每个question,会提供几十或上百个文档片段,只有一个文档包含有用信息,gold_document_id表示含有有用信息的文档序号,注意文档是从1开始编号。
FairytaleQA-translated-french
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the French machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-french.Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedACL-SRW-2025
Dataset Components
The dataset is partitioned into three discrete tables stored in CSV or Parquet format:
Questions
Recipes
Evaluation Results
Each component is described in detail below.
Questions
area
domain
question_number
An integer index uniquely identifying each question inside the knowledge domain.
translation_method
English, Google Translate, GPT-3.5-Turbo, GPT-4o, Human
question
option_a, option_b, option_c, option_d
Recipes
area… See the full description on the dataset page: https://huggingface.co/datasets/Translated-MMLU-Blind-Review/ACL-SRW-2025.FairytaleQA-translated-romanian
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.FairytaleQA-translated-italian
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Italian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-italian.persian-qa-translated
Dataset Card for "persian-qa-translated"
More Information needed
