CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Helsinki-NLP /fineweb-edu-translated Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 36,704,000 documents with over 28 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages. In the v1.1 release, additional translations… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/fineweb-edu-translated.texttranslation1B<n<10B16 likes211k downloads5mo agoHugging Face02Helsinki-NLP /nemotron-cc-translated Helsinki-NLP/nemotron-cc-translated nemotron-cc-tanslated is a collection of automatically translated documents from nemotron-cc taken out of the high-quality subset. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 156,431,999 documents with over 70 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 2.4 trillion tokens and the translated documents are aligned across all languages. v1.1… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/nemotron-cc-translated.texttranslation1B<n<10B5 likes55k downloads5mo agoHugging Face03openeurollm /Dolci-Instruct-SFT-translatedtexttext-generation1M<n<10M3 likes1.5k downloads3mo agoHugging Face04openeurollm /Dolci-Think-SFT-translated Dolci-Think-SFT-translated Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations. Columns Each row is a translated conversation plus the result of a post-translation quality filter: id — source record id. messages — the translated conversation (list of {content, role}). filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.tabulartext-generation1M<n<10M0 likes847 downloads17h agoHugging Face05openeurollm /nemotron-cc-10K-sample-translated Translated Nemotron-cc-hq samples This dataset contains translated samples from https://huggingface.co/datasets/spyysalo/nemotron-cc-10K-sample Currently, the following are available, we will add other models and languages: Model Languages Gemma-3-4b-it ["Bulgarian", "Czech", "Danish", "German", "Estonian", "Finnish", "French", "Croatian", "Dutch"] EuroLLM-9B-Instruct ["Bulgarian", "Czech", "Danish", "German", "Greek", "Estonia", "Finnish", "French", "Irish"… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/nemotron-cc-10K-sample-translated.texttext-generation100K<n<1M1 likes174 downloads1y agoHugging Face06soketlabs /bhasha-wiki-translated Bhasha Wikipedia Translated Translated wikipedia articles Dataset Details Dataset is being updated Dataset Description We have translated 6.185 million English wikipedia articles into 6 Indic languages. The translations were done using IndicTrans2 model. Curated by: Soket AI labs Language(s) (NLP): Hindi, Bengali, Gujarati, Tamil, Kannada, Urdu License: cc-by-sa-4.0 Uses For pretraining or Fine tuning for Indic language models Dataset… See the full description on the dataset page: https://huggingface.co/datasets/soketlabs/bhasha-wiki-translated.texttext-generation100K<n<1M3 likes134 downloads2y agoHugging Face075CD-AI /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K8 likes108 downloads2y agoHugging Face085CD-AI /Vietnamese-nampdn-ai-tiny-webtext-gg-translatedtextquestion-answering1M<n<10M10 likes98 downloads3y agoHugging Face09tellang /yeji-bazi-translated-ko ██████╗ █████╗ ███████╗██╗ ████████╗██████╗ █████╗ ███╗ ██╗███████╗ ██╔══██╗██╔══██╗╚══███╔╝██║ ╚══██╔══╝██╔══██╗██╔══██╗████╗ ██║██╔════╝ ██████╔╝███████║ ███╔╝ ██║ ██║ ██████╔╝███████║██╔██╗ ██║███████╗ ██╔══██╗██╔══██║ ███╔╝ ██║ ██║ ██╔══██╗██╔══██║██║╚██╗██║╚════██║ ██████╔╝██║ ██║███████╗██║ ██║ ██║ ██║██║ ██║██║ ╚████║███████║ ╚═════╝ ╚═╝ ╚═╝╚══════╝╚═╝ ╚═╝ ╚═╝ ╚═╝╚═╝ ╚═╝╚═╝ ╚═══╝╚══════╝ ⚡ MASSIVE TRANSLATION CORPUS… See the full description on the dataset page: https://huggingface.co/datasets/tellang/yeji-bazi-translated-ko.texttext-generation100K<n<1M1 likes78 downloads8mo agoHugging Face10lubzo /marathi-alpaca-cleaned-translated Marathi Alpaca Cleaned Translated A Marathi translation of the 51,760-row Alpaca-Cleaned instruction-tuning dataset — Unsloth's hosted fork of yahma/alpaca-cleaned, which fixes hallucinations, empty outputs, and formatting errors found in the original Stanford Alpaca-52k dataset. Translated using Meta's facebook/nllb-200-distilled-600M model. Built to reproduce and evaluate the Marathi instruction-tuning experiment from Khade et al., CHiPSAL 2025. The original paper translated… See the full description on the dataset page: https://huggingface.co/datasets/lubzo/marathi-alpaca-cleaned-translated.texttext-generation10K<n<100K0 likes71 downloads6d agoHugging Face11AI-Sweden-Models /Dolci-Instruct-SFT-translated Dolci-Instruct-SFT-translated (Swedish) This dataset is a Swedish machine translation of the openeurollm/Dolci-Instruct-SFT-translated dataset, originally created as part of the OpenEuroLLM project. Dataset details Examples: 494,841 multi-turn conversations Language: Swedish (sv-SE) Format: Chat/messages format (id, messages) License: Apache 2.0 Translation All English source texts were machine-translated to Swedish using Google Gemma 3 27B-IT (w8a8_fp8… See the full description on the dataset page: https://huggingface.co/datasets/AI-Sweden-Models/Dolci-Instruct-SFT-translated.texttext-generation100K<n<1M0 likes70 downloads6mo agoHugging Face12math-across-languages /gsm8k-translated Multilingual GSM8K Translations This dataset contains machine-translated versions of GSM8K in these languages: French (fr) German (de) Hindi (hi) Dataset Structure For each language, we provide the original GSM8K train and test splits: train: 7,473 samples test: 1,319 samples Each sample consists of a question and an answer. The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.textquestion-answering10K<n<100K0 likes62 downloads3mo agoHugging Face13benjleite /FairytaleQA-translated-spanish Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Spanish machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-spanish.textquestion-answering10K<n<100K0 likes58 downloads1y agoHugging Face14joujiboi /bluemoon-fandom-1-1-rp-jp-translated bluemoon-fandom-1-1-rp-jp-translated A subset of Squish42/bluemoon-fandom-1-1-rp-cleaned translated to Japanese using command-r-08-2024. Misc. info I used openrouter's api for inference with command-r-08-2024. Doing so is roughly 4x quicker than running the model locally, doesn't use up 95% of my vram, and doesn't make my 3090 as loud as my neighbours. I decided to use command-r-08-2024 because it is completely uncensored for nsfw translation and provides translation… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/bluemoon-fandom-1-1-rp-jp-translated.tabulartext-generationn<1K3 likes53 downloads1y agoHugging Face15pulipakav-1 /translated-babylm-telugu Translated BabyLM — Telugu (translated-babylm-telugu) Dataset Description This dataset is a Telugu translation of the English BabyLM 2026 corpus, produced using IndicTrans2, a state-of-the-art neural machine translation model developed by AI4Bharat for Indic languages. The dataset is intended for training and evaluating language models on Telugu, following the BabyLM challenge setup. Translated by: IndicTrans2 (ai4bharat/indictrans2-en-indic-1B) Source language:… See the full description on the dataset page: https://huggingface.co/datasets/pulipakav-1/translated-babylm-telugu.texttext-generation10M<n<100M0 likes52 downloads5mo agoHugging Face16blastai /Open_o1_sft_Pro_translated_jp 概要 このデータセットはOpen_o1_sft_ProデータセットをQwen社のQwen2.5-14B-Instructを用いて日本語に翻訳したものになります。 テンプレート テンプレートは以下です。 {"conversations": [{"role": "user", "content": "入力"}, {"role": "assistant", "thought": "思考", "content": "出力"}, ...], "id": id(整数), "dataset": "元データセットの名前"} ライセンス ライセンスは元データセットに準じます。 謝辞 データセットの製作者様,Qwenの開発者様,計算資源を貸してくださったVolt mindの皆様に感謝します。 texttext-generation10K<n<100K9 likes45 downloads2y agoHugging Face175CD-AI /Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedtexttext-generation100K<n<1M4 likes44 downloads3y agoHugging Face18benjleite /FairytaleQA-translated-ptBR Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Brazilian Portuguese (pt-BR) machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-ptBR.textquestion-answering10K<n<100K3 likes42 downloads1y agoHugging Face19kristaller486 /hermes-3-dataset-ru-translated-prompts Переведенные промты из hermes-3-dataset Модель-переводчик Gemma-3-27b-it. Переведены все промты. Multi-turn промты переведены с учетом контекста англоязычного ответа. Будет полезно для создания крупных русскоязычных инструктивных датасетов или Online RL. Translated prompts from hermes-3-dataset Translator model: Gemma-3-27b-it. All prompts have been translated. Multi-turn prompts were translated considering the context of the English response. This will be useful… See the full description on the dataset page: https://huggingface.co/datasets/kristaller486/hermes-3-dataset-ru-translated-prompts.texttext-generation100K<n<1M2 likes42 downloads1y agoHugging Face20pulipakav-1 /translated-babylm-hindi Translated BabyLM — Hindi (translated-babylm-hindi) Dataset Description This dataset is a Hindi translation of the English BabyLM 2026 corpus, produced using IndicTrans2, a state-of-the-art neural machine translation model developed by AI4Bharat for Indic languages. The dataset is intended for training and evaluating language models on Hindi, following the BabyLM challenge setup. Translated by: IndicTrans2 (ai4bharat/indictrans2-en-indic-1B) Source language:… See the full description on the dataset page: https://huggingface.co/datasets/pulipakav-1/translated-babylm-hindi.texttext-generation10M<n<100M0 likes42 downloads5mo agoHugging Face21joshbarua /s1K-1.1-Translated s1K-1.1-Translated This dataset contains translated versions of the s1K-1.1 dataset across multiple languages. Languages The dataset contains the following language subsets: Zh, Fr, Ja, Af, Th, Lv, Mr, Te, Sw, En Translation Method This dataset was created using Gemini 2.0 Flash for automatic translation. Dataset Structure Each language subset contains conversational data in the following format: { 'conversations': [ {'from': 'human'… See the full description on the dataset page: https://huggingface.co/datasets/joshbarua/s1K-1.1-Translated.texttext-generation10K<n<100K3 likes41 downloads1y agoHugging Face225CD-AI /Vietnamese-nvidia-OpenMathInstruct-1-50k-gg-translatedtexttext-generation10K<n<100K7 likes39 downloads3y agoHugging Face23benjleite /FairytaleQA-translated-ptPT Dataset Card for FairytaleQA-translated-ptPT Dataset Summary This repository contains the European Portuguese (pt-PT) machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-ptPT.textquestion-answering10K<n<100K0 likes37 downloads1y agoHugging Face24yuyijiong /multi-doc-qa-zh-translated 中文多文档QA数据集 从togethercomputer/Long-Data-Collections中的多文档QA任务,使用谷歌翻译机翻成中文得到。 任务:给定多个参考文档和一个问题,只有一个文档包含有用信息,模型需要根据参考文档回答问题,并指出哪个文档包含有用信息。 对于每个question,会提供几十或上百个文档片段,只有一个文档包含有用信息,gold_document_id表示含有有用信息的文档序号,注意文档是从1开始编号。 texttext-generation10K<n<100K11 likes35 downloads1y agoHugging Face25benjleite /FairytaleQA-translated-french Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the French machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-french.textquestion-answering10K<n<100K1 likes35 downloads1y agoHugging Face26ChaosAIVision /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K0 likes34 downloads9mo agoHugging Face27Translated-MMLU-Blind-Review /ACL-SRW-2025 Dataset Components The dataset is partitioned into three discrete tables stored in CSV or Parquet format: Questions Recipes Evaluation Results Each component is described in detail below. Questions area domain question_number An integer index uniquely identifying each question inside the knowledge domain. translation_method English, Google Translate, GPT-3.5-Turbo, GPT-4o, Human question option_a, option_b, option_c, option_d Recipes area… See the full description on the dataset page: https://huggingface.co/datasets/Translated-MMLU-Blind-Review/ACL-SRW-2025.tabularquestion-answering100K<n<1M0 likes32 downloads1y agoHugging Face28benjleite /FairytaleQA-translated-romanian Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Romanian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-romanian.textquestion-answering10K<n<100K1 likes31 downloads1y agoHugging Face29benjleite /FairytaleQA-translated-italian Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Italian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-italian.textquestion-answering10K<n<100K0 likes30 downloads1y agoHugging Face30pourmand1376 /persian-qa-translated Dataset Card for "persian-qa-translated" More Information needed textquestion-answering100K<n<1M4 likes28 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.