CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Helsinki-NLP /fineweb-edu-translated Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 36,704,000 documents with over 28 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages. In the v1.1 release, additional translations… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/fineweb-edu-translated.texttranslation1B<n<10B16 likes206k downloads5mo agoHugging Face02Helsinki-NLP /nemotron-cc-translated Helsinki-NLP/nemotron-cc-translated nemotron-cc-tanslated is a collection of automatically translated documents from nemotron-cc taken out of the high-quality subset. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 156,431,999 documents with over 70 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 2.4 trillion tokens and the translated documents are aligned across all languages. v1.1… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/nemotron-cc-translated.texttranslation1B<n<10B5 likes55k downloads5mo agoHugging Face03fromozu /ebook-translate-queuetextn<1K3 likes8.5k downloads5mo agoHugging Face04MBZUAI /human_translated_arabic_mmlutext10K<n<100K4 likes2.4k downloads2y agoHugging Face05openeurollm /Dolci-Instruct-SFT-translatedtexttext-generation1M<n<10M3 likes1.6k downloads3mo agoHugging Face06qvac /TranslatePsy-AfriSLM-Synthetic-Mix TranslatePsy-AfriSLM Synthetic Mix TranslatePsy-AfriSLM Synthetic Mix is a quality-filtered synthetic parallel corpus for machine translation between English and 19 Sub-Saharan African languages. It contains 215,653,192 bidirectional training examples and was selected as the primary African translation component used to post-train the TranslatePsy-AfriSLM model family. The dataset accompanies the EMNLP 2026 paper TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource… See the full description on the dataset page: https://huggingface.co/datasets/qvac/TranslatePsy-AfriSLM-Synthetic-Mix.texttranslation100M<n<1B0 likes1.3k downloads28d agoHugging Face07aman4014 /translated-german-english-asr Translated German-English ASR Dataset A large-scale, multi-source German speech dataset with paired English translations, designed for training and evaluating German Automatic Speech Recognition (ASR), Speech Translation, and Text-to-Speech (TTS) systems. This dataset is a curated mixture of well-established open-source German and multilingual speech corpora, all unified under a common schema with German audio, original German transcriptions, and English translations.… See the full description on the dataset page: https://huggingface.co/datasets/aman4014/translated-german-english-asr.audioautomatic-speech-recognition1M<n<10M4 likes1.1k downloads5mo agoHugging Face08thaottn /datacomp-medium-pool-translatedimage100M<n<1B0 likes1k downloads1y agoHugging Face09AkiraChisaka /sizefetish-jp2cn-translated-texttext10K<n<100K7 likes977 downloads2y agoHugging Face10ai4bharat /wiki-translatetext1M<n<10M8 likes976 downloads2y agoHugging Face11ChavyvAkvar /Translate-Preparedtext100K<n<1M0 likes798 downloads1y agoHugging Face12AkiraChisaka /sizefetish-jp2cn-sakura-translated-collectiontext1M<n<10M3 likes789 downloads1y agoHugging Face13llama-lang-adapt /Translated-GSM8Ktext100K<n<1M0 likes761 downloads2y agoHugging Face14openeurollm /Dolci-Think-SFT-translated Dolci-Think-SFT-translated Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations. Columns Each row is a translated conversation plus the result of a post-translation quality filter: id — source record id. messages — the translated conversation (list of {content, role}). filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.tabulartext-generation1M<n<10M0 likes693 downloads9d agoHugging Face15OALL /AlGhafa-Arabic-LLM-Benchmark-Translatedtabular10K<n<100K2 likes646 downloads2y agoHugging Face16Mwanzau /Tumbuka_Text_Corpus_Translated_Gutenberg Tumbuka Text Corpus - Translated Gutenberg Dataset Description This dataset contains a large-scale collection of Tumbuka text, primarily consisting of machine-translated literary works from the Project Gutenberg library. It is designed to support Natural Language Processing (NLP) research for Tumbuka, a Bantu language spoken in Malawi, Zambia, and Tanzania. Dataset Summary Language: Tumbuka (tum) Source: Project Gutenberg Content: Translated… See the full description on the dataset page: https://huggingface.co/datasets/Mwanzau/Tumbuka_Text_Corpus_Translated_Gutenberg.text100M<n<1B0 likes619 downloads2mo agoHugging Face17FinancialSupport /Nemotron-CC-Translated-Diverse-QA-ittext10M<n<100M0 likes600 downloads2mo agoHugging Face18dequeirozrodriguez /translated_Zyda_2text1M<n<10M0 likes597 downloads7mo agoHugging Face19ItalianNarratives /megamatt-translated-ITtext100K<n<1M0 likes576 downloads16d agoHugging Face20matiss /P3-Latvian-translategemma-27bThis is an automatically translated version of P3 (Public Pool of Prompts) using translategemma-27b. Languages The data in P3-Latvian-Full are in Latvian (BCP-47 lv). Dataset Structure Data Instances An example of "train" looks as follows: { 'answer_choices': ['mobilais tālrunis', 'televīzija', 'ledusskapis', 'lidmašīna'], 'inputs_pretokenized': 'Kura tehnoloģija tika izstrādāta pavisam nesen? Iespējas: - mobilais tālrunis - televizors - ledusskapis -… See the full description on the dataset page: https://huggingface.co/datasets/matiss/P3-Latvian-translategemma-27b.text100K<n<1M0 likes564 downloads7mo agoHugging Face215CD-AI /Vietnamese-THUIR-T2Ranking-gg-translated 📚 5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated 📝 Overview Vietnamese-THUIR-T2Ranking-gg-translated is a large-scale dataset for passage ranking in Vietnamese.It is translated from the original THUIR/T2Ranking [1] using Google Translate, inspired by the approach of mMARCO [2].The dataset aims to provide a large-scale dataset for research and applications in Information Retrieval (IR) in Vietnamese. In IR, passage ranking is an essential and challenging task… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated.tabulartext-retrieval100M<n<1B22 likes535 downloads1y agoHugging Face22vidore /li-vdr-translatedimage100K<n<1M0 likes514 downloads1y agoHugging Face23ItalianNarratives /cranemath-translated-ITtext100K<n<1M0 likes504 downloads16d agoHugging Face24bytel0rd /yoruba_audio_translatedThis is a copy of odunola/Yoruba_translate_preprocessed, the only difference is, it's already splitted into train & test. Awesome credits to her, her license applies too. audiotranslation10K<n<100K3 likes490 downloads2y agoHugging Face25AdnanElAssadi /Google-Translated_Turkish_GPQA_Datasettabular1K<n<10K0 likes485 downloads2y agoHugging Face26youjunhyeok /smoltalk-ko-translate 번역 결과에 토큰이 반복된 결과들이 포함되어 있습니다. 필터링 후 재업로드 하겠습니다. Z 알고리즘을 사용해 결과를 필터링 하였으며 {subset}_filtered 로 업로드하였습니다. 필터링 후 결과 subset 전 후 split/train 4205413 4162254 split/test 221249 218830 merge/train 1043917 1034473 merge/test 54948 54430 HuggingFaceTB/smoltalk 데이터셋의 subset:all을 nayohan/llama3-instrucTrans-enko-8b 모델을 사용해 번역했습니다. 원본의 messages 중 4096 token 이 넘어가는 content가 있다면 해당 레코드는 번역하지 않았습니다. texttext-generation10M<n<100M5 likes410 downloads2y agoHugging Face27alexandrainst /ragtruth-translated-hallucinations RAGTruth Translated Hallucinations Multilingual machine translation of RAGTruth into 31 European languages, preserving RAGTruth's word-level hallucination-span annotations. RAGTruth is a corpus of LLM responses to retrieval-augmented generation (RAG) tasks in which humans marked the exact spans that are hallucinated (unsupported by, or contradicting, the provided context). Here both the RAG prompt and the response are translated into each target language, and the annotated… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/ragtruth-translated-hallucinations.texttoken-classification100K<n<1M2 likes400 downloads1mo agoHugging Face28lewington /laion2B-multi-joined-translated-to-en-smolimage10M<n<100M2 likes390 downloads2y agoHugging Face29KrorngAI /DDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translateDisclaimer: The original dataset can be found here. It is published by Digital Divide Data Cambodia (DDD-Cambodia). License: Khmer ASR Cultural Dataset's license is Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0). Please attribute Digital Divide Data if you use this dataset in any way. Objective of this dataset Add English translation: a new column en_translate is added to the original dataset (only from parquet 000 to 159 of the original… See the full description on the dataset page: https://huggingface.co/datasets/KrorngAI/DDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translate.audioautomatic-speech-recognition10K<n<100K0 likes375 downloads2mo agoHugging Face30ericrisco /gsm8k-translated-spanish Dataset Card for GSM8K (Spanish Version) Dataset Summary GSM8K (Grade School Math 8K) is a dataset containing 8.5K high-quality math word problems designed to assess multi-step mathematical reasoning in language models. This version is a Spanish translation of the original openai/gsm8k dataset, preserving the structure and objectives of the original English dataset. Key features of the dataset: Problems require between 2 and 8 steps to solve. Solutions involve sequences… See the full description on the dataset page: https://huggingface.co/datasets/ericrisco/gsm8k-translated-spanish.text1K<n<10K1 likes374 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.