CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Helsinki-NLP /fineweb-edu-translated Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 36,704,000 documents with over 28 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages. In the v1.1 release, additional translations… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/fineweb-edu-translated.texttranslation1B<n<10B16 likes211k downloads5mo agoHugging Face02Helsinki-NLP /nemotron-cc-translated Helsinki-NLP/nemotron-cc-translated nemotron-cc-tanslated is a collection of automatically translated documents from nemotron-cc taken out of the high-quality subset. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 156,431,999 documents with over 70 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 2.4 trillion tokens and the translated documents are aligned across all languages. v1.1… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/nemotron-cc-translated.texttranslation1B<n<10B5 likes55k downloads5mo agoHugging Face03MBZUAI /human_translated_arabic_mmlutext10K<n<100K4 likes1.9k downloads2y agoHugging Face04openeurollm /Dolci-Instruct-SFT-translatedtexttext-generation1M<n<10M3 likes1.5k downloads3mo agoHugging Face05aman4014 /translated-german-english-asr Translated German-English ASR Dataset A large-scale, multi-source German speech dataset with paired English translations, designed for training and evaluating German Automatic Speech Recognition (ASR), Speech Translation, and Text-to-Speech (TTS) systems. This dataset is a curated mixture of well-established open-source German and multilingual speech corpora, all unified under a common schema with German audio, original German transcriptions, and English translations.… See the full description on the dataset page: https://huggingface.co/datasets/aman4014/translated-german-english-asr.audioautomatic-speech-recognition1M<n<10M4 likes1.1k downloads5mo agoHugging Face06thaottn /datacomp-medium-pool-translatedimage100M<n<1B0 likes1k downloads1y agoHugging Face07AkiraChisaka /sizefetish-jp2cn-translated-texttext10K<n<100K7 likes940 downloads2y agoHugging Face08openeurollm /Dolci-Think-SFT-translated Dolci-Think-SFT-translated Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations. Columns Each row is a translated conversation plus the result of a post-translation quality filter: id — source record id. messages — the translated conversation (list of {content, role}). filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.tabulartext-generation1M<n<10M0 likes847 downloads12h agoHugging Face09AkiraChisaka /sizefetish-jp2cn-sakura-translated-collectiontext1M<n<10M3 likes796 downloads1y agoHugging Face10llama-lang-adapt /Translated-GSM8Ktext100K<n<1M0 likes773 downloads2y agoHugging Face11ItalianNarratives /megamatt-translated-ITtext100K<n<1M0 likes724 downloads18d agoHugging Face12dequeirozrodriguez /translated_Zyda_2text1M<n<10M0 likes646 downloads7mo agoHugging Face13OALL /AlGhafa-Arabic-LLM-Benchmark-Translatedtabular10K<n<100K2 likes622 downloads2y agoHugging Face14vidore /li-vdr-translatedimage100K<n<1M0 likes586 downloads1y agoHugging Face15FinancialSupport /Nemotron-CC-Translated-Diverse-QA-ittext10M<n<100M0 likes571 downloads2mo agoHugging Face165CD-AI /Vietnamese-THUIR-T2Ranking-gg-translated 📚 5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated 📝 Overview Vietnamese-THUIR-T2Ranking-gg-translated is a large-scale dataset for passage ranking in Vietnamese.It is translated from the original THUIR/T2Ranking [1] using Google Translate, inspired by the approach of mMARCO [2].The dataset aims to provide a large-scale dataset for research and applications in Information Retrieval (IR) in Vietnamese. In IR, passage ranking is an essential and challenging task… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated.tabulartext-retrieval100M<n<1B22 likes558 downloads1y agoHugging Face17ItalianNarratives /cranemath-translated-ITtext100K<n<1M0 likes542 downloads18d agoHugging Face18bytel0rd /yoruba_audio_translatedThis is a copy of odunola/Yoruba_translate_preprocessed, the only difference is, it's already splitted into train & test. Awesome credits to her, her license applies too. audiotranslation10K<n<100K3 likes493 downloads2y agoHugging Face19AdnanElAssadi /Google-Translated_Turkish_GPQA_Datasettabular1K<n<10K0 likes485 downloads2y agoHugging Face20alexandrainst /ragtruth-translated-hallucinations RAGTruth Translated Hallucinations Multilingual machine translation of RAGTruth into 31 European languages, preserving RAGTruth's word-level hallucination-span annotations. RAGTruth is a corpus of LLM responses to retrieval-augmented generation (RAG) tasks in which humans marked the exact spans that are hallucinated (unsupported by, or contradicting, the provided context). Here both the RAG prompt and the response are translated into each target language, and the annotated… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/ragtruth-translated-hallucinations.texttoken-classification100K<n<1M2 likes395 downloads1mo agoHugging Face21ericrisco /gsm8k-translated-spanish Dataset Card for GSM8K (Spanish Version) Dataset Summary GSM8K (Grade School Math 8K) is a dataset containing 8.5K high-quality math word problems designed to assess multi-step mathematical reasoning in language models. This version is a Spanish translation of the original openai/gsm8k dataset, preserving the structure and objectives of the original English dataset. Key features of the dataset: Problems require between 2 and 8 steps to solve. Solutions involve sequences… See the full description on the dataset page: https://huggingface.co/datasets/ericrisco/gsm8k-translated-spanish.text1K<n<10K1 likes376 downloads1y agoHugging Face22Mwanzau /Tumbuka_Text_Corpus_Translated_Gutenberg Tumbuka Text Corpus - Translated Gutenberg Dataset Description This dataset contains a large-scale collection of Tumbuka text, primarily consisting of machine-translated literary works from the Project Gutenberg library. It is designed to support Natural Language Processing (NLP) research for Tumbuka, a Bantu language spoken in Malawi, Zambia, and Tanzania. Dataset Summary Language: Tumbuka (tum) Source: Project Gutenberg Content: Translated… See the full description on the dataset page: https://huggingface.co/datasets/Mwanzau/Tumbuka_Text_Corpus_Translated_Gutenberg.text100M<n<1B0 likes367 downloads2mo agoHugging Face23MariaIsabel /PROMISE_NFR_translated Dataset Summary Published version of PROMISE NFR translated to Spanish used for paper 'Requirements Classification Using FastText and BETO in Spanish Documents' Languages Spanish Dataset Structure Data Fields Project: Project's Identifier. Requirement: Description of the software requirement. Label: Label of the requirement: F (functional requirement) and NF (non-functional requirement). Dataset Creation Initial Data Collection and… See the full description on the dataset page: https://huggingface.co/datasets/MariaIsabel/PROMISE_NFR_translated.texttext-classificationn<1K0 likes356 downloads3y agoHugging Face24lewington /laion2B-multi-joined-translated-to-en-smolimage10M<n<100M2 likes306 downloads2y agoHugging Face25NetherlandsForensicInstitute /s2orc-citation-pairs-translated-nlThis is a Dutch version of the S2ORC: The Semantic Scholar Open Research Corpus. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity10M<n<100M0 likes284 downloads2y agoHugging Face265CD-AI /Vietnamese-lmms-lab-LLaVA-Video-178K-gg-translated Dataset Card for 5CD-AI/Vietnamese-lmms-lab-LLaVA-Video-178K-gg-translated This translated dataset includes: LLaVA-Video-178K: 178,509 caption entries, 960,791 open-ended QA (question and answer) items, and 196,198 multiple-choice QA items. The video source of the original dataset is in this repo: lmms-lab/LLaVA-Video-178K textvisual-question-answering1M<n<10M1 likes273 downloads2y agoHugging Face27CohereLabs /dolly-machine-translated-v2 Dolly Machine Translated (v2) Dataset Description Dolly Machine Translated (v2) is a multilingual evaluation-only release built from a curated subset of Databricks Dolly 15k prompts. It contains the original English prompts plus machine translations in 66 non-English languages, with the English source prompts included as the en config for reference. Each language is provided as a separate config (subset). All language codes use ISO 639-1 two-letter codes. Each row carries… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/dolly-machine-translated-v2.text10K<n<100K2 likes261 downloads5mo agoHugging Face28Sara237 /gsm8k-translatedtext10K<n<100K0 likes222 downloads2y agoHugging Face29DeepPavlov /xrisawoz-translatedtext10K<n<100K0 likes210 downloads2mo agoHugging Face30cointegrated /nli-rus-translated-v2021 Dataset Card for "nli-rus-translated-v2021" This dataset was introduced in the Habr post "Нейросети для Natural Language Inference (NLI): логические умозаключения на русском языке". It is composed from various English NLI datasets automatically translated into Russian. Here are the sizes of the source datasets included into different splits: source train dev test add_one_rte 4991 387 0 anli_r1 16946 1000 1000 anli_r2 45460 1000 1000 anli_r3 100459 1200 1200 copa… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/nli-rus-translated-v2021.tabulartext-classification1M<n<10M1 likes207 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.