datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-edu-translated
Helsinki-NLP/fineweb-edu-translated
fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu.
Translations are based on OPUS-MT and HPLT-MT models.
The data in v1.0 covers 36,704,000 documents with over 28 billion space-searated tokens of English data translated into 36 languages.
The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages.
In the v1.1 release, additional translations… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/fineweb-edu-translated.nemotron-cc-translated
Helsinki-NLP/nemotron-cc-translated
nemotron-cc-tanslated is a collection of automatically translated documents from nemotron-cc taken out of the high-quality subset.
Translations are based on OPUS-MT and HPLT-MT models.
The data in v1.0 covers 156,431,999 documents with over 70 billion space-searated tokens of English data translated into 36 languages.
The total v1.0 data set includes over 2.4 trillion tokens and the translated documents are aligned across all languages.
v1.1… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/nemotron-cc-translated.human_translated_arabic_mmluDolci-Instruct-SFT-translatedtranslated-german-english-asr
Translated German-English ASR Dataset
A large-scale, multi-source German speech dataset with paired English translations, designed for training and evaluating German Automatic Speech Recognition (ASR), Speech Translation, and Text-to-Speech (TTS) systems. This dataset is a curated mixture of well-established open-source German and multilingual speech corpora, all unified under a common schema with German audio, original German transcriptions, and English translations.… See the full description on the dataset page: https://huggingface.co/datasets/aman4014/translated-german-english-asr.datacomp-medium-pool-translatedsizefetish-jp2cn-translated-textDolci-Think-SFT-translated
Dolci-Think-SFT-translated
Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations.
Columns
Each row is a translated conversation plus the result of a post-translation quality filter:
id — source record id.
messages — the translated conversation (list of {content, role}).
filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.sizefetish-jp2cn-sakura-translated-collectionTranslated-GSM8Kmegamatt-translated-ITtranslated_Zyda_2AlGhafa-Arabic-LLM-Benchmark-Translatedli-vdr-translatedNemotron-CC-Translated-Diverse-QA-itVietnamese-THUIR-T2Ranking-gg-translated
📚 5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated
📝 Overview
Vietnamese-THUIR-T2Ranking-gg-translated is a large-scale dataset for passage ranking in Vietnamese.It is translated from the original THUIR/T2Ranking [1] using Google Translate, inspired by the approach of mMARCO [2].The dataset aims to provide a large-scale dataset for research and applications in Information Retrieval (IR) in Vietnamese.
In IR, passage ranking is an essential and challenging task… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated.cranemath-translated-ITyoruba_audio_translatedThis is a copy of odunola/Yoruba_translate_preprocessed, the only difference is, it's already splitted into train & test. Awesome credits to her, her license applies too.
Google-Translated_Turkish_GPQA_Datasetragtruth-translated-hallucinations
RAGTruth Translated Hallucinations
Multilingual machine translation of
RAGTruth into 31 European languages,
preserving RAGTruth's word-level hallucination-span annotations. RAGTruth is a corpus of
LLM responses to retrieval-augmented generation (RAG) tasks in which humans marked the
exact spans that are hallucinated (unsupported by, or contradicting, the provided
context). Here both the RAG prompt and the response are translated into each target
language, and the annotated… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/ragtruth-translated-hallucinations.gsm8k-translated-spanish
Dataset Card for GSM8K (Spanish Version)
Dataset Summary
GSM8K (Grade School Math 8K) is a dataset containing 8.5K high-quality math word problems designed to assess multi-step mathematical reasoning in language models. This version is a Spanish translation of the original openai/gsm8k dataset, preserving the structure and objectives of the original English dataset.
Key features of the dataset:
Problems require between 2 and 8 steps to solve.
Solutions involve sequences… See the full description on the dataset page: https://huggingface.co/datasets/ericrisco/gsm8k-translated-spanish.Tumbuka_Text_Corpus_Translated_Gutenberg
Tumbuka Text Corpus - Translated Gutenberg
Dataset Description
This dataset contains a large-scale collection of Tumbuka text, primarily consisting of machine-translated literary works from the Project Gutenberg library. It is designed to support Natural Language Processing (NLP) research for Tumbuka, a Bantu language spoken in Malawi, Zambia, and Tanzania.
Dataset Summary
Language: Tumbuka (tum)
Source: Project Gutenberg
Content: Translated… See the full description on the dataset page: https://huggingface.co/datasets/Mwanzau/Tumbuka_Text_Corpus_Translated_Gutenberg.PROMISE_NFR_translated
Dataset Summary
Published version of PROMISE NFR translated to Spanish used for paper 'Requirements Classification Using FastText and BETO in Spanish Documents'
Languages
Spanish
Dataset Structure
Data Fields
Project: Project's Identifier.
Requirement: Description of the software requirement.
Label: Label of the requirement: F (functional requirement) and NF (non-functional requirement).
Dataset Creation
Initial Data Collection and… See the full description on the dataset page: https://huggingface.co/datasets/MariaIsabel/PROMISE_NFR_translated.laion2B-multi-joined-translated-to-en-smols2orc-citation-pairs-translated-nlThis is a Dutch version of the S2ORC: The Semantic Scholar Open Research Corpus. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
Vietnamese-lmms-lab-LLaVA-Video-178K-gg-translated
Dataset Card for 5CD-AI/Vietnamese-lmms-lab-LLaVA-Video-178K-gg-translated
This translated dataset includes:
LLaVA-Video-178K: 178,509 caption entries, 960,791 open-ended QA (question and answer) items, and 196,198 multiple-choice QA items.
The video source of the original dataset is in this repo: lmms-lab/LLaVA-Video-178K
dolly-machine-translated-v2
Dolly Machine Translated (v2)
Dataset Description
Dolly Machine Translated (v2) is a multilingual evaluation-only release built from a curated subset of Databricks Dolly 15k prompts. It contains the original English prompts plus machine translations in 66 non-English languages, with the English source prompts included as the en config for reference.
Each language is provided as a separate config (subset). All language codes use ISO 639-1 two-letter codes. Each row carries… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/dolly-machine-translated-v2.gsm8k-translatedxrisawoz-translatednli-rus-translated-v2021
Dataset Card for "nli-rus-translated-v2021"
This dataset was introduced in the Habr post
"Нейросети для Natural Language Inference (NLI): логические умозаключения на русском языке".
It is composed from various English NLI datasets automatically translated into Russian.
Here are the sizes of the source datasets included into different splits:
source
train
dev
test
add_one_rte
4991
387
0
anli_r1
16946
1000
1000
anli_r2
45460
1000
1000
anli_r3
100459
1200
1200
copa… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/nli-rus-translated-v2021.
