datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vietnamese-Locutusque-function-calling-chatml-gg-translatedVietnamese-alpaca-gpt4-gg-translatedVietnamese-Salesforce-xlam-function-calling-60k-gg-translatedVietnamese-nampdn-ai-tiny-webtext-gg-translateds1_translated
S1K Multilingual Translation Dataset
This dataset contains translations of the simplescaling/s1K dataset into 20 languages.
Languages Included
High-Resource Languages
Chinese (Simplified) - zh
Spanish - es
French - fr
German - de
Japanese - ja
Arabic - ar
Russian - ru
Portuguese - pt
Medium-Resource Languages
Korean - ko
Vietnamese - vi
Thai - th
Polish - pl
Dutch - nl
Turkish - tr
Low-Resource Languages
Swahili - sw
Bengali - bn
Urdu… See the full description on the dataset page: https://huggingface.co/datasets/JiayiHe/s1_translated.Vietnamese-395k-meta-math-MetaMathQA-gg-translatedVietnamese-ShareGPT4Video-ShareGPT4Video-gg-translatedVietnamese-ShareGPT4Vision-gg-translatedVietnamese-TriviaQA-RC-gg-translatedVietnamese-liuhaotian-llava_v1_5_mix665k-gg-translatedVietnamese-ComplexWebQuestions-gg-translatedgsm8k-translated
Multilingual GSM8K Translations
This dataset contains machine-translated versions of GSM8K in these languages:
French (fr)
German (de)
Hindi (hi)
Dataset Structure
For each language, we provide the original GSM8K train and test splits:
train: 7,473 samples
test: 1,319 samples
Each sample consists of a question and an answer.
The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.FairytaleQA-translated-spanish
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Spanish machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-spanish.Vietnamese-OpenGVLab-ShareGPT-4o-gg-translatedTranslated
MedInjection-FR — Translated Subset 🌍
Summary
The Translated component of MedInjection-FR adapts large-scale English biomedical instruction datasets into French through high-quality automatic translation.It represents the most extensive part of the collection, comprising 416 401 instruction–response pairs, and provides a bridge between English biomedical resources and French medical instruction tuning.
This subset was designed to ensure broad domain coverage while… See the full description on the dataset page: https://huggingface.co/datasets/MedInjection/Translated.semeval-2016-absa-reviews-english-translated-stanford-alpaca
Dataset Card for Dataset Name
Derived from eastwind/semeval-2016-absa-reviews-arabic using Helsinki-NLP/opus-mt-tc-big-ar-en
Vietnamese-meta-math-MetaMathQA-40K-gg-translatedVietnamese-microsoft-orca-math-word-problems-200k-gg-translatedVietnamese-cosmos-qa-gg-translatedVietnamese-LLaVA-Instruct-150K-gg-translatedVietnamese-Openorca-Multiplechoice-gg-translatedVietnamese-beyond-rlhf-reward-single-round-gg-translatedFairytaleQA-translated-ptBR
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Brazilian Portuguese (pt-BR) machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-ptBR.s1K-1.1-Translated
s1K-1.1-Translated
This dataset contains translated versions of the s1K-1.1 dataset across multiple languages.
Languages
The dataset contains the following language subsets: Zh, Fr, Ja, Af, Th, Lv, Mr, Te, Sw, En
Translation Method
This dataset was created using Gemini 2.0 Flash for automatic translation.
Dataset Structure
Each language subset contains conversational data in the following format:
{
'conversations': [
{'from': 'human'… See the full description on the dataset page: https://huggingface.co/datasets/joshbarua/s1K-1.1-Translated.Vietnamese-nvidia-OpenMathInstruct-1-50k-gg-translatedFairytaleQA-translated-ptPT
Dataset Card for FairytaleQA-translated-ptPT
Dataset Summary
This repository contains the European Portuguese (pt-PT) machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-ptPT.multi-doc-qa-zh-translated
中文多文档QA数据集
从togethercomputer/Long-Data-Collections中的多文档QA任务,使用谷歌翻译机翻成中文得到。
任务:给定多个参考文档和一个问题,只有一个文档包含有用信息,模型需要根据参考文档回答问题,并指出哪个文档包含有用信息。
对于每个question,会提供几十或上百个文档片段,只有一个文档包含有用信息,gold_document_id表示含有有用信息的文档序号,注意文档是从1开始编号。
FairytaleQA-translated-french
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the French machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-french.MedQA-USMLE-back-translatedMedQA dataset perturbed using back-translation technique with BAT
Vietnamese-Salesforce-xlam-function-calling-60k-gg-translated
