datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
megamatt-translated-ITcranemath-translated-ITs2orc-citation-pairs-translated-nlThis is a Dutch version of the S2ORC: The Semantic Scholar Open Research Corpus. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
Vietnamese-lmms-lab-LLaVA-Video-178K-gg-translated
Dataset Card for 5CD-AI/Vietnamese-lmms-lab-LLaVA-Video-178K-gg-translated
This translated dataset includes:
LLaVA-Video-178K: 178,509 caption entries, 960,791 open-ended QA (question and answer) items, and 196,198 multiple-choice QA items.
The video source of the original dataset is in this repo: lmms-lab/LLaVA-Video-178K
african-translated-alpaca
Citation
Please cite the stanford_alpaca project
@misc{alpaca,
author = {Rohan Taori and Ishaan Gulrajani and Tianyi Zhang and Yann Dubois and Xuechen Li and Carlos Guestrin and Percy Liang and Tatsunori B. Hashimoto },
title = {Stanford Alpaca: An Instruction-following LLaMA model},
year = {2023},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/tatsu-lab/stanford_alpaca}},
}
Vietnamese-alpaca-gpt4-gg-translatedVietnamese-Salesforce-xlam-function-calling-60k-gg-translatedVietnamese-nampdn-ai-tiny-webtext-gg-translatedVietnamese-395k-meta-math-MetaMathQA-gg-translatedtranslated-cnn-dailymailtranslated-xwikisVietnamese-ShareGPT4Vision-gg-translateddolmino-math-translated-ITVietnamese-liuhaotian-llava_v1_5_mix665k-gg-translatedVietnamese-Intel-orca_dpo_pairs-gg-translatedFairytaleQA-translated-spanish
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Spanish machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-spanish.squad_v2_french_translatedUsing Google Translation, we have translated SQuAD 2.0 dataset into multiple languages.
Here is the translated dataset of SQuAD 2.0 in French language.
Shared by Pragnakalp Techlabs
tinymath-pot-translated-ITtranslated-glaive-function-callTranslated
MedInjection-FR — Translated Subset 🌍
Summary
The Translated component of MedInjection-FR adapts large-scale English biomedical instruction datasets into French through high-quality automatic translation.It represents the most extensive part of the collection, comprising 416 401 instruction–response pairs, and provides a bridge between English biomedical resources and French medical instruction tuning.
This subset was designed to ensure broad domain coverage while… See the full description on the dataset page: https://huggingface.co/datasets/MedInjection/Translated.open_math_instruct_v2_translated_african_languagesThis is a set of 41k nvidia/OpenMathInstruct-2 questions translated into 9 African languages using Azure/GPT-4o.
We shuffle the dataset and then randomly sample a question without replacement, and then equally sample a language and then we translate the question and answer to that language.
Vietnamese-meta-math-MetaMathQA-40K-gg-translatedOpen_o1_sft_Pro_translated_jp
概要
このデータセットはOpen_o1_sft_ProデータセットをQwen社のQwen2.5-14B-Instructを用いて日本語に翻訳したものになります。
テンプレート
テンプレートは以下です。
{"conversations": [{"role": "user", "content": "入力"}, {"role": "assistant", "thought": "思考",
"content": "出力"}, ...],
"id": id(整数),
"dataset": "元データセットの名前"}
ライセンス
ライセンスは元データセットに準じます。
謝辞
データセットの製作者様,Qwenの開発者様,計算資源を貸してくださったVolt mindの皆様に感謝します。
Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedVietnamese-cosmos-qa-gg-translatedVietnamese-LLaVA-Instruct-150K-gg-translatedVietnamese-Openorca-Multiplechoice-gg-translatedcroco-translated-datawiki-atomic-edits-translated-nlThis is a Dutch version of the Wiki Atomic Edits dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
Vietnamese-beyond-rlhf-reward-single-round-gg-translated
