CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
015CD-AI /Vietnamese-Locutusque-function-calling-chatml-gg-translatedtextquestion-answering100K<n<1M27 likes139 downloads2y agoHugging Face025CD-AI /Vietnamese-alpaca-gpt4-gg-translatedtextquestion-answering10K<n<100K20 likes111 downloads3y agoHugging Face035CD-AI /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K8 likes108 downloads2y agoHugging Face045CD-AI /Vietnamese-nampdn-ai-tiny-webtext-gg-translatedtextquestion-answering1M<n<10M10 likes98 downloads3y agoHugging Face05JiayiHe /s1_translated S1K Multilingual Translation Dataset This dataset contains translations of the simplescaling/s1K dataset into 20 languages. Languages Included High-Resource Languages Chinese (Simplified) - zh Spanish - es French - fr German - de Japanese - ja Arabic - ar Russian - ru Portuguese - pt Medium-Resource Languages Korean - ko Vietnamese - vi Thai - th Polish - pl Dutch - nl Turkish - tr Low-Resource Languages Swahili - sw Bengali - bn Urdu… See the full description on the dataset page: https://huggingface.co/datasets/JiayiHe/s1_translated.texttranslation10K<n<100K0 likes95 downloads11mo agoHugging Face065CD-AI /Vietnamese-395k-meta-math-MetaMathQA-gg-translatedtextquestion-answering100K<n<1M61 likes85 downloads3y agoHugging Face075CD-AI /Vietnamese-ShareGPT4Video-ShareGPT4Video-gg-translatedtextvisual-question-answering10K<n<100K0 likes80 downloads2y agoHugging Face085CD-AI /Vietnamese-ShareGPT4Vision-gg-translatedtextvisual-question-answering100K<n<1M3 likes78 downloads2y agoHugging Face095CD-AI /Vietnamese-TriviaQA-RC-gg-translatedtextquestion-answering3 likes76 downloads3y agoHugging Face105CD-AI /Vietnamese-liuhaotian-llava_v1_5_mix665k-gg-translatedtextvisual-question-answering100K<n<1M0 likes73 downloads2y agoHugging Face115CD-AI /Vietnamese-ComplexWebQuestions-gg-translatedquestion-answering10K<n<100K4 likes65 downloads3y agoHugging Face12math-across-languages /gsm8k-translated Multilingual GSM8K Translations This dataset contains machine-translated versions of GSM8K in these languages: French (fr) German (de) Hindi (hi) Dataset Structure For each language, we provide the original GSM8K train and test splits: train: 7,473 samples test: 1,319 samples Each sample consists of a question and an answer. The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.textquestion-answering10K<n<100K0 likes62 downloads3mo agoHugging Face13benjleite /FairytaleQA-translated-spanish Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Spanish machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-spanish.textquestion-answering10K<n<100K0 likes58 downloads1y agoHugging Face145CD-AI /Vietnamese-OpenGVLab-ShareGPT-4o-gg-translatedtextvisual-question-answering10K<n<100K0 likes50 downloads2y agoHugging Face15MedInjection /Translated MedInjection-FR — Translated Subset 🌍 Summary The Translated component of MedInjection-FR adapts large-scale English biomedical instruction datasets into French through high-quality automatic translation.It represents the most extensive part of the collection, comprising 416 401 instruction–response pairs, and provides a bridge between English biomedical resources and French medical instruction tuning. This subset was designed to ensure broad domain coverage while… See the full description on the dataset page: https://huggingface.co/datasets/MedInjection/Translated.textquestion-answering100K<n<1M0 likes49 downloads11mo agoHugging Face16srinivasbilla /semeval-2016-absa-reviews-english-translated-stanford-alpaca Dataset Card for Dataset Name Derived from eastwind/semeval-2016-absa-reviews-arabic using Helsinki-NLP/opus-mt-tc-big-ar-en texttext-classification10K<n<100K3 likes45 downloads3y agoHugging Face175CD-AI /Vietnamese-meta-math-MetaMathQA-40K-gg-translatedtextquestion-answering10K<n<100K16 likes45 downloads3y agoHugging Face185CD-AI /Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedtexttext-generation100K<n<1M4 likes44 downloads3y agoHugging Face195CD-AI /Vietnamese-cosmos-qa-gg-translatedtextquestion-answering10K<n<100K6 likes43 downloads3y agoHugging Face205CD-AI /Vietnamese-LLaVA-Instruct-150K-gg-translatedtextvisual-question-answering100K<n<1M27 likes43 downloads3y agoHugging Face215CD-AI /Vietnamese-Openorca-Multiplechoice-gg-translatedtabularquestion-answering10K<n<100K2 likes43 downloads2y agoHugging Face225CD-AI /Vietnamese-beyond-rlhf-reward-single-round-gg-translatedtextquestion-answering10K<n<100K6 likes42 downloads3y agoHugging Face23benjleite /FairytaleQA-translated-ptBR Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Brazilian Portuguese (pt-BR) machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-ptBR.textquestion-answering10K<n<100K3 likes42 downloads1y agoHugging Face24joshbarua /s1K-1.1-Translated s1K-1.1-Translated This dataset contains translated versions of the s1K-1.1 dataset across multiple languages. Languages The dataset contains the following language subsets: Zh, Fr, Ja, Af, Th, Lv, Mr, Te, Sw, En Translation Method This dataset was created using Gemini 2.0 Flash for automatic translation. Dataset Structure Each language subset contains conversational data in the following format: { 'conversations': [ {'from': 'human'… See the full description on the dataset page: https://huggingface.co/datasets/joshbarua/s1K-1.1-Translated.texttext-generation10K<n<100K3 likes41 downloads1y agoHugging Face255CD-AI /Vietnamese-nvidia-OpenMathInstruct-1-50k-gg-translatedtexttext-generation10K<n<100K7 likes39 downloads3y agoHugging Face26benjleite /FairytaleQA-translated-ptPT Dataset Card for FairytaleQA-translated-ptPT Dataset Summary This repository contains the European Portuguese (pt-PT) machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-ptPT.textquestion-answering10K<n<100K0 likes37 downloads1y agoHugging Face27yuyijiong /multi-doc-qa-zh-translated 中文多文档QA数据集 从togethercomputer/Long-Data-Collections中的多文档QA任务,使用谷歌翻译机翻成中文得到。 任务:给定多个参考文档和一个问题,只有一个文档包含有用信息,模型需要根据参考文档回答问题,并指出哪个文档包含有用信息。 对于每个question,会提供几十或上百个文档片段,只有一个文档包含有用信息,gold_document_id表示含有有用信息的文档序号,注意文档是从1开始编号。 texttext-generation10K<n<100K11 likes35 downloads1y agoHugging Face28benjleite /FairytaleQA-translated-french Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the French machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-french.textquestion-answering10K<n<100K1 likes35 downloads1y agoHugging Face29Detsutut /MedQA-USMLE-back-translatedMedQA dataset perturbed using back-translation technique with BAT tabularquestion-answering10K<n<100K0 likes34 downloads2y agoHugging Face30ChaosAIVision /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K0 likes34 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.