CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ItalianNarratives /megamatt-translated-ITtext100K<n<1M0 likes724 downloads18d agoHugging Face02ItalianNarratives /cranemath-translated-ITtext100K<n<1M0 likes542 downloads18d agoHugging Face03NetherlandsForensicInstitute /s2orc-citation-pairs-translated-nlThis is a Dutch version of the S2ORC: The Semantic Scholar Open Research Corpus. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity10M<n<100M0 likes284 downloads2y agoHugging Face045CD-AI /Vietnamese-lmms-lab-LLaVA-Video-178K-gg-translated Dataset Card for 5CD-AI/Vietnamese-lmms-lab-LLaVA-Video-178K-gg-translated This translated dataset includes: LLaVA-Video-178K: 178,509 caption entries, 960,791 open-ended QA (question and answer) items, and 196,198 multiple-choice QA items. The video source of the original dataset is in this repo: lmms-lab/LLaVA-Video-178K textvisual-question-answering1M<n<10M1 likes273 downloads2y agoHugging Face05masakhane /african-translated-alpaca Citation Please cite the stanford_alpaca project @misc{alpaca, author = {Rohan Taori and Ishaan Gulrajani and Tianyi Zhang and Yann Dubois and Xuechen Li and Carlos Guestrin and Percy Liang and Tatsunori B. Hashimoto }, title = {Stanford Alpaca: An Instruction-following LLaMA model}, year = {2023}, publisher = {GitHub}, journal = {GitHub repository}, howpublished = {\url{https://github.com/tatsu-lab/stanford_alpaca}}, } text1M<n<10M1 likes119 downloads2y agoHugging Face065CD-AI /Vietnamese-alpaca-gpt4-gg-translatedtextquestion-answering10K<n<100K20 likes111 downloads3y agoHugging Face075CD-AI /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K8 likes108 downloads2y agoHugging Face085CD-AI /Vietnamese-nampdn-ai-tiny-webtext-gg-translatedtextquestion-answering1M<n<10M10 likes98 downloads3y agoHugging Face095CD-AI /Vietnamese-395k-meta-math-MetaMathQA-gg-translatedtextquestion-answering100K<n<1M61 likes85 downloads3y agoHugging Face10mesolitica /translated-cnn-dailymailtext100K<n<1M1 likes82 downloads4y agoHugging Face11mesolitica /translated-xwikistext10K<n<100K0 likes82 downloads4y agoHugging Face125CD-AI /Vietnamese-ShareGPT4Vision-gg-translatedtextvisual-question-answering100K<n<1M3 likes78 downloads2y agoHugging Face13ItalianNarratives /dolmino-math-translated-ITtext100K<n<1M0 likes77 downloads18d agoHugging Face145CD-AI /Vietnamese-liuhaotian-llava_v1_5_mix665k-gg-translatedtextvisual-question-answering100K<n<1M0 likes73 downloads2y agoHugging Face155CD-AI /Vietnamese-Intel-orca_dpo_pairs-gg-translatedtext10K<n<100K35 likes66 downloads2y agoHugging Face16benjleite /FairytaleQA-translated-spanish Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Spanish machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-spanish.textquestion-answering10K<n<100K0 likes58 downloads1y agoHugging Face17pragnakalp /squad_v2_french_translatedUsing Google Translation, we have translated SQuAD 2.0 dataset into multiple languages. Here is the translated dataset of SQuAD 2.0 in French language. Shared by Pragnakalp Techlabs textn<1K1 likes57 downloads4y agoHugging Face18ItalianNarratives /tinymath-pot-translated-ITtext100K<n<1M0 likes54 downloads18d agoHugging Face19mesolitica /translated-glaive-function-calltext100K<n<1M0 likes50 downloads3y agoHugging Face20MedInjection /Translated MedInjection-FR — Translated Subset 🌍 Summary The Translated component of MedInjection-FR adapts large-scale English biomedical instruction datasets into French through high-quality automatic translation.It represents the most extensive part of the collection, comprising 416 401 instruction–response pairs, and provides a bridge between English biomedical resources and French medical instruction tuning. This subset was designed to ensure broad domain coverage while… See the full description on the dataset page: https://huggingface.co/datasets/MedInjection/Translated.textquestion-answering100K<n<1M0 likes49 downloads11mo agoHugging Face21taresco /open_math_instruct_v2_translated_african_languagesThis is a set of 41k nvidia/OpenMathInstruct-2 questions translated into 9 African languages using Azure/GPT-4o. We shuffle the dataset and then randomly sample a question without replacement, and then equally sample a language and then we translate the question and answer to that language. text10K<n<100K0 likes46 downloads9mo agoHugging Face225CD-AI /Vietnamese-meta-math-MetaMathQA-40K-gg-translatedtextquestion-answering10K<n<100K16 likes45 downloads3y agoHugging Face23blastai /Open_o1_sft_Pro_translated_jp 概要 このデータセットはOpen_o1_sft_ProデータセットをQwen社のQwen2.5-14B-Instructを用いて日本語に翻訳したものになります。 テンプレート テンプレートは以下です。 {"conversations": [{"role": "user", "content": "入力"}, {"role": "assistant", "thought": "思考", "content": "出力"}, ...], "id": id(整数), "dataset": "元データセットの名前"} ライセンス ライセンスは元データセットに準じます。 謝辞 データセットの製作者様,Qwenの開発者様,計算資源を貸してくださったVolt mindの皆様に感謝します。 texttext-generation10K<n<100K9 likes45 downloads2y agoHugging Face245CD-AI /Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedtexttext-generation100K<n<1M4 likes44 downloads3y agoHugging Face255CD-AI /Vietnamese-cosmos-qa-gg-translatedtextquestion-answering10K<n<100K6 likes43 downloads3y agoHugging Face265CD-AI /Vietnamese-LLaVA-Instruct-150K-gg-translatedtextvisual-question-answering100K<n<1M27 likes43 downloads3y agoHugging Face275CD-AI /Vietnamese-Openorca-Multiplechoice-gg-translatedtabularquestion-answering10K<n<100K2 likes43 downloads2y agoHugging Face28jjzha /croco-translated-datatext100K<n<1M0 likes43 downloads4mo agoHugging Face29NetherlandsForensicInstitute /wiki-atomic-edits-translated-nlThis is a Dutch version of the Wiki Atomic Edits dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity10M<n<100M1 likes42 downloads2y agoHugging Face305CD-AI /Vietnamese-beyond-rlhf-reward-single-round-gg-translatedtextquestion-answering10K<n<100K6 likes42 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.