CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01uilab /BLEnD BLEnD This is the official repository of BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages (Submitted to NeurIPS 2024 Datasets and Benchmarks Track). 24/12/05: Updated translation errors25/05/02: Updated multiple choice questions file (v1.1)26/09/15: Added new data collected for SemEval-2026 Task 7, covering 17 additional language-culture pairs (semeval-annotations, semeval-questions, and semeval split of multiple-choice-questions)… See the full description on the dataset page: https://huggingface.co/datasets/uilab/BLEnD.textquestion-answering100K<n<1M16 likes1.6k downloads9d agoHugging Face02taidng /UIT-ViQuAD2.0 Vietnamese Question Answering Dataset Dataset Card for UIT-ViQuAD2.0 Dataset Summary The HF version for Vietnamese QA dataset created by Nguyen et al. (2020) and released in the shared task. The original UIT-ViQuAD contains over 23,000 QA pairs based on 174 Vietnamese Wikipedia articles. UIT-ViQuAD2.0 adds over 12K unanswerable questions for the same passage. The dataset has been processed to remove a few duplicated questions and answers. Version 2.0 contains… See the full description on the dataset page: https://huggingface.co/datasets/taidng/UIT-ViQuAD2.0.textquestion-answering10K<n<100K21 likes562 downloads2y agoHugging Face03macpaw-research /UiPad UiPad - UI Parsing and Accessibility Dataset 📌 Dataset status: stable release. UiPad was built for the IASA Champ 2024 Challenge and is a complete, fixed research artifact. No further updates are planned. Curated by: MacPaw Way Ltd. Language(s): Mostly EN, UA License: MIT Overview UiPad is a dataset created for the IASA Champ 2024 Challenge, focusing on the accessibility and interface understanding of MacOS applications. With growing interest in AI-driven user interface… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/UiPad.imagequestion-answering1K<n<10K16 likes211 downloads1mo agoHugging Face04DataScience-UIBK /PopMCQ 🎯 PopMCQ Does your model pick the famous answer or the correct one? 📌 Overview PopMCQ renders the same question six ways. The question and the correct answer never change — only how popular the three distractors are. That makes option popularity an independent variable, so an accuracy swing across S1–S6 is attributable to popularity rather than to question difficulty. The swings are large. Under the hardest setting (S2), models pick a popular-but-wrong… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/PopMCQ.tabularquestion-answering10K<n<100K0 likes166 downloads14d agoHugging Face05UIS-Digger /UIS-QA UIS-QA: A Benchmark for Unindexed Information Seeking Figure 1. UIS problem. Standard agents (bottom) rely on indexed information and often fail or hallucinate; UIS-capable agents (top) use additional tools to excavate unindexed information and solve UIS tasks. If .figs do not load, see the paper. 🔔 News [2026.03.10] 🎉 We release the UIS-QA dataset and the paper (ICLR 2026, arXiv) today! 📋 Dataset Description Homepage Paper… See the full description on the dataset page: https://huggingface.co/datasets/UIS-Digger/UIS-QA.imagequestion-answeringn<1K1 likes117 downloads7mo agoHugging Face06LAIA-UIB /SOAR SOAR: System Operations and Reasoning Benchmark System Operations and Reasoning (SOAR) is a human-curated benchmark designed to evaluate the ability of large language models (LLMs) to interpret and answer questions over IT operational data, with a primary focus on Unix/Linux-style system administration contexts.It focuses on how well models can interpret, extract, and reason over operational data — including logs, configurations, and system outputs — and on how question… See the full description on the dataset page: https://huggingface.co/datasets/LAIA-UIB/SOAR.tabularquestion-answeringn<1K0 likes44 downloads20d agoHugging Face07uitnlp /vicoqagated Vietnamese Conversational Machine Comprehension dataset (UIT-ViCoQA). This datset is used for Conversational Machine Comprehension task in Vietnamese. The UIT-ViCoQA dataset consists of 10,000 questions with answers over 2,000 conversations about health news articles. Source code The source code available at: https://github.com/sonlam1102/vicoqa-cmc. Publication Please cite this paper if you use our dataset @inproceedings{luu2021conversational… See the full description on the dataset page: https://huggingface.co/datasets/uitnlp/vicoqa.textquestion-answering1K<n<10K0 likes29 downloads29d agoHugging Face08tuanquocbd /UIT-ViQuAD2.0 Vietnamese Question Answering Dataset Dataset Card for UIT-ViQuAD2.0 Dataset Summary The HF version for Vietnamese QA dataset created by Nguyen et al. (2020) and released in the shared task. The original UIT-ViQuAD contains over 23,000 QA pairs based on 174 Vietnamese Wikipedia articles. UIT-ViQuAD2.0 adds over 12K unanswerable questions for the same passage. The dataset has been processed to remove a few duplicated questions and answers. Version… See the full description on the dataset page: https://huggingface.co/datasets/tuanquocbd/UIT-ViQuAD2.0.textquestion-answering10K<n<100K0 likes24 downloads3mo agoHugging Face09uitnlp /vimmrc2.0gated ViMMRC 2.0 The Vietnamese Multiple-choice reading comprehension dataset version 2 (ViMMRC 2.0) The dataset is freely available for research purposes only. Users need to sign the data agreement before receiving the dataset. More information, please visit the NLP@UIT research group: https://nlp.uit.edu.vn/ The original Github for the dataset (including source code): https://github.com/sonlam1102/vimmrc2 Usage from datasets import load_dataset train =… See the full description on the dataset page: https://huggingface.co/datasets/uitnlp/vimmrc2.0.textquestion-answeringn<1K1 likes21 downloads8mo agoHugging Face10nhphuc210 /UIT-ViQuAD2.0 Vietnamese Question Answering Dataset Dataset Card for UIT-ViQuAD2.0 Dataset Summary The HF version for Vietnamese QA dataset created by Nguyen et al. (2020) and released in the shared task. The original UIT-ViQuAD contains over 23,000 QA pairs based on 174 Vietnamese Wikipedia articles. UIT-ViQuAD2.0 adds over 12K unanswerable questions for the same passage. The dataset has been processed to remove a few duplicated questions and answers. Version… See the full description on the dataset page: https://huggingface.co/datasets/nhphuc210/UIT-ViQuAD2.0.textquestion-answering10K<n<100K0 likes13 downloads3mo agoHugging Face11dim014 /ui-form-user-manual-generation-dataset-rus UI Form User Manual Generation Dataset (Russian) Dataset Description This dataset was developed on the basis of 'yahma/alpaca-cleaned' dataset. It contains examples of generating user guides for interface forms in Russian. Each example includes a description of the UI form elements and corresponding step-by-step instructions for completing it. Data Structure The dataset is in JSON format, and contains three fields: instruction — system instruction input —… See the full description on the dataset page: https://huggingface.co/datasets/dim014/ui-form-user-manual-generation-dataset-rus.textquestion-answering1K<n<10K0 likes10 downloads10mo agoHugging Face12uitnlp /vlogqagated Dataset Card for Dataset Name VlogQA: Question-answering based on the transcript from VLOG Videos in Vietnamese Language about travel and food. Dataset Details VlogQA consists of 10,076 question-answer pairs based on 1,230 transcript documents sourced from YouTube You need to sign the agreement and send to the author via the email in "Dataset Contact" to access to the dataset. The agreement from can be found at:… See the full description on the dataset page: https://huggingface.co/datasets/uitnlp/vlogqa.textquestion-answering1K<n<10K2 likes9 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.