CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KamiKrafton /agentvidbench AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents Agentic Video Understanding Benchmark — 100 multiple-choice video QA questions, 26 options each (A-Z; ~3.8% random baseline) by Seoyeon An*, Hyeonseo Jang*, Minsu Kim*, Chanho Lee, Younghan Park, Kangwook Lee (KRAFTON AI) Layout . ├── README.md ├── questions.jsonl # 100 rows — one per question ├── videos.jsonl # 71 rows — one per unique video ├── videos/… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench.textvideo-text-to-textn<1K0 likes258 downloads2mo agoHugging Face02kamruzzaman-asif /bangla-instruction-dataset 🧠 Bangla Instruction Dataset This dataset repository consolidates high-quality instruction-tuning data from multiple popular sources, structured for easy use in training and evaluating instruction-following models. 📚 Dataset Splits The dataset is organized into the following splits: Split Name Source Dataset Description OdiaGenAI OdiaGenAI/all_combined_bengali_252k A large-scale collection of diverse Bangla instructions and responses. chrononeel… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/bangla-instruction-dataset.texttext-generation1M<n<10M1 likes219 downloads1y agoHugging Face03KamiKrafton /agentvidbench-sample AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above. Sample selection The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench-sample.textvideo-text-to-textn<1K0 likes128 downloads2mo agoHugging Face04touati-kamel /DziriEval DziriEval : Benchmark d'Évaluation des LLMs en Dialecte Algérien (Darja) DziriEval est le benchmark académique natif de questions-réponses à choix multiples (QCM) conçu spécifiquement pour évaluer les capacités de compréhension, de raisonnement et de connaissance culturelle des grands modèles de langage (LLMs) sur le dialecte algérien (Darja). Ce jeu de données a été construit et vérifié manuellement afin de refléter la richesse linguistique, culturelle et quotidienne de… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/DziriEval.textmultiple-choice1K<n<10K1 likes61 downloads3mo agoHugging Face05Yazeed-Kamel /hashem-dataset Hashem AI – Jordanian Legal FAQ Dataset Overview This dataset contains a curated collection of frequently asked questions (FAQ) and authoritative legal answers related to Jordanian laws and regulations.It is designed to support Arabic-language legal question answering, retrieval-augmented generation (RAG), and legal knowledge systems. The dataset focuses on providing clear, neutral, and legally accurate explanations based strictly on official Jordanian legislation and… See the full description on the dataset page: https://huggingface.co/datasets/Yazeed-Kamel/hashem-dataset.question-answering2 likes38 downloads8mo agoHugging Face06Kamitor /TestDataQuora Question Answer Dataset (Quora-QuAD) contains 56,402 question-answer pairs scraped from Quora. Usage: For instructions on fine-tuning a model (Flan-T5) with this dataset, please check out the article: https://www.toughdata.net/blog/post/finetune-flan-t5-question-answer-quora-dataset textquestion-answering10K<n<100K0 likes15 downloads2y agoHugging Face07kamarko /ccnews-french-subsetCredits and Attribution: This dataset is derived from the Common Crawl dataset (https://huggingface.co/datasets/stanford-oval/ccnews). The data has been transformed and filtered to achieve the current format. For license information, please refer to https://commoncrawl.org/terms-of-use texttext-classification1K<n<10K0 likes11 downloads2y agoHugging Face08Kamyar-zeinalipour /MultiHint-9K MultiHint-9K MultiHint-9K is a multilingual dataset of 10,032 machine-verified, human-calibrated QA-Hint tuples across English, Italian, and Farsi, produced by a human-AI agency pipeline with adversarial quality assurance. 📄 Paper: MultiHint-9K: Orchestrating Human and AI Agency for Scalable Pedagogical Hint Generation — HAI-Agency @ AIED 2026💻 Code: github.com/KamyarZeinalipour/MultiHint-9K Dataset Summary Each example contains a Wikipedia passage, a… See the full description on the dataset page: https://huggingface.co/datasets/Kamyar-zeinalipour/MultiHint-9K.question-answering10K<n<100K0 likes10 downloads4mo agoHugging Face09kammavidya /educationgatedtextquestion-answeringn<1K0 likes9 downloads3y agoHugging Face10kammavidya /NGITgatedtextquestion-answeringn<1K0 likes4 downloads3y agoHugging Face11kammavidya /MQAgatedtextquestion-answeringn<1K0 likes3 downloads3y agoHugging Face12kammavidya /AIgated Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/kammavidya/AI.question-answering0 likes2 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.