CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01databricks /officeqagated OfficeQA Dataset Summary OfficeQA is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over historical U.S. Treasury Bulletin documents (1939–2025), which contain dense financial tables, charts, and narrative text. OfficeQA is designed to test retrieval, tool use, and multi-step reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa.documentquestion-answeringn<1K26 likes9.1k downloads2mo agoHugging Face02bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K195 likes7.9k downloads2y agoHugging Face03databricks /officeqa-pro-v2gated OfficeQA Pro v2 Dataset Summary OfficeQA Pro v2 is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over two centuries of U.S. Federal Accounts of Receipts and Expenditures reporting (1793–2024) — Combined Statements of Receipts, Outlays, and Balances of the United States Government, together with earlier… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa-pro-v2.documentquestion-answeringn<1K17 likes2.6k downloads2mo agoHugging Face04bitext /Bitext-retail-ecommerce-llm-chatbot-training-dataset Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.textquestion-answering10K<n<100K19 likes1.5k downloads2y agoHugging Face05bitext /Bitext-events-ticketing-llm-chatbot-training-dataset Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes968 downloads2y agoHugging Face06snorkelai /finqa-data SnorkelFinance Expert-verified financial QA dataset for evaluating AI agents on tool-calling and reasoning over SEC 10-K filings. Overview SnorkelFinance is a benchmark of 290 questions across 20 companies spanning 5 industry verticals. Questions are created from 10-K filing documents and verified by Snorkel's network of financial experts on a 5-point scale for realism and accuracy. Agents don't have direct access to the documents. Instead, they must plan and use provided… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/finqa-data.textquestion-answeringn<1K1 likes821 downloads7mo agoHugging Face07habedi /stack-exchange-dataset Overview This dataset consists of three TSV files, namely: cs.tsv, ds.tsv, and p.tsv. Each file includes the data for the questions asked on a Stack Exchange (SE) question-answering community, from the creation of the community until May 2021. cs.tsv --> Computer Science SE ds.csv --> Data Science SE p.csv --> Political Science SE File Structure Each file has the following columns: id: the question id title: the title of the question body: the body or text of the… See the full description on the dataset page: https://huggingface.co/datasets/habedi/stack-exchange-dataset.tabulartext-classification10K<n<100K11 likes380 downloads7mo agoHugging Face08facebook /EgoAVU_data [CVPR2026 HIGHLIGHT] EgoAVU, [ICASSP2026 Oral] Exploring Audio Hallucination in Egocentric Video Understanding Official Implementation of EgoAVU: Egocentric Audio-Visual Understanding and Exploring Audio Hallucination in Egocentric Video Understanding See our github for the code and setup instructions. Check out our homepage, paper (CVPR) and paper (ICASSP) for more information. We introduce EgoAVU, a scalable and automated data engine to enable egocentric audio–visual… See the full description on the dataset page: https://huggingface.co/datasets/facebook/EgoAVU_data.tabularquestion-answering1M<n<10M14 likes272 downloads5mo agoHugging Face09NoirZangetsu /Flutter-Code-with-Questions-Dataset-Turkish Flutter Code with Questions Dataset (Turkish) 📦 Dataset Name: flutter_code_with_questions Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir. 📁 Dataset Format Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.textquestion-answering1K<n<10K0 likes264 downloads2mo agoHugging Face10bitext /Bitext-telco-llm-chatbot-training-dataset Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes254 downloads2y agoHugging Face11NoirZangetsu /Flutter-Code-with-Questions-Dataset-English 🧠 Flutter Code with Questions Dataset (English) This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development. 📂 Dataset Structure The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes: A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.textquestion-answering1K<n<10K3 likes222 downloads2mo agoHugging Face12bitext /Bitext-insurance-llm-chatbot-training-dataset Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.textquestion-answering10K<n<100K8 likes216 downloads2y agoHugging Face13McGill-NLP /statcan-dialogue-dataset-retrieval Statcan Dialogue Dataset (Processed for Retrieval Tasks) This is a variant of the Statcan Dialogue Dataset, which we processed specifically for multilingual retrieval (english, french). It contains everything in CSVs, rather than having metadata hosted separately. Quickstart from datasets import load_dataset repo = 'McGill-NLP/statcan-dialogue-dataset-retrieval' # load english queries, training split queries_en = load_dataset(repo, 'queries_english', split='train') #… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/statcan-dialogue-dataset-retrieval.textquestion-answering10K<n<100K1 likes197 downloads2y agoHugging Face14maum-ai /KOFFVQA_Data About this data KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language KOFFVQA is a general-purpose VLM benchmark in the Korean language. For more information, refer to our leaderboard page and the official evaluation code. This contains the data for the benchmark consisting of images, their corresponding questions, and response grading criteria. The benchmark focuses on free-form visual question answering, evaluating the… See the full description on the dataset page: https://huggingface.co/datasets/maum-ai/KOFFVQA_Data.textvisual-question-answeringn<1K2 likes188 downloads1y agoHugging Face15bitext /Bitext-travel-llm-chatbot-training-dataset Bitext - Travel Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Travel] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-travel-llm-chatbot-training-dataset.textquestion-answering10K<n<100K4 likes184 downloads2y agoHugging Face16k-master /k-beauty-ai-citation-dataset K-Beauty AI Citation Dataset Open dataset mapping Korean K-beauty entities (ingredients, skin concerns, use cases, brands) and answer-style guides to citation-shaped external references. Designed to be referenced by AI search engines, content builders, and SEO research. Canonical source: https://kbeautyanswers.com/dataset/ License: CC BY 4.0 Maintainer: K-Beauty Answers (site) Initial release: 2026-05-23 What's in it 128 entities (37 ingredients + 18 skin… See the full description on the dataset page: https://huggingface.co/datasets/k-master/k-beauty-ai-citation-dataset.texttext-classificationn<1K0 likes176 downloads3mo agoHugging Face17binhtran23 /vecombot-dataset VECOM — Bộ dữ liệu thị trường Thương mại điện tử Việt Nam (VEComBot) Bộ dữ liệu phụ lục cho đồ án tốt nghiệp VEComBot — hệ thống Đa tác tử (Multi-Agent System) phân tích và tổng hợp thị trường Thương mại điện tử Việt Nam (VECOM). Đây là kho tài liệu nguồn và corpus đã qua xử lý (figure-aware) được nạp vào PostgreSQL/pgvector để phục vụ cả nhánh MAS lẫn nhánh baseline naive RAG. Mục đích: dùng cho nghiên cứu học thuật và tái lập kết quả đồ án. Các báo cáo gốc là ấn phẩm công… See the full description on the dataset page: https://huggingface.co/datasets/binhtran23/vecombot-dataset.imagequestion-answeringn<1K0 likes162 downloads3mo agoHugging Face18kguo2 /MolPuzzle_data MolPuzzle: A Multimodal Benchmark for Molecular Structure Elucidation Dataset Description: The MolPuzzle dataset is a newly developed resource designed to challenge Large Language Models with multi-modal, multi-step reasoning tasks (molecular structure elucidation). This dataset consists of 217 diverse and intricate structure elucidation challenges that require LLMs to demonstrate advanced reasoning capabilities, integrating multimodal data and deep chemical understanding… See the full description on the dataset page: https://huggingface.co/datasets/kguo2/MolPuzzle_data.imagequestion-answering10K<n<100K0 likes142 downloads2y agoHugging Face19me-aas /pentesting-dataset Dataset Card for Penetration Testing Dataset This dataset card aims to provide essential information about the Penetration Testing Dataset, which includes various resources and scripts useful for penetration testing and cybersecurity research. Dataset Details Dataset Description The Penetration Testing Dataset is a collection of scripts, tools, and vulnerability data designed for cybersecurity professionals to facilitate penetration testing tasks.… See the full description on the dataset page: https://huggingface.co/datasets/me-aas/pentesting-dataset.tabularquestion-answering1K<n<10K0 likes139 downloads4mo agoHugging Face20aumghag /Data-Analytics-Digital-Marketing-Project-Management-QA_DBtextquestion-answeringn<1K4 likes135 downloads2y agoHugging Face21boapro /pentesting-dataset Dataset Card for Penetration Testing Dataset This dataset card aims to provide essential information about the Penetration Testing Dataset, which includes various resources and scripts useful for penetration testing and cybersecurity research. Dataset Details Dataset Description The Penetration Testing Dataset is a collection of scripts, tools, and vulnerability data designed for cybersecurity professionals to facilitate penetration testing tasks. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/boapro/pentesting-dataset.tabularquestion-answering1K<n<10K3 likes132 downloads1y agoHugging Face22KadamParth /Ncert_datasettabularquestion-answering100K<n<1M4 likes129 downloads1y agoHugging Face23Canstralian /pentesting_datasetgated Dataset Card for Penetration Testing Dataset This dataset card aims to provide essential information about the Penetration Testing Dataset, which includes various resources and scripts useful for penetration testing and cybersecurity research. Dataset Details Dataset Description The Penetration Testing Dataset is a collection of scripts, tools, and vulnerability data designed for cybersecurity professionals to facilitate penetration testing tasks. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Canstralian/pentesting_dataset.tabularquestion-answering1K<n<10K12 likes112 downloads2y agoHugging Face24sarahwei /cyber_MITRE_CTI_dataset_v15This dataset is a specialized resource designed for training and evaluating question-answering models in the context of Cyber Threat Intelligence (CTI), specifically targeting the identification of tactics and techniques based on natural language descriptions of cyber-attacks. The dataset is derived from the MITRE ATT&CK framework (version 15) and contains annotated pairs of sentences and their corresponding tactics and techniques. The primary goal is to assist automated systems in… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/cyber_MITRE_CTI_dataset_v15.textquestion-answering10K<n<100K6 likes99 downloads2y agoHugging Face25bitext /Bitext-mortgage-loans-llm-chatbot-training-dataset Bitext - Mortgage and Loans Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Mortgage and Loans] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-mortgage-loans-llm-chatbot-training-dataset.textquestion-answering10K<n<100K5 likes95 downloads2y agoHugging Face26Pradeep016 /career-guidance-qa-dataset Dataset Card for Career Guidance Dataset Dataset Overview This dataset provides career guidance information for a variety of career roles. It includes questions and answers related to career roles such as "Data Scientist," "Software Engineer," "Product Manager," and many more. The dataset covers aspects like job responsibilities, required skills, career progression, salary expectations, and work environment. It is intended for use in building chatbot applications for… See the full description on the dataset page: https://huggingface.co/datasets/Pradeep016/career-guidance-qa-dataset.textquestion-answering1K<n<10K7 likes87 downloads2y agoHugging Face27tarekmasryo /rag-qa-logs-corpus-data 🧠📚 RAG QA Logs & Corpus (Synthetic) 🧪 Multi-table synthetic RAG telemetry for quality, hallucinations, latency, and cost A production-style, privacy-safe synthetic dataset that mimics telemetry exported from a real RAG system — from corpus → chunks → retrieval events → eval runs. ✅ Fully synthetic (no real users / orgs / PII). ⚡ Quick facts Total rows: 103,255 across 6 linked tables Labels (in eval_runs): is_correct, hallucination_flag, faithfulness_label… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/rag-qa-logs-corpus-data.tabularquestion-answering100K<n<1M2 likes86 downloads8mo agoHugging Face28Taylor658 /deep-space-optical-chip-thermal-dataset 🚀 Deep Space Optical Chip Thermal Dataset 🪐 🌡️ 40,000 scenario-based prompt and response pairs on thermal mitigation for photonic chips in scientific instruments aboard deep-space probes, covering refractive index drift, waveguide misalignment, and thermal stress across materials, instruments, and environments. ⚠️ Disclaimer: All entries are synthetically generated. Material coefficients are drawn from published typical values, but no row is based on mission logs or flight… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/deep-space-optical-chip-thermal-dataset.tabulartext-generation10K<n<100K2 likes82 downloads11d agoHugging Face29bitext /Bitext-wealth-management-llm-chatbot-training-dataset Bitext - Wealth Management Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Wealth Management] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-wealth-management-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes81 downloads2y agoHugging Face30KBayoud /MoroccanHistory-QA-Datasettextquestion-answering1K<n<10K3 likes79 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.