datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
officeqa
OfficeQA
Dataset Summary
OfficeQA is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents.
The benchmark consists of question–answer pairs that require reasoning over historical U.S. Treasury Bulletin documents (1939–2025), which contain dense financial tables, charts, and narrative text. OfficeQA is designed to test retrieval, tool use, and multi-step reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.officeqa-pro-v2
OfficeQA Pro v2
Dataset Summary
OfficeQA Pro v2 is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents.
The benchmark consists of question–answer pairs that require reasoning over two centuries of U.S. Federal Accounts of Receipts and Expenditures reporting (1793–2024) — Combined Statements of Receipts, Outlays, and Balances of the United States Government, together with earlier… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa-pro-v2.Bitext-retail-ecommerce-llm-chatbot-training-dataset
Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.Bitext-events-ticketing-llm-chatbot-training-dataset
Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.finqa-data
SnorkelFinance
Expert-verified financial QA dataset for evaluating AI agents on tool-calling and reasoning over SEC 10-K filings.
Overview
SnorkelFinance is a benchmark of 290 questions across 20 companies spanning 5 industry verticals. Questions are created from 10-K filing documents and verified by Snorkel's network of financial experts on a 5-point scale for realism and accuracy.
Agents don't have direct access to the documents. Instead, they must plan and use provided… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/finqa-data.stack-exchange-dataset
Overview
This dataset consists of three TSV files, namely: cs.tsv, ds.tsv, and p.tsv.
Each file includes the data for the questions asked on a Stack Exchange (SE) question-answering community, from the creation of the community until May 2021.
cs.tsv --> Computer Science SE
ds.csv --> Data Science SE
p.csv --> Political Science SE
File Structure
Each file has the following columns:
id: the question id
title: the title of the question
body: the body or text of the… See the full description on the dataset page: https://huggingface.co/datasets/habedi/stack-exchange-dataset.EgoAVU_data
[CVPR2026 HIGHLIGHT] EgoAVU, [ICASSP2026 Oral] Exploring Audio Hallucination in Egocentric Video Understanding
Official Implementation of EgoAVU: Egocentric Audio-Visual Understanding and Exploring Audio Hallucination in Egocentric Video Understanding
See our github for the code and setup instructions.
Check out our homepage, paper (CVPR) and paper (ICASSP) for more information.
We introduce EgoAVU, a scalable and automated data engine to enable egocentric audio–visual… See the full description on the dataset page: https://huggingface.co/datasets/facebook/EgoAVU_data.Flutter-Code-with-Questions-Dataset-Turkish
Flutter Code with Questions Dataset (Turkish)
📦 Dataset Name: flutter_code_with_questions
Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir.
📁 Dataset Format
Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.Bitext-telco-llm-chatbot-training-dataset
Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.Flutter-Code-with-Questions-Dataset-English
🧠 Flutter Code with Questions Dataset (English)
This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development.
📂 Dataset Structure
The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes:
A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.Bitext-insurance-llm-chatbot-training-dataset
Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.statcan-dialogue-dataset-retrieval
Statcan Dialogue Dataset (Processed for Retrieval Tasks)
This is a variant of the Statcan Dialogue Dataset, which we processed specifically for multilingual retrieval (english, french). It contains everything in CSVs, rather than having metadata hosted separately.
Quickstart
from datasets import load_dataset
repo = 'McGill-NLP/statcan-dialogue-dataset-retrieval'
# load english queries, training split
queries_en = load_dataset(repo, 'queries_english', split='train') #… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/statcan-dialogue-dataset-retrieval.KOFFVQA_Data
About this data
KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
KOFFVQA is a general-purpose VLM benchmark in the Korean language. For more information, refer to our leaderboard page and the official evaluation code.
This contains the data for the benchmark consisting of images, their corresponding questions, and response grading criteria. The benchmark focuses on free-form visual question answering, evaluating the… See the full description on the dataset page: https://huggingface.co/datasets/maum-ai/KOFFVQA_Data.Bitext-travel-llm-chatbot-training-dataset
Bitext - Travel Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Travel] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-travel-llm-chatbot-training-dataset.k-beauty-ai-citation-dataset
K-Beauty AI Citation Dataset
Open dataset mapping Korean K-beauty entities (ingredients, skin concerns, use cases, brands) and answer-style guides to citation-shaped external references. Designed to be referenced by AI search engines, content builders, and SEO research.
Canonical source: https://kbeautyanswers.com/dataset/
License: CC BY 4.0
Maintainer: K-Beauty Answers (site)
Initial release: 2026-05-23
What's in it
128 entities (37 ingredients + 18 skin… See the full description on the dataset page: https://huggingface.co/datasets/k-master/k-beauty-ai-citation-dataset.vecombot-dataset
VECOM — Bộ dữ liệu thị trường Thương mại điện tử Việt Nam (VEComBot)
Bộ dữ liệu phụ lục cho đồ án tốt nghiệp VEComBot — hệ thống Đa tác tử (Multi-Agent
System) phân tích và tổng hợp thị trường Thương mại điện tử Việt Nam (VECOM). Đây là
kho tài liệu nguồn và corpus đã qua xử lý (figure-aware) được nạp vào PostgreSQL/pgvector
để phục vụ cả nhánh MAS lẫn nhánh baseline naive RAG.
Mục đích: dùng cho nghiên cứu học thuật và tái lập kết quả đồ án. Các báo cáo gốc là
ấn phẩm công… See the full description on the dataset page: https://huggingface.co/datasets/binhtran23/vecombot-dataset.MolPuzzle_data
MolPuzzle: A Multimodal Benchmark for Molecular Structure Elucidation
Dataset Description:
The MolPuzzle dataset is a newly developed resource designed to challenge Large Language Models with multi-modal, multi-step reasoning tasks (molecular structure elucidation). This dataset consists of 217 diverse and intricate structure elucidation challenges that require LLMs to demonstrate advanced reasoning capabilities, integrating multimodal data and deep chemical understanding… See the full description on the dataset page: https://huggingface.co/datasets/kguo2/MolPuzzle_data.pentesting-dataset
Dataset Card for Penetration Testing Dataset
This dataset card aims to provide essential information about the Penetration Testing Dataset, which includes various resources and scripts useful for penetration testing and cybersecurity research.
Dataset Details
Dataset Description
The Penetration Testing Dataset is a collection of scripts, tools, and vulnerability data designed for cybersecurity professionals to facilitate penetration testing tasks.… See the full description on the dataset page: https://huggingface.co/datasets/me-aas/pentesting-dataset.Data-Analytics-Digital-Marketing-Project-Management-QA_DBpentesting-dataset
Dataset Card for Penetration Testing Dataset
This dataset card aims to provide essential information about the Penetration Testing Dataset, which includes various resources and scripts useful for penetration testing and cybersecurity research.
Dataset Details
Dataset Description
The Penetration Testing Dataset is a collection of scripts, tools, and vulnerability data designed for cybersecurity professionals to facilitate penetration testing tasks. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/boapro/pentesting-dataset.Ncert_datasetpentesting_dataset
Dataset Card for Penetration Testing Dataset
This dataset card aims to provide essential information about the Penetration Testing Dataset, which includes various resources and scripts useful for penetration testing and cybersecurity research.
Dataset Details
Dataset Description
The Penetration Testing Dataset is a collection of scripts, tools, and vulnerability data designed for cybersecurity professionals to facilitate penetration testing tasks. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Canstralian/pentesting_dataset.cyber_MITRE_CTI_dataset_v15This dataset is a specialized resource designed for training and evaluating question-answering models in the context of Cyber Threat Intelligence (CTI), specifically targeting the identification of tactics and techniques based on natural language descriptions of cyber-attacks. The dataset is derived from the MITRE ATT&CK framework (version 15) and contains annotated pairs of sentences and their corresponding tactics and techniques. The primary goal is to assist automated systems in… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/cyber_MITRE_CTI_dataset_v15.Bitext-mortgage-loans-llm-chatbot-training-dataset
Bitext - Mortgage and Loans Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Mortgage and Loans] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-mortgage-loans-llm-chatbot-training-dataset.career-guidance-qa-dataset
Dataset Card for Career Guidance Dataset
Dataset Overview
This dataset provides career guidance information for a variety of career roles. It includes questions and answers related to career roles such as "Data Scientist," "Software Engineer," "Product Manager," and many more. The dataset covers aspects like job responsibilities, required skills, career progression, salary expectations, and work environment. It is intended for use in building chatbot applications for… See the full description on the dataset page: https://huggingface.co/datasets/Pradeep016/career-guidance-qa-dataset.rag-qa-logs-corpus-data
🧠📚 RAG QA Logs & Corpus (Synthetic)
🧪 Multi-table synthetic RAG telemetry for quality, hallucinations, latency, and cost
A production-style, privacy-safe synthetic dataset that mimics telemetry exported from a real RAG system — from corpus → chunks → retrieval events → eval runs.
✅ Fully synthetic (no real users / orgs / PII).
⚡ Quick facts
Total rows: 103,255 across 6 linked tables
Labels (in eval_runs): is_correct, hallucination_flag, faithfulness_label… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/rag-qa-logs-corpus-data.deep-space-optical-chip-thermal-dataset
🚀 Deep Space Optical Chip Thermal Dataset 🪐
🌡️ 40,000 scenario-based prompt and response pairs on thermal mitigation for photonic chips in scientific instruments aboard deep-space probes, covering refractive index drift, waveguide misalignment, and thermal stress across materials, instruments, and environments.
⚠️ Disclaimer: All entries are synthetically generated. Material coefficients are drawn from published typical values, but no row is based on mission logs or flight… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/deep-space-optical-chip-thermal-dataset.Bitext-wealth-management-llm-chatbot-training-dataset
Bitext - Wealth Management Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Wealth Management] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-wealth-management-llm-chatbot-training-dataset.MoroccanHistory-QA-Dataset
