datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.Bitext-retail-ecommerce-llm-chatbot-training-dataset
Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.Bitext-events-ticketing-llm-chatbot-training-dataset
Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.stack-exchange-dataset
Overview
This dataset consists of three TSV files, namely: cs.tsv, ds.tsv, and p.tsv.
Each file includes the data for the questions asked on a Stack Exchange (SE) question-answering community, from the creation of the community until May 2021.
cs.tsv --> Computer Science SE
ds.csv --> Data Science SE
p.csv --> Political Science SE
File Structure
Each file has the following columns:
id: the question id
title: the title of the question
body: the body or text of the… See the full description on the dataset page: https://huggingface.co/datasets/habedi/stack-exchange-dataset.Bitext-telco-llm-chatbot-training-dataset
Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.Flutter-Code-with-Questions-Dataset-Turkish
Flutter Code with Questions Dataset (Turkish)
📦 Dataset Name: flutter_code_with_questions
Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir.
📁 Dataset Format
Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.Bitext-insurance-llm-chatbot-training-dataset
Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.Flutter-Code-with-Questions-Dataset-English
🧠 Flutter Code with Questions Dataset (English)
This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development.
📂 Dataset Structure
The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes:
A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.statcan-dialogue-dataset-retrieval
Statcan Dialogue Dataset (Processed for Retrieval Tasks)
This is a variant of the Statcan Dialogue Dataset, which we processed specifically for multilingual retrieval (english, french). It contains everything in CSVs, rather than having metadata hosted separately.
Quickstart
from datasets import load_dataset
repo = 'McGill-NLP/statcan-dialogue-dataset-retrieval'
# load english queries, training split
queries_en = load_dataset(repo, 'queries_english', split='train') #… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/statcan-dialogue-dataset-retrieval.Bitext-travel-llm-chatbot-training-dataset
Bitext - Travel Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Travel] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-travel-llm-chatbot-training-dataset.k-beauty-ai-citation-dataset
K-Beauty AI Citation Dataset
Open dataset mapping Korean K-beauty entities (ingredients, skin concerns, use cases, brands) and answer-style guides to citation-shaped external references. Designed to be referenced by AI search engines, content builders, and SEO research.
Canonical source: https://kbeautyanswers.com/dataset/
License: CC BY 4.0
Maintainer: K-Beauty Answers (site)
Initial release: 2026-05-23
What's in it
128 entities (37 ingredients + 18 skin… See the full description on the dataset page: https://huggingface.co/datasets/k-master/k-beauty-ai-citation-dataset.vecombot-dataset
VECOM — Bộ dữ liệu thị trường Thương mại điện tử Việt Nam (VEComBot)
Bộ dữ liệu phụ lục cho đồ án tốt nghiệp VEComBot — hệ thống Đa tác tử (Multi-Agent
System) phân tích và tổng hợp thị trường Thương mại điện tử Việt Nam (VECOM). Đây là
kho tài liệu nguồn và corpus đã qua xử lý (figure-aware) được nạp vào PostgreSQL/pgvector
để phục vụ cả nhánh MAS lẫn nhánh baseline naive RAG.
Mục đích: dùng cho nghiên cứu học thuật và tái lập kết quả đồ án. Các báo cáo gốc là
ấn phẩm công… See the full description on the dataset page: https://huggingface.co/datasets/binhtran23/vecombot-dataset.pentesting-dataset
Dataset Card for Penetration Testing Dataset
This dataset card aims to provide essential information about the Penetration Testing Dataset, which includes various resources and scripts useful for penetration testing and cybersecurity research.
Dataset Details
Dataset Description
The Penetration Testing Dataset is a collection of scripts, tools, and vulnerability data designed for cybersecurity professionals to facilitate penetration testing tasks.… See the full description on the dataset page: https://huggingface.co/datasets/me-aas/pentesting-dataset.Ncert_datasetpentesting-dataset
Dataset Card for Penetration Testing Dataset
This dataset card aims to provide essential information about the Penetration Testing Dataset, which includes various resources and scripts useful for penetration testing and cybersecurity research.
Dataset Details
Dataset Description
The Penetration Testing Dataset is a collection of scripts, tools, and vulnerability data designed for cybersecurity professionals to facilitate penetration testing tasks. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/boapro/pentesting-dataset.pentesting_dataset
Dataset Card for Penetration Testing Dataset
This dataset card aims to provide essential information about the Penetration Testing Dataset, which includes various resources and scripts useful for penetration testing and cybersecurity research.
Dataset Details
Dataset Description
The Penetration Testing Dataset is a collection of scripts, tools, and vulnerability data designed for cybersecurity professionals to facilitate penetration testing tasks. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Canstralian/pentesting_dataset.Bitext-mortgage-loans-llm-chatbot-training-dataset
Bitext - Mortgage and Loans Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Mortgage and Loans] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-mortgage-loans-llm-chatbot-training-dataset.cyber_MITRE_CTI_dataset_v15This dataset is a specialized resource designed for training and evaluating question-answering models in the context of Cyber Threat Intelligence (CTI), specifically targeting the identification of tactics and techniques based on natural language descriptions of cyber-attacks. The dataset is derived from the MITRE ATT&CK framework (version 15) and contains annotated pairs of sentences and their corresponding tactics and techniques. The primary goal is to assist automated systems in… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/cyber_MITRE_CTI_dataset_v15.Bitext-wealth-management-llm-chatbot-training-dataset
Bitext - Wealth Management Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Wealth Management] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-wealth-management-llm-chatbot-training-dataset.career-guidance-qa-dataset
Dataset Card for Career Guidance Dataset
Dataset Overview
This dataset provides career guidance information for a variety of career roles. It includes questions and answers related to career roles such as "Data Scientist," "Software Engineer," "Product Manager," and many more. The dataset covers aspects like job responsibilities, required skills, career progression, salary expectations, and work environment. It is intended for use in building chatbot applications for… See the full description on the dataset page: https://huggingface.co/datasets/Pradeep016/career-guidance-qa-dataset.MoroccanHistory-QA-DatasetBitext-hospitality-llm-chatbot-training-dataset
Bitext - Hospitality Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [hospitality] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-hospitality-llm-chatbot-training-dataset.deep-space-optical-chip-thermal-dataset
🚀 Deep Space Optical Chip Thermal Dataset 🪐
🌡️ 40,000 scenario-based prompt and response pairs on thermal mitigation for photonic chips in scientific instruments aboard deep-space probes, covering refractive index drift, waveguide misalignment, and thermal stress across materials, instruments, and environments.
⚠️ Disclaimer: All entries are synthetically generated. Material coefficients are drawn from published typical values, but no row is based on mission logs or flight… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/deep-space-optical-chip-thermal-dataset.xai-questions-datasetExplore the questions users have for robots across a diverse set of situations!
You can read the paper here: What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics!
from datasets import load_dataset
dataset = load_dataset("lwachowiak/xai-questions-dataset")
dataset['train'][0]
The analysis code can be found on GitHub
Paper Abstract
With the increased use of large language models and conversational interfaces in human–robot… See the full description on the dataset page: https://huggingface.co/datasets/lwachowiak/xai-questions-dataset.Bitext-media-llm-chatbot-training-dataset
Bitext - Media Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [media] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-media-llm-chatbot-training-dataset.Dense-Information-Science-Physics-Dataset
Dense Information With Multiple Fine-tuned Variations
This dataaset has multiple for each input to learn how to express the same answer in different ways
Dataset Structure
The dataset contains two columns:
Column
Description
input
A science or quantum-physics question
output
A conversational answer to the question
Example:
{
"input": "What is quantum entanglement?",
"output": "Quantum entanglement is when two quantum systems share one… See the full description on the dataset page: https://huggingface.co/datasets/StarpowerTechnology/Dense-Information-Science-Physics-Dataset.datasets
LogicPoison: Logical Attacks on Graph Retrieval-Augmented Generation
This repository contains the datasets for LogicPoison, a logical poisoning framework for Graph-based Retrieval-Augmented Generation (GraphRAG) systems.
Paper: LogicPoison: Logical Attacks on Graph Retrieval-Augmented Generation
GitHub Repository: Jord8061/logicPoison
Overview
LogicPoison targets the topological integrity of knowledge graphs used in GraphRAG. Instead of injecting false content… See the full description on the dataset page: https://huggingface.co/datasets/Jord8061/datasets.Bitext-restaurants-llm-chatbot-training-dataset
Bitext - Restaurants Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [restaurants] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-restaurants-llm-chatbot-training-dataset.turkish_llm_finetune_dataset_4_topics
Turkish LLM Finetune Dataset - 4 Topics
This dataset is designed to fine-tune the T3 AI Turkish LLM. It was created by Barathan Aslan, Ömer Faruk Çelik, and Batuhan Kalem for the T3 AI Hackathon. The dataset focuses on four distinct topics: Agriculture, Sustainability, Turkish Education Sytem, and Turkish Law System.
Contributors
Barathan Aslan (https://huggingface.co/barathanasln)
Batuhan Kalem(https://huggingface.co/Pancarsuyu)
Ömer Faruk Çelik… See the full description on the dataset page: https://huggingface.co/datasets/barathanasln/turkish_llm_finetune_dataset_4_topics.Financial_Context_DatasetThis dataset contains over 50,000 samples of user financial queries paired with their corresponding structured data requests (context). It was created to facilitate the creation of the Financial Agent LLM for accurate data extraction and query answering.
How to load the Dataset
You can load the dataset using the code below:
from datasets import load_dataset
ds = load_dataset("Chaitanya14/Financial_Context_Dataset")
Dataset Construction
Diverse Query Sources… See the full description on the dataset page: https://huggingface.co/datasets/Chaitanya14/Financial_Context_Dataset.
