CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K195 likes8.2k downloads2y agoHugging Face02bitext /Bitext-retail-ecommerce-llm-chatbot-training-dataset Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.textquestion-answering10K<n<100K19 likes1.6k downloads2y agoHugging Face03bitext /Bitext-events-ticketing-llm-chatbot-training-dataset Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes1.1k downloads2y agoHugging Face04tokyotech-llm /Swallow-Nemotron-Post-Training-Dataset-v1 Swallow-Nemotron-Post-Training-Dataset-v1 The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below. Dataset Construction The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528. However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.texttext-generation1M<n<10M6 likes772 downloads7mo agoHugging Face05hoang14 /3112_llm_70b_trainingtext1M<n<10M0 likes642 downloads2y agoHugging Face06bitext /Bitext-retail-banking-llm-chatbot-training-dataset Bitext - Retail Banking Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail Banking] sector can be easily achieved using our two-step approach to LLM Fine-Tuning.… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-banking-llm-chatbot-training-dataset.textquestion-answering10K<n<100K17 likes339 downloads2y agoHugging Face07bitext /Bitext-telco-llm-chatbot-training-dataset Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes259 downloads2y agoHugging Face08quranlab /islamic-llm-training QuranLab — Qur'an and Hadith Training Mix Training-ready data derived from the QuranLab corpora: continued-pretraining text, grounded instruction data, preference pairs, verifiable prompts, retrieval pairs and a held-out evaluation set — all built on the same verse and ḥadīth keys as quranlab/quran and quranlab/hadith. QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/islamic-llm-training.texttext-generation1M<n<10M1 likes233 downloads2mo agoHugging Face09bitext /Bitext-insurance-llm-chatbot-training-dataset Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.textquestion-answering10K<n<100K8 likes221 downloads2y agoHugging Face10bitext /Bitext-travel-llm-chatbot-training-dataset Bitext - Travel Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Travel] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-travel-llm-chatbot-training-dataset.textquestion-answering10K<n<100K4 likes181 downloads2y agoHugging Face11ogulcanaydogan /Turkish-LLM-v10-Training Turkish LLM Training Dataset v10 A curated corpus of 144,022 Turkish instruction-completion pairs used to train the Turkish LLM Family. Dataset Description This dataset was created to address the scarcity of high-quality Turkish instruction-following data for language model fine-tuning. It covers a broad range of topics including: Science & Technology (physics, chemistry, biology, computer science) History & Geography (Turkish and world history, geography) General… See the full description on the dataset page: https://huggingface.co/datasets/ogulcanaydogan/Turkish-LLM-v10-Training.texttext-generation100K<n<1M3 likes104 downloads7mo agoHugging Face12bitext /Bitext-mortgage-loans-llm-chatbot-training-dataset Bitext - Mortgage and Loans Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Mortgage and Loans] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-mortgage-loans-llm-chatbot-training-dataset.textquestion-answering10K<n<100K5 likes102 downloads2y agoHugging Face13Faramir /Bitext-customer-support-llm-chatbot-training-dataset-spanish Spanish Customer Support LLM Chatbot Training Dataset Spanish-language adaptation of the Bitext Customer Support LLM Chatbot Training Dataset. This dataset is intended for training and evaluating Spanish-language customer-support chatbots and instruction-following large language models. Dataset Details Dataset Description This dataset is a Spanish translation and adaptation of the original Bitext Customer Support LLM Chatbot Training Dataset. The… See the full description on the dataset page: https://huggingface.co/datasets/Faramir/Bitext-customer-support-llm-chatbot-training-dataset-spanish.text10K<n<100K0 likes101 downloads16d agoHugging Face14bytedance-research /llm-training-alert-trace llm-training-alert-trace Dataset 📖 Overview llm-training-alert-trace is an open-source dataset comprising anonymized production-scale alert traces from large-scale Large Language Model (LLM) training jobs. It captures the interplay between system alert events and node health status in high-performance computing environments. This dataset is designed to foster research in: AIOps & Reliability Engineering: Automated fault diagnosis and root cause analysis. System… See the full description on the dataset page: https://huggingface.co/datasets/bytedance-research/llm-training-alert-trace.text100K<n<1M0 likes94 downloads21d agoHugging Face15bitext /Bitext-wealth-management-llm-chatbot-training-dataset Bitext - Wealth Management Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Wealth Management] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-wealth-management-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes90 downloads2y agoHugging Face16Ashu9675 /space-llm-training-data Space LLM Training Data (~1.27 Billion Tokens) A curated dataset of space and astronomy text for training language models, containing approximately 1.27 billion tokens collected from academic papers, arXiv abstracts, and educational web content. Dataset Summary File Size Est. Tokens Source jsalt_astroph_full.txt 2.88 GB ~862M 271K full astrophysics papers (abstract + introduction + conclusions) arxiv_astro_full.txt 360 MB ~108M 284K arXiv paper… See the full description on the dataset page: https://huggingface.co/datasets/Ashu9675/space-llm-training-data.texttext-generation1M<n<10M1 likes90 downloads4mo agoHugging Face17LLM-CLEM /Training-fr-basetext10K<n<100K0 likes84 downloads10mo agoHugging Face18bitext /Bitext-hospitality-llm-chatbot-training-dataset Bitext - Hospitality Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [hospitality] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-hospitality-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes83 downloads2y agoHugging Face19bitext /Bitext-media-llm-chatbot-training-dataset Bitext - Media Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [media] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-media-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes76 downloads2y agoHugging Face20Kubermatic /cncf-raw-data-for-llm-training CNCF Raw Data for LLM Training Description This dataset, named cncf-raw-data-for-llm-training, consists of markdown (MD) and PDF content extracted from various project repositories within the CNCF (Cloud Native Computing Foundation) landscape. The data was collected by fetching MD and PDF files from different CNCF project repositories and converting them into JSON format. This dataset is intended as raw data for training large language models (LLMs). The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/Kubermatic/cncf-raw-data-for-llm-training.text10K<n<100K0 likes64 downloads2y agoHugging Face21Kubermatic /cncf-question-and-answer-dataset-for-llm-training CNCF QA Dataset for LLM Tuning Description This dataset, named cncf-qa-dataset-for-llm-tuning, is designed for fine-tuning large language models (LLMs) and is formatted in a question-answer (QA) style. The data is sourced from PDF and markdown (MD) files extracted from various project repositories within the CNCF (Cloud Native Computing Foundation) landscape. These files were processed and converted into a QA format to be fed into the LLM model. The dataset includes the… See the full description on the dataset page: https://huggingface.co/datasets/Kubermatic/cncf-question-and-answer-dataset-for-llm-training.text10K<n<100K3 likes64 downloads2y agoHugging Face22bitext /Bitext-restaurants-llm-chatbot-training-dataset Bitext - Restaurants Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [restaurants] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-restaurants-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes62 downloads2y agoHugging Face23llmtraining-scraper /discord-messages Discord Messages Dataset Description This dataset contains 6.2 million anonymized messages extracted from public Discord servers. All personal identifying information (user IDs, server IDs, channel IDs, timestamps) has been removed. Only the raw message text remains. The data is formatted as plain text with one message per line, making it ideal for: Language model pre-training Fine-tuning chatbots Sentiment analysis Toxicity detection Slang and language evolution… See the full description on the dataset page: https://huggingface.co/datasets/llmtraining-scraper/discord-messages.texttext-generation1M<n<10M0 likes61 downloads2mo agoHugging Face24abhi23457 /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/abhi23457/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes50 downloads21d agoHugging Face25CJJones /Multiturn_Microcontroller-Arduino-LLM-TrainingThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. Dataset Description Repository: Arduino Project Generator Code Author: Cameron Jones Dataset Summary This dataset has no affiliation with the Arduino brand. It contains synthetic conversation examples generated by a Java-based Arduino project suggestion system. Each conversation follows a structured format where:A user interacts with a bot… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Multiturn_Microcontroller-Arduino-LLM-Training.text100K<n<1M4 likes48 downloads7mo agoHugging Face26krystv /code-corpus-llm-training Code Corpus for LLM Training Manually collected from top open-source repositories across: video/graphics editors, browsers, terminals, UI/UX, Qt/QML, Flutter, Rust, Python, ethical hacking, system-level, game engines, web frameworks, and more. Stats Records: 240,378 Raw text: 2,156,908,643 chars (~2.01 GB) Domains: 20 Domains web_ui: 32,354 records cpp: 29,792 records kotlin_android: 19,476 records ui_ux_design: 19,382 records rust: 15,440 records python:… See the full description on the dataset page: https://huggingface.co/datasets/krystv/code-corpus-llm-training.text100K<n<1M0 likes46 downloads5mo agoHugging Face27Talhat /customer_support_llm_training_traintest_splitThis dataset has been prepared for finetuning and downloaded from the repository "bitext/Bitext-customer-support-llm-chatbot-training-dataset". The dataset is the same as "Talhat/customer_support_llm_training" but divided into train_test split text10K<n<100K0 likes43 downloads2y agoHugging Face28CJJones /Elementary_Math_Word_Problems_LLM_Training_Short Dataset Card for Math Problem Generator Dataset Summary This dataset contains a subset of 100,000 procedurally generated math word problems, covering various mathematical concepts and difficulty levels. The problems were generated using a Java program that creates contextual word problems with detailed solutions and explanations. 🔗 Full dataset available on Gumroad The full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more?… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Elementary_Math_Word_Problems_LLM_Training_Short.textquestion-answering10K<n<100K0 likes42 downloads7mo agoHugging Face29strova-ai /resume-conversations-llm-training 📄 Resume Conversations for LLM Training High-quality conversational dataset for building AI that understands resumes, careers, and professional growth.Created and maintained by Syncora.ai. ✅ Overview This dataset provides resume-related conversations in a structured JSONL format, ideal for developers and AI practitioners working on chatbots, career advisory tools, or LLM fine-tuning. It includes realistic Q&A on career development, technology trends, and professional… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/resume-conversations-llm-training.texttext-generationn<1K3 likes37 downloads1y agoHugging Face30CJJones /Synthetic_Java_Dialog_And_Programs_LLM_TrainingThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. Java Programming Examples Dataset Dataset Description This dataset contains 8 distinct Java programs with 10 conversational examples each, synthetically generated from a larger dataset of 80+ programs. Each program has 10,000 variants, providing a diverse set of Java code examples covering various programming… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Synthetic_Java_Dialog_And_Programs_LLM_Training.textquestion-answering1K<n<10K1 likes36 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.