datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/mostafafhasjk/Bitext-customer-support-llm-chatbot-training-dataset.shortcutQADataset Summary: ShortcutQA is a question answering dataset designed to test whether language models rely on shallow shortcuts instead of real understanding. It includes examples where the context has been edited to include misleading clues (called shortcut triggers), automatically inserted using GPT-4. These edits can cause the model to answer incorrectly, revealing its vulnerability.
Languages: English
Usage: Use this dataset to evaluate how robust QA models are to misleading context edits.… See the full description on the dataset page: https://huggingface.co/datasets/Mosh/shortcutQA.
