datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-agent-eval
Arabic Agent Eval — Dataset Card
An open, installable Arabic function-calling benchmark with dialect splits.
Dataset summary
51 evaluation items spanning 6 categories and 5 dialects of Arabic, testing whether large language models can (a) select the right tool, (b) extract arguments from natural Arabic instructions, (c) preserve Arabic text in tool arguments instead of transliterating, and (d) understand dialectal framing.
Supported tasks… See the full description on the dataset page: https://huggingface.co/datasets/Mosescreates/arabic-agent-eval.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/mostafafhasjk/Bitext-customer-support-llm-chatbot-training-dataset.Kos_Mos_project_data
Model / Dataset Overview | 项目概述
Kos-Mos v0.1.0 is an early research release from the Kos-Mos Project, which explores role-oriented and persona-aligned large language models.
This dataset supports training models that maintain a consistent identity and expressive style under minimal prompting.
Kos-Mos v0.1.0 是 Kos-Mos 项目 的早期研究版本,聚焦于角色化与人格对齐的大语言模型。
该数据集用于支持在低提示词条件下保持稳定身份与表达风格的模型训练。
Long-Term Goal | 长期目标
The long-term goal of the Kos-Mos Project is to investigate… See the full description on the dataset page: https://huggingface.co/datasets/Hengzongshu/Kos_Mos_project_data.shortcutQADataset Summary: ShortcutQA is a question answering dataset designed to test whether language models rely on shallow shortcuts instead of real understanding. It includes examples where the context has been edited to include misleading clues (called shortcut triggers), automatically inserted using GPT-4. These edits can cause the model to answer incorrectly, revealing its vulnerability.
Languages: English
Usage: Use this dataset to evaluate how robust QA models are to misleading context edits.… See the full description on the dataset page: https://huggingface.co/datasets/Mosh/shortcutQA.
