datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MuSR
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
Creating murder mysteries that require multi-step reasoning with commonsense using ChatGPT!
By: Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, and Greg Durrett.
View the dataset on our custom viewer and project website!
Check out the paper. Appeared at ICLR 2024 as a spotlight presentation!
Git Repo with the source data, how to recreate the dataset (and create new ones!) here
tmmluplus
TMMLU+ : Large scale traditional chinese massive multitask language understanding
iKala presents TMMLU+, a large-scale benchmark for evaluating LLM capabilities in Traditional Chinese, with content primarily reflecting Taiwan's linguistic, educational, and professional contexts. It covers 66 subjects, from elementary to professional domains, and is approximately six times larger than TMMLU with broader, more balanced coverage.
TMMLU+ v1.1 improves benchmark quality through… See the full description on the dataset page: https://huggingface.co/datasets/ikala/tmmluplus.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.taxbench-au
TaxBench-AU
A benchmark for testing whether AI agents can calculate Australian tax.
TaxBench-AU contains 156 Australian tax calculation questions, presented as multiple-choice (4-option) worked tax problems. The benchmark is designed to test whether an AI agent can read the facts, apply the right Australian tax rule for the relevant income year, do the calculation, and choose the correct answer.
The Kaggle mirror is published as Agent Tax Exam for Australian Tax.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/Pn101/taxbench-au.TruthfulQA
Dataset Card for TruthfulQA
Dataset Summary
TruthfulQA: Measuring How Models Mimic Human Falsehoods
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/domenicrosati/TruthfulQA.customer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits… See the full description on the dataset page: https://huggingface.co/datasets/Tobi-Bueck/customer-support-tickets.thai-onet-m6-exam
Thai O-Net Exams Dataset
Overview
The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems.
Dataset Source
Thai National Institute of Educational Testing Service (NIETS)
Maintainer
Dr. Kobkrit Viriyayudhakorn… See the full description on the dataset page: https://huggingface.co/datasets/matichon/thai-onet-m6-exam.function_calling_extended
Trelis Function Calling Dataset
UPDATE: As of Dec 5th 2023, there is a v3 of this dataset now available from here.
Allows models to be fine-tuned for function-calling.
The dataset is human generated and does not make use of Llama 2 or OpenAI!
Contains 59 training and 17 test rows
Based on eight functions: search_bing, search_arxiv, save_chat, read_json_file, list_files, get_current_weather, delete_file, clear_chat
Access this dataset by purchasing a license HERE.
Alternatively… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/function_calling_extended.Bitext-retail-ecommerce-llm-chatbot-training-dataset
Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.thai-onet-m6-exam
Thai O-Net Exams Dataset
Overview
The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems.
Dataset Source
Thai National Institute of Educational Testing Service (NIETS)
Maintainer
Dr. Kobkrit… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-onet-m6-exam.Bitext-events-ticketing-llm-chatbot-training-dataset
Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.Traditional-Chinese-Medicine-Multiple_choice_question
Discription
This dataset is sourced from the website of the Ministry of Examination, R.O.C (Taiwan) and contains past exam questions from the national Traditional Chinese Medicine examinations in Taiwan. The exam comprises six subjects. This dataset specifically includes questions from two subjects, including the History of Traditional Chinese Medicine, Basic Theories of Traditional Chinese Medicine, Neijing, Nanjing, Traditional Chinese Medicine Prescription Studies, and… See the full description on the dataset page: https://huggingface.co/datasets/Liavan/Traditional-Chinese-Medicine-Multiple_choice_question.lc_quad_synth
LC-QuAD 2.0-synth
Dataset Summary
This dataset is an updated version of the LC-QuAD 2.0 dataset which includes LLM-based natural language translations of the corresponding wikidata queries. It also includes
verifier scores for the LLM translations and the original translations indicating the probability that the translation is correct (for details see our linked GitHub Repository).
It contains 19000 examples of queries and translations. It can be used for training and… See the full description on the dataset page: https://huggingface.co/datasets/timschwa/lc_quad_synth.trivia-qa-20k
Dataset Description
This dataset, trivia-qa-20k, is a cleaned and simplified version of the popular mandarjoshi/trivia_qa dataset. It contains 20,000 high-quality, simple question-answer pairs in English.
The data is structured for ease of use in fine-tuning models for straightforward question-answering tasks, where context or evidence is not required.
Note: This is a cleaned and simplified version of the mandarjoshi/trivia_qa dataset.
Dataset Structure
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/prajwalmani/trivia-qa-20k.WikiRAG-TR
Dataset Summary
WikiRAG-TR is a dataset of 6K (5999) question and answer pairs which synthetically created from introduction part of Turkish Wikipedia Articles. The dataset is created to be used for Turkish Retrieval-Augmented Generation (RAG) tasks.
Dataset Information
Number of Instances: 5999 (5725 synthetically generated question-answer pairs, 274 augmented negative samples)
Dataset Size: 20.5 MB
Language: Turkish
Dataset License: apache-2.0
Dataset Category:… See the full description on the dataset page: https://huggingface.co/datasets/Metin/WikiRAG-TR.ShareChat
ShareChat: A Dataset of Chatbot Conversations in the Wild
Paper | Github
This dataset contains 142,808 real-world user conversations across multiple conversational AI platforms (ChatGPT, Claude, Gemini, Grok, and Perplexity).
The dataset is collected and processed for research purposes to understand usage patterns, topic distributions, and behavioral characteristics across different AI platforms.
Update
5 Apr 2026: The dataset is updated with an additional column for… See the full description on the dataset page: https://huggingface.co/datasets/tucnguyen/ShareChat.TwoHopFactThis is the dataset introduced in the paper Do Large Language Models Latently Perform Multi-Hop Reasoning?.
Code: https://github.com/google-deepmind/latent-multi-hop-reasoning.
or-bench-toxic-all
OR-Bench: An Over-Refusal Benchmark for Large Language Models
This dataset constains highly toxic prompts, use with caution!!!
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.ARCHIVE-TEXT-URLS
Internet Archive English Text URLs Dataset
Dataset Description
This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials.
Dataset Summary
Total Rows: 11,151,637
Language: English
Source: Internet Archive
Format: CSV with metadata and direct text file URLs
Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.Security-TTP-Mapping
The Security Attack Pattern (TTP) Recognition or Mapping Task
We share in this repo the MITRE ATT&CK mapping datasets, with training, validation and test splits.
The datasets can be considered as an emerging and challenging multilabel classification NLP task, with over 600 hierarchical classes.
NOTE: due to their security nature, these datasets contain textual information about malware and other security aspects.
Datasets
TRAM
This dataset belongs to CTID… See the full description on the dataset page: https://huggingface.co/datasets/tumeteor/Security-TTP-Mapping.question-type-and-complexity
Question Type and Complexity (QTC) Dataset
Dataset Overview
The Question Type and Complexity (QTC) dataset is a comprehensive resource for linguistics/NLP research focusing on question classification and linguistic complexity analysis across multiple languages. It contains questions from two distinct sources (TyDi QA and Universal Dependencies v2.15), automatically annotated with question types (polar/content) and a set of linguistic complexity features.
Key Features:
2… See the full description on the dataset page: https://huggingface.co/datasets/rokokot/question-type-and-complexity.Trueque-Benchmark-beta-0.1
🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture
🌐 Language versions: Español | Português
⚠️ Official Disclaimer: Beta Release (v0.1)
Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America.
Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.Bitext-telco-llm-chatbot-training-dataset
Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.Target-QA
🎯 Target-QA: The First QA Dataset Benchmarking Target Priorization Based on DepMap
📑 Dataset Summary
Target-QA is derived from the DepMap multi-omics and CRISPR screening cohorts, harmonized via BioMedGraphica.It enables multi-modal reasoning by combining numeric evidence, topological knowledge and language context for CRISPR target prioritization.
This dataset supports the training and benchmarking of… See the full description on the dataset page: https://huggingface.co/datasets/FuhaiLiAiLab/Target-QA.Flutter-Code-with-Questions-Dataset-Turkish
Flutter Code with Questions Dataset (Turkish)
📦 Dataset Name: flutter_code_with_questions
Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir.
📁 Dataset Format
Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.Bitext-insurance-llm-chatbot-training-dataset
Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.Reverse-hybrid-train-no-persona-meanReverse-alpha-suppression-task-boostOriginal-alpha-suppression-task-boosttiny-singleturn-chat-ko
