datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.llm-election-data-2024
Data Release for Large-Scale, Longitudinal Survey of Large Language Models (LLMs) During the 2024 US Elections
Overview
This repository contains the questions asked of and responses given by LLMs during the 2024 US elections, collected for a longitudinal survey conducted from July 23, 2024 to November 12, 2024. The study is described in detail in the paper "Large-Scale, Longitudinal Study of Large Language Models During the 2024 US Election Season" by Sarah H. Cen, Andrew… See the full description on the dataset page: https://huggingface.co/datasets/sarahcen/llm-election-data-2024.Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.KoHRM-Text-1.4B-sft-lora-data
KoHRM-Text-1.4B SFT and LoRA Prepared Data
This dataset repo stores curated KoHRM SFT/LoRA subsets in the same tokenized
HRM-Text V1Dataset format used by training. It is intended for quick behavior
alignment experiments after KoHRM pretraining.
Model repo:
https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B
Code repo:
https://github.com/LLM-OS-Models/KoHRM-text
Format
Each folder is a prepared V1Dataset:
<dataset-name>/
metadata.json
tokenizer_info.json… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-sft-lora-data.agent-llm-traces-v2
Exgentic Agent LLM Traces v2 — Agent Chat Only
OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it.
This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text
Dataset Overview
A collection of 27 domains (“topics”) and 3100 question-answer pair.
Each topic comes with average 117 QA pairs.Every QA entry comes with:
references: one or more source files the answer is extracted from
time with each reference comes the starting and ending time the answer is extracted from the reference
video_files: the video files where the answer can be found
(future) video title & description from metadata.csv
File structure
You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.Bitext-retail-ecommerce-llm-chatbot-training-dataset
Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.serbian-llm-benchmark
Serbian LLM Evaluation Dataset
Welcome to the Serbian LLM Evaluation Dataset, your one-stop solution for evaluating Serbian Language Models (LLMs) like never before! This comprehensive toolkit empowers you to measure model performance across diverse domains in Serbian, ensuring your models are smarter, faster, and more intuitive. Whether you're a researcher, developer, or just an enthusiast—this dataset is tailor-made to help your LLM thrive.
🔍 What's Inside?
This… See the full description on the dataset page: https://huggingface.co/datasets/datatab/serbian-llm-benchmark.Bitext-events-ticketing-llm-chatbot-training-dataset
Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.NacidCette dataset est le corpus d'entraînement principal du SLM lam (lam-1, lam-2, lam-3, et supérieur.),
il est agrandit régulièrement pour le développement des nouvelles itérations et versions du SLM lam et plus largement, de la série de SLM lamina.
Nous ne cherchons pas à faire de lam, un LLM toute de suite en une traite, mais un SLM progressif
🛑 : Lam, sur toutes ses iterations et modèles (Lam-1, Lam-2, Lam-3 , et supérieur etc...), sont des créations de Clemylia, et du studio LES-IA-ETOILES.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-CLEM/Nacid.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.DEBATE
DEBATE: Diverse Multi-Agent Debates
This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework".
Citation
comming soon.
HalluEditBench
Can Knowledge Editing Really Correct Hallucinations?
Respository Oveview: This repository contains the code, results and dataset for the paper "Can Knowledge Editing Really Correct Hallucinations? (ICLR 2025)"
TLDR: We proposed HalluEditBench to holistically benchmark knowledge editing methods in correcting real-world hallucinations on five dimensions including Efficacy, Generalization, Portability, Locality, and Robustness. We find that their effectiveness could be far from what… See the full description on the dataset page: https://huggingface.co/datasets/llm-editing/HalluEditBench.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our leaderboard at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.JGLUEJGLUE, Japanese General Language Understanding Evaluation, is built to measure the general NLU ability in Japanese. JGLUE has been constructed from scratch without translation. We hope that JGLUE will facilitate NLU research in Japanese.Global-LLMs-Replies
Global LLMs Replies
GPT-4o
-> 74,644 rows
mixtral-8x22b
-> 13,129 rows
claude-3-haiku
-> 3,871 rows
hermes3-uk
Dataset Card for Hermes 3 Ukrainian Fixed Conversations
Dataset Description
Dataset Summary
hermes3-uk-fixed is a Ukrainian translation of the [NousResearch/Hermes-3-Dataset]. The translation was produced with Gemma 3 27B (instruction-tuned). During preparation we removed all system prompts and normalized the message roles and content to match the common schema we use across our dialog datasets.
Languages
Ukrainian (uk)
Dataset Structure
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/hermes3-uk.spai-ss6-llm-1b-thai-corpus
Thai Medical And Health Corpus
Thai public medical and health web corpus collected for research and LLM dataset
experimentation, with optional imported Thai medical/health datasets from
Hugging Face stored as separate configs.
Public Web Corpus
Config: default
Split: train
Records: 3660 deduplicated articles
Columns: 16
Format: Parquet
Latest collection profile: free_1000
Latest generated at: 2026-06-06T17:41:38.787978+00:00
Source And Method
The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.CBT-Bench
CBT-Bench Dataset
Overview
CBT-Bench is a benchmark dataset designed to evaluate the proficiency of Large Language Models (LLMs) in assisting cognitive behavior therapy (CBT). The dataset is organized into three levels, each focusing on different key aspects of CBT, including basic knowledge recitation, cognitive model understanding, and therapeutic response generation. The goal is to assess how well LLMs can support various stages of professional mental health care… See the full description on the dataset page: https://huggingface.co/datasets/Psychotherapy-LLM/CBT-Bench.Multi-turn_Long-context_Benchmark_for_LLMs
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
Arxiv: https://www.arxiv.org/abs/2507.13681
Huggingface: https://huggingface.co/papers/2507.13681
Introduction
LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios.
Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.legal-llm-benchmark
Legal LLM Benchmark Dataset
Safety-Utility Trade-offs in Legal AI: An LLM Evaluation Across 12 Models
Quick Start
from datasets import load_dataset
# Core datasets
questions = load_dataset("marvintong/legal-llm-benchmark", "questions")
phase1_evals = load_dataset("marvintong/legal-llm-benchmark", "phase1_evaluations")
phase3_evals = load_dataset("marvintong/legal-llm-benchmark", "phase3_evaluations")
# Additional helpful datasets
contracts =… See the full description on the dataset page: https://huggingface.co/datasets/marvintong/legal-llm-benchmark.Twin-2K-500-Mega-Study
Twin-2K-500-Mega-Study Dataset
GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study
To see more details for how to process these data, please refer to this GitHub repository.
This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants).
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.assist-llm-function-calling
Function Calling dataset for Assist LLM for Home Assistant
This dataset is generated by using other conversation agent pipelines as teachers
from the deivce-actions-v2 dataset.
This dataset is used to support fine tuning of llama based models.
See Device Actions for a notebook for construction of this dataset and the device-actions dataset.
hermes3-en-fixed
Dataset Card for Hermes 3 Fixed Conversations
Dataset Description
Dataset Summary
hermes3-en-fixed is a [NousResearch/Hermes-3-Dataset]. During preparation we removed all system prompts and normalized the message roles and content to match the common schema we use across our dialog datasets.
Languages
English (en)
Dataset Structure
Data Fields
conversations: list of messages in a dialog (array of objects)
from: normalized sender role — user or assistant… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/hermes3-en-fixed.Grounded_3D_LLM_with_Referent_Tokens_Dataset
Grounded 3D-LLM Dataset
For detailed information and resources, please visit the following links:
Paper
Arxiv
Project Website
Dataset Access
Code
We are in the process of releasing our data incrementally:
Processed ScanNet200 PCD(~7G):
Each .npyfile represents a N*12 array with the following structure:
coordinates, color, normals, segments, labels = (
points[:, :3],
points[:, 3:6],
points[:, 6:9],
points[:, 9]… See the full description on the dataset page: https://huggingface.co/datasets/ShuaiYang03/Grounded_3D_LLM_with_Referent_Tokens_Dataset.jyotish-llm-sft
Jyotish-LLM SFT Dataset
1,000,000 grounded instruction-tuning examples for Vedic astrology chart
interpretation. Every example pairs a deterministically-computed birth
chart with a response whose every claim traces back to either a chart
fact or a cited classical yoga rule — built specifically so a model
trained on it learns to cite, not free-associate.
Total examples
1,000,000
Train / Validation / Test
959,625 / 20,575 / 19,800
Underlying unique birth charts… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/jyotish-llm-sft.or-bench-toxic-all
OR-Bench: An Over-Refusal Benchmark for Large Language Models
This dataset constains highly toxic prompts, use with caution!!!
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.Bitext-retail-banking-llm-chatbot-training-dataset
Bitext - Retail Banking Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail Banking] sector can be easily achieved using our two-step approach to LLM Fine-Tuning.… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-banking-llm-chatbot-training-dataset.wiki-instruction-dialogs
Dataset Card for Ukrainian Wiki Instruction Dialogs
Dataset Description
Dataset Summary
This dataset is a processed version of cleansed Wikipedia part of the lang-uk UberText 2.0 Corpus. The answers were generating using Gemma 3 12B. The goal for this dataset is to provide question answering knowledge grounding for NER, summarization, paraphrase, and simplification tasks.
Languages
Ukrainian (uk)
Data Fields
task: Task type
instruction: Text containing task… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/wiki-instruction-dialogs.
