datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CXM_Arena
Dataset Card for CXM Arena Benchmark Suite
Dataset Description
This dataset, "CXM Arena Benchmark Suite," is a comprehensive collection designed to evaluate various AI capabilities within the Customer Experience Management (CXM) domain. It consolidates five distinct tasks into a unified benchmark, enabling robust testing of models and pipelines in business contexts. The entire suite was synthetically generated using advanced large language models, primarily… See the full description on the dataset page: https://huggingface.co/datasets/sprinklr-huggingface/CXM_Arena.arenaCK-Arena
CK-Arena Dataset
Overview
This is the official dataset for CK-Arena, a multi-agent benchmark designed to evaluate whether large language models (LLMs) truly master concept-level knowledge.
CK-Arena operationalises concept understanding through a language-based social deduction game (Undercover): LLM players receive closely related word concepts and must describe their assigned concept naturally, while LLM judges score each statement. By running many games… See the full description on the dataset page: https://huggingface.co/datasets/Xushuhaha/CK-Arena.llm-jp-chatbot-arena-conversations
LLM-jp Chatbot Arena Conversations Dataset
This dataset contains approximately 1,000 conversations with pairwise human preferences, most of which are in Japanese.
The data was collected during the trial phase of the LLM-jp Chatbot Arena (January–February 2025), where users compared responses from two different models in a head-to-head format.
Each sample includes a question ID, the names of the two models, their conversation transcripts, the user's vote, an anonymized user ID, a… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-chatbot-arena-conversations.ru-arena-hard
ru-arena-hard
This is translated version of arena-hard-auto dataset for evaluation LLMs. The translation of the original dataset was done manually. In addition, content of each task in dataset was reviewed, the correctness of the task statement and compliance with moral and ethical standards were assessed. Thus, this dataset allows you to evaluate the abilities of language models to support the Russian language.
Overview of the Dataset
Original dataset:… See the full description on the dataset page: https://huggingface.co/datasets/t-tech/ru-arena-hard.rag-qa-arena
RAG QA Arena Annotated Dataset
A comprehensive multi-domain question-answering dataset with citation annotations designed for evaluating Retrieval-Augmented Generation (RAG) systems, featuring faithful answers with proper source attribution across 6 specialized domains.
🎯 Dataset Overview
This annotated version of the RAG QA Arena dataset includes citation information and gold document IDs, making it ideal for evaluating not just answer accuracy but also answer grounding… See the full description on the dataset page: https://huggingface.co/datasets/rajistics/rag-qa-arena.SAGEO-Arena
SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization
SAGEO Arena is a benchmark for evaluating Search-Augmented Generative Engine Optimization (SAGEO) — the practice of optimizing web documents to improve their visibility in AI-generated responses.
Contents
This dataset releases the queries and Google Custom Search API results used to construct the SAGEO Arena corpus. Please follow the crawler instructions in the GitHub… See the full description on the dataset page: https://huggingface.co/datasets/yonsei-dli/SAGEO-Arena.phi3-arena-short-dpo
Dataset Summary
DPO (Direct Policy Optimization) dataset of normal and short answers generated from lmsys/chatbot_arena_conversations dataset using microsoft/Phi-3-mini-4k-instruct model.
Generated using ShortGPT project.
EldenRingQA
🗡️ Elden Ring QA Dataset
A domain-specific question-answering dataset for Elden Ring, covering weapons, bosses, armors, spells, NPCs, locations, creatures, skills, and ashes of war — including cross-entity boss vulnerability analysis and per-build weapon recommendations.
Uses
Intended Uses
Fine-tuning language models for Elden Ring domain-specific QA
Training instruction-following models on structured game knowledge
Retrieval-augmented generation (RAG)… See the full description on the dataset page: https://huggingface.co/datasets/ArenaRune/EldenRingQA.chatbot-arena-ja-calm2-7b-chat-experimental_dedupedchatbot-arena-ja-calm2-7b-chatからpromptが一致するデータを削除したデータセットです。
arenaarenaarena
