datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cje-chatbot-arena
CJE Chatbot Arena Dataset
Dataset from Causal Judge Evaluation experiments on Chatbot Arena data.
Dataset Structure
cje_dataset.jsonl - Main dataset with judge scores and oracle labels (4,961 prompts)
prompts.jsonl - Original Chatbot Arena prompts
responses/ - Model responses for each policy variant
logprobs/ - Token logprobs for importance sampling estimators
Policies
5 system prompt variants evaluated:
base - No system prompt
clone - "Respond exactly as… See the full description on the dataset page: https://huggingface.co/datasets/elandy/cje-chatbot-arena.llm-jp-chatbot-arena-conversations
LLM-jp Chatbot Arena Conversations Dataset
This dataset contains approximately 1,000 conversations with pairwise human preferences, most of which are in Japanese.
The data was collected during the trial phase of the LLM-jp Chatbot Arena (January–February 2025), where users compared responses from two different models in a head-to-head format.
Each sample includes a question ID, the names of the two models, their conversation transcripts, the user's vote, an anonymized user ID, a… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-chatbot-arena-conversations.
