datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
music-off-policy-evaluation-benchmark
Music Off-Policy Evaluation Dataset
Music Off-Policy Evaluation Dataset is a dataset designed for Off-Policy Evaluation (OPE) research. It contains logged interactions from the home page of Amazon Music.
Use cases:
Benchmarking OPE estimators
Evaluating counterfactual ranking policies offline
License
Music Off-Policy Evaluation Benchmark © 2026 by Amazon is licensed under Creative Commons Attribution-NonCommercial 4.0 International.… See the full description on the dataset page: https://huggingface.co/datasets/amazon/music-off-policy-evaluation-benchmark.toxicity_benchmark-evaluationgovreport_evaluation_benchmarkCrab-role-playing-evaluation-benchmark
📄 Paper
|
📄 Github
💬 Role-playing Model
|
💬 Role-palying Evaluation Model
💬 Training Dataset
|
💬 Evaluation Benchmark
|
💬 Annotated Role-playing Evaluation Dataset
|
💬 Human-preference Dataset
1. Introduction
This is the dataset used for evalauating a role‑playing LLM.
More details can be seen at GitHub and Crab… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/Crab-role-playing-evaluation-benchmark.benchmark-evaluation-meter
Shaer-AI/benchmark-evaluation-meter
This dataset extends Shaer-AI/benchmark-evaluation by adding BiLSTM-based meter probability distributions.
Added columns
*_meter_dist_json: JSON with:
pred
pred_prob
probs (probability per meter label)
