CoolFace
10 results

evaluation-benchmark

amazon /music-off-policy-evaluation-benchmark Music Off-Policy Evaluation Dataset Music Off-Policy Evaluation Dataset is a dataset designed for Off-Policy Evaluation (OPE) research. It contains logged interactions from the home page of Amazon Music. Use cases: Benchmarking OPE estimators Evaluating counterfactual ranking policies offline License Music Off-Policy Evaluation Benchmark © 2026 by Amazon is licensed under Creative Commons Attribution-NonCommercial 4.0 International.… See the full description on the dataset page: https://huggingface.co/datasets/amazon/music-off-policy-evaluation-benchmark.1M<n<10M0 likes1.4k downloads2mo agoHugging Faceminhhien0811 /java_evaluation_benchmarkstext1K<n<10K0 likes548 downloads4mo agoHugging Faceplnguyen2908 /AudioVisual-Benchmark-Evaluation AudioVisual Benchmark Evaluation — evaluation subsets Item-id lists for the audio-visual benchmark subsets used in our reported evaluation tables. Layout <benchmark>/eval_subset.csv item ids evaluated in the paper <benchmark>/media_index.csv id -> media filename(s) <benchmark>/media/ the media files those ids refer to eval_subset.csv holds a single id column keyed to the source benchmark (question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.audiomultiple-choice10K<n<100K0 likes214 downloads25d agoHugging Facellmsql-bench /benchmark-evaluation-resultstext1M<n<10M0 likes65 downloads7mo agoHugging Faceaiagentkarl /agent-evaluation-benchmark Agent Evaluation Benchmark A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories. Overview This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more. Categories Category Test Cases Description Weather 5 Forecasts, UV index, climate history Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.texttext-generationn<1K0 likes54 downloads6mo agoHugging FacePlumloom /evaluation-reliability-benchmark Plumloom Evaluation Reliability Benchmark Public results from Plumloom’s research on reliability in single-turn AI chat evaluations. This dataset contains 45 controlled single-turn chat cases used to test whether repeated evaluation produces more stable model choices than one-shot evaluation. It also includes aggregate results from 38 measurement-reliability evaluation runs. Private prompts, rubrics, model responses, and production execution details are intentionally excluded. n<1K1 likes32 downloads2mo agoHugging Face