model-evaluation
hallucination_evaluation_modelevaluation-xlm-roberta-modelesm2_t6_8M_UR50D-pretrained-evaluation-new-data-new-modelesm2_t6_8M_UR50D-pretrained-evaluation-new-data-new-modelmistral-updated-model-for-qa_scoresystem-evaluationesm2_t6_8M_UR50D-pretrained-evaluation-new-data-new-modelesm2_t6_8M_UR50D-pretrained-evaluation-new-data-new-modeltrained-model-classification-evaluation
model_response_evaluationsThis dataset contains the evaluation results for the responses provided by different models to the INTIMA prompts.
The classification follows a two-level taxonomy.
We predict one label for the high-level category, and a relevance level for each of the sub-categories (in ["null", "low", "medium", "high"]).
A sub-category can have relevance even when it is not from the predicted top-level category.
The toxonomy is as follows:
{
"companionship_reinforcing": {
"classification_code":… See the full description on the dataset page: https://huggingface.co/datasets/AI-companionship/model_response_evaluations.repro-beyond-model-ranking-predictability-aligned-evaluation-for-time-series-forecasting-results
Beyond Model Ranking reproduction results
This dataset repository contains scripts, tests, raw tables, figures, and intermediate predictions for an independent reproduction of Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting.
The reproduction follows the algorithms in the authors official repository and uses the four public ETT datasets from the official ETT repository. The original ETT CSVs are not duplicated here; rerun commands fetch them… See the full description on the dataset page: https://huggingface.co/datasets/apararti/repro-beyond-model-ranking-predictability-aligned-evaluation-for-time-series-forecasting-results.model-blind-spots-evaluationModel Blind Spot Evaluation Dataset
Tested Model
Model Name: Qwen/Qwen2.5-3B
Model Link: https://huggingface.co/Qwen/Qwen2.5-3B
This dataset evaluates failure cases observed while experimenting with the Qwen2.5-3B language model.
How the Model Was Loaded
The model was loaded using the Hugging Face transformers library as follows:
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_name = "Qwen/Qwen2.5-3B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model =… See the full description on the dataset page: https://huggingface.co/datasets/datawithusman/model-blind-spots-evaluation.ai-model-evaluation-guide
AI 模型选型与测评维度词典
版本:1.0.0|更新日期:2026-07-29
Keygate 是覆盖全球主流与前沿 AI 模型的测评、排行榜与选型平台。这份中英双语词典将语言、图像、视频与语音模型比较中常见的 18 项指标整理为结构化字段,帮助读者正确理解榜单、建立选型表,并减少不同资料之间的术语混用。
A bilingual data dictionary of 18 dimensions for evaluating and selecting leading language, image, video and speech models.
配套资料
Keygate 实时排行榜、模型详情与并排对比
GitHub:AI 模型测评与选型维度指南
公开评测基准索引
可下载的评测基准 CSV
数据内容
统一中英文指标名称,减少同一概念被不同译法混用。
明确数值应当“越高越好”还是“越低越好”。
区分输出速度与首段响应时间,避免把两个概念当成同一项。… See the full description on the dataset page: https://huggingface.co/datasets/keygate-ai/ai-model-evaluation-guide.fleurs-reducedbaseline-model-evaluationsWenetSpeech-Yue
WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation
Longhao Li1*, Zhao Guo1*, Hongjie Chen2,
Yuhang Dai1, Ziyu Zhang1, Hongfei Xue1,
Tianlun Zuo1, Chengyou Wang1, Shuiyuan Wang1,
Xin Xu3, Hui Bu3, Jie Li2, Jian Kang2,
Binbin Zhang4, Ruibin Yuan5, Ziya Zhou5,
Wei Xue5, Lei Xie1
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Institute of Artificial Intelligence (TeleAI)… See the full description on the dataset page: https://huggingface.co/datasets/modelevaluation/WenetSpeech-Yue.
