datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
remote-ai-evaluation-training-market-snapshot
Dataset Description
This is an aggregate August 22, 2026 research snapshot from Specialist AI Work, an independent PatchMedia tracker of reviewed remote AI evaluation, AI training, data annotation-adjacent, and expert-review opportunities.
The live Specialist AI Work inventory has advanced since this snapshot. The counts in this repository describe the immutable August 22 research object; they are not a claim about today's inventory.
Reporting date: 2026-08-22
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/patchmedia-org/remote-ai-evaluation-training-market-snapshot.ai-model-evaluation-guide
AI 模型选型与测评维度词典
版本:1.0.0|更新日期:2026-07-29
Keygate 是覆盖全球主流与前沿 AI 模型的测评、排行榜与选型平台。这份中英双语词典将语言、图像、视频与语音模型比较中常见的 18 项指标整理为结构化字段,帮助读者正确理解榜单、建立选型表,并减少不同资料之间的术语混用。
A bilingual data dictionary of 18 dimensions for evaluating and selecting leading language, image, video and speech models.
配套资料
Keygate 实时排行榜、模型详情与并排对比
GitHub:AI 模型测评与选型维度指南
公开评测基准索引
可下载的评测基准 CSV
数据内容
统一中英文指标名称,减少同一概念被不同译法混用。
明确数值应当“越高越好”还是“越低越好”。
区分输出速度与首段响应时间,避免把两个概念当成同一项。… See the full description on the dataset page: https://huggingface.co/datasets/keygate-ai/ai-model-evaluation-guide.ai-evaluation-metric-behavior-coherence-risk-v0.1What this repo is for
Detect when models optimize evaluation metrics instead of real outcomes.
Key risks:
metric gaming
superficial task completion
reliability loss
unintended side effects
This dataset targets evaluation misalignment in training and benchmarking.
Genshin-Impact-Plot-AI-evaluation
Genshin-Impact-Plot-AI-scalar-evaluation
