datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prompt-quality-vs-response-compliance
Prompt Quality vs Response Compliance — per-item results
Per-item numeric results from a study that measures prompt quality and response
compliance as separate constructs, rather than treating a model's response
score as a proxy for the person's prompting skill.
A four-agent evaluator on a locally hosted Qwen 2.5-7B-Instruct judge scores each
prompt as an artifact before any response exists, then scores the resulting
response twice on identical items: once against criteria… See the full description on the dataset page: https://huggingface.co/datasets/shahoismael/prompt-quality-vs-response-compliance.prompt-response-llmrouterbench
