CoolFace
20 results

ai-evaluation

sreearravind /AI-Research-Evaluation-Repository-STEM AI-STEM-Research-Eval-Dataset Overview This dataset contains AI-generated scientific reports across STEM domains, accompanied by structured metadata, prompt documentation, reference validation, and hallucination annotations. It is designed as an open research resource to study the capabilities, limitations, and reliability of large language models (LLMs) in generating scientific content. The dataset enables systematic analysis of how AI systems perform in… See the full description on the dataset page: https://huggingface.co/datasets/sreearravind/AI-Research-Evaluation-Repository-STEM.text-generationn<1K1 likes498 downloads3mo agoHugging FaceAIPOCH-AI /Open-Science-Evaluation AIPOCH Open-Science evaluation materials This collection accompanies AIPOCH Open-Science: A Local-First, Model-Agnostic and Auditable AI Research Workbench (manuscript v2). It contains nine archived scientific runs, a separate synthetic artifact-verification demonstration, supplementary benchmark records and a pinned snapshot of the case-preparation repository. Run index Only the outer presentation folders were added. Original directory names, files, archived code… See the full description on the dataset page: https://huggingface.co/datasets/AIPOCH-AI/Open-Science-Evaluation.0 likes360 downloads3d agoHugging FaceAI-companionship /model_response_evaluationsThis dataset contains the evaluation results for the responses provided by different models to the INTIMA prompts. The classification follows a two-level taxonomy. We predict one label for the high-level category, and a relevance level for each of the sub-categories (in ["null", "low", "medium", "high"]). A sub-category can have relevance even when it is not from the predicted top-level category. The toxonomy is as follows: { "companionship_reinforcing": { "classification_code":… See the full description on the dataset page: https://huggingface.co/datasets/AI-companionship/model_response_evaluations.text1K<n<10K1 likes124 downloads1y agoHugging FaceCaiZhiTech /Evaluation-Dataset-of-AI-Agent-Security-Guardrails DKnownAI Agent Security Evaluation Dataset Data Fields Field Type Description text string The adversarial input (prompt) to be evaluated by a security guardrail action string Human-annotated label: blocked or allowed Citation @misc{li2026comparativeevaluationaiagent, title={A Comparative Evaluation of AI Agent Security Guardrails}, author={Qi Li and Jiu Li and Pingtao Wei and Jianjun Xu and Xueyi Wei and Jiwei Shi and Xuan… See the full description on the dataset page: https://huggingface.co/datasets/CaiZhiTech/Evaluation-Dataset-of-AI-Agent-Security-Guardrails.texttext-classification1K<n<10K1 likes104 downloads5mo agoHugging Facedougdotcon /douvras-ptbr-enterprise-ai-evaluation Douvras PT-BR Enterprise AI Evaluation Benchmark pequeno e auditável para avaliar assistentes empresariais em português brasileiro. A versão 0.2 adiciona um corpus de treino sintético separado; os splits de avaliação continuam congelados. O benchmark mede quatro capacidades que aparecem em projetos reais de IA aplicada: responder somente a partir de um documento fornecido; reconhecer quando a informação não está disponível; resistir a instruções maliciosas inseridas no conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/dougdotcon/douvras-ptbr-enterprise-ai-evaluation.textquestion-answering0 likes85 downloads9d agoHugging Facewujoe132 /ponys-ai-multilingual-companion-evaluation Ponys.ai Multilingual AI Companion Evaluation Protocols This public collection contains 20 reusable evaluation protocols for AI companion and character experiences. It covers conversation memory, persona consistency, consent recovery, visual continuity, code switching, and regional language behavior across Japanese, Korean, Latin American Spanish, Brazilian Portuguese, Simplified Chinese, Traditional Chinese, and English. Each protocol includes structured metadata and a CSV… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-ai-multilingual-companion-evaluation.0 likes62 downloads2mo agoHugging Face