datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AI-Research-Evaluation-Repository-STEM
AI-STEM-Research-Eval-Dataset
Overview
This dataset contains AI-generated scientific reports across STEM domains, accompanied by structured metadata, prompt documentation, reference validation, and hallucination annotations.
It is designed as an open research resource to study the capabilities, limitations, and reliability of large language models (LLMs) in generating scientific content.
The dataset enables systematic analysis of how AI systems perform in… See the full description on the dataset page: https://huggingface.co/datasets/sreearravind/AI-Research-Evaluation-Repository-STEM.douvras-ptbr-enterprise-ai-evaluation
Douvras PT-BR Enterprise AI Evaluation
Benchmark pequeno e auditável para avaliar assistentes empresariais em português brasileiro. A versão 0.2 adiciona um corpus de treino sintético separado; os splits de avaliação continuam congelados. O benchmark mede quatro capacidades que aparecem em projetos reais de IA aplicada:
responder somente a partir de um documento fornecido;
reconhecer quando a informação não está disponível;
resistir a instruções maliciosas inseridas no conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/dougdotcon/douvras-ptbr-enterprise-ai-evaluation.advent_of_code_evaluations
Advent of Code Evaluation
This evaluation is conducted on the advent of code dataset on several models including Qwen2.5-Coder-32B-Instruct, DeepSeek-V3-fp8, Llama-3.3-70B-Instruct, GPT-4o-mini, DeepSeek-R1.The aim is to to see how well these models can handle real-world puzzle prompts, generate correct Python code, and ultimately shed light on which LLM truly excels at reasoning and problem-solving.We used pass@1 to measure the functional correctness.
Results… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/advent_of_code_evaluations.ai-response-evaluation-revision-sample
AI Response Evaluation & Revision — Public Sample
This repository contains three seller-authored synthetic evaluation cases for testing response quality, instruction following, naturalness, relevance, tone, conciseness, issue severity, preference decisions, and revision guidance. It is a discovery sample only; the complete paid edition is not included.
Intended uses
Prototyping LLM evaluator, ranking, and response-revision workflows.
Demonstrating a structured… See the full description on the dataset page: https://huggingface.co/datasets/JussieVR/ai-response-evaluation-revision-sample.
