AI evaluations
model_response_evaluationsThis dataset contains the evaluation results for the responses provided by different models to the INTIMA prompts.
The classification follows a two-level taxonomy.
We predict one label for the high-level category, and a relevance level for each of the sub-categories (in ["null", "low", "medium", "high"]).
A sub-category can have relevance even when it is not from the predicted top-level category.
The toxonomy is as follows:
{
"companionship_reinforcing": {
"classification_code":… See the full description on the dataset page: https://huggingface.co/datasets/AI-companionship/model_response_evaluations.advent_of_code_evaluations
Advent of Code Evaluation
This evaluation is conducted on the advent of code dataset on several models including Qwen2.5-Coder-32B-Instruct, DeepSeek-V3-fp8, Llama-3.3-70B-Instruct, GPT-4o-mini, DeepSeek-R1.The aim is to to see how well these models can handle real-world puzzle prompts, generate correct Python code, and ultimately shed light on which LLM truly excels at reasoning and problem-solving.We used pass@1 to measure the functional correctness.
Results… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/advent_of_code_evaluations.pythia.summary.evaluations
