datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llama-4-eval-logs-and-scores
Dataset Card for llama-4-eval-logs-and-scores
This repository contains the detailed evaluation results of Llama 4 models, tested using Twinkle Eval, a robust and efficient AI evaluation tool developed by Twinkle AI. Each entry includes per-question scores across multiple benchmark suites.
Dataset Details
Dataset Description
This dataset provides the complete evaluation logs and per-question scores of various Llama 4 models, including Scout and… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/llama-4-eval-logs-and-scores.Llama4_SFTLlama_4_Maverick_Distilled_5k
Llama_4_Maverick_Distilled – 5,000 Reasoning Traces
High-quality distilled dataset built to mirror the thinking and reasoning traces of Llama 4 Maverick class models. Created for training student LLMs to reproduce Llama_4_Maverick_Distilled style step-by-step reasoning.
Version: 1.0Date: May 24, 2026Source: Synthetic generation by Meta AI
Purpose
This dataset captures the explicit chain-of-thought pattern characteristic of Llama_4_Maverick_Distilled: state the… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Llama_4_Maverick_Distilled_5k.dataset_dpo_llama4llama_4_fsdptrain_dpo_llama4llama405B-generated-small
