datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
details_meta-llama__Llama-3.1-8B-Instruct_private
Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct.
The dataset is composed of 78 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 20 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_meta-llama__Llama-3.1-8B-Instruct_private.Taur_CoT_Analysis_Project___meta-llama__Meta-Llama-3.1-8B-Instructllama-3.1-8b-Instruct_sae_repsLlama-3.1-8B-Instruct-evals
Dataset Card for Llama-3.1-8B-Instruct Evaluation Result Details
This dataset contains the Meta evaluation result details for Llama-3.1-8B-Instruct. The dataset has been created from 30 evaluation tasks. These tasks are human_eval, gorilla_api_bench__huggingface, mmlu_pro, infinite_bench, api_bank, human_eval_plus, ifeval__loose, mmlu__0_shot__cot, nih__multi_needle, multilingual_mmlu_de, mmlu, gsm8k, mgsm, multilingual_mmlu_fr, multilingual_mmlu_pt, math_hard… See the full description on the dataset page: https://huggingface.co/datasets/meta-llama/Llama-3.1-8B-Instruct-evals.meta-llama_Llama-3.1-8B-Instruct-jdgfct-HarmlessnessMagpie-Llama-3.1-8B-Instruct-UnfilteredDataset generated using meta-llama/Llama-3.1-8B-Instruc with the MAGPIE codebase.
The filtered dataset can be found here: /HiTZ/Magpie-Llama-3.1-8B-Instruct-Filtered
System prompts used
General
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nCutting Knowledge Date: December 2023\nToday Date: 26 Jul 2024\n\n<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n
Code
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-8B-Instruct-Unfiltered.meta-llama_Llama-3.1-8B-Instruct-jdgfct-Readabilitymeta-llama_Llama-3.1-8B-Instruct-jdgfct-Completenesstweet_topic_Llama-3.1-8B-Instruct_vocab_2000_last20_newsgroups_Llama-3.1-8B-Instruct_vocab_2000_laststackoverflow_Llama-3.1-8B-Instruct_vocab_2000_lastdetails_meta-llama__Llama-3.1-8B-Instruct_v2
Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_meta-llama__Llama-3.1-8B-Instruct_v2.MMLU-Pro_Llama-3.1-8B-Instruct_gPRM_train
MMLU-Pro_Llama-3.1-8B-Instruct_gPRM_train
Llama-3.1-8B-Instruct-eval-logs-and-scoresLlama-3.1-Taiwan-8B-Instruct-eval-logs-and-scoresMMLU-Pro_Llama-3.1-8B-Instruct_gORM_train
MMLU-Pro_Llama-3.1-8B-Instruct_gORM_train
details_Dampfinchen__Llama-3.1-8B-Ultra-Instruct
Dataset Card for Evaluation run of Dampfinchen/Llama-3.1-8B-Ultra-Instruct
Dataset automatically created during the evaluation run of model Dampfinchen/Llama-3.1-8B-Ultra-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Dampfinchen__Llama-3.1-8B-Ultra-Instruct.details_gaverfraxz__Meta-Llama-3.1-8B-Instruct-HalfAbliterated-TIES
Dataset Card for Evaluation run of gaverfraxz/Meta-Llama-3.1-8B-Instruct-HalfAbliterated-TIES
Dataset automatically created during the evaluation run of model gaverfraxz/Meta-Llama-3.1-8B-Instruct-HalfAbliterated-TIES.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_gaverfraxz__Meta-Llama-3.1-8B-Instruct-HalfAbliterated-TIES.Llama_3.1-8B-Instruct-Self-CalibrationThe official repository which contains the code and pre-trained models/datasets for our paper Efficient Test-Time Scaling via Self-Calibration.
🔥 Updates
[2025-3-3]: We released our paper.
[2025-2-25]: We released our codes, models and datasets.
🏴 Overview
We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Llama_3.1-8B-Instruct-Self-Calibration.llama-3.1-8b-instruct-atlas
llama-3.1-8b-instruct-atlas
Llama-3.1-8B-Instruct-uPRM-T80-adapters-best_of_n-completionsGPQA_with_Llama_3.1_8B_Instruct_v1
GPQA with Llama-3.1-8B-Instruct
This dataset contains 646 graduate-level science questions from the GPQA benchmark with 100 candidate responses generated by Llama-3.1-8B-Instruct for each problem. Each response has been evaluated for correctness using a mixture of GPT-4o-mini and procedural Python code to robustly parse different answer formats, and scored by multiple reward models (scalar values) and LM judges (boolean verdicts).
Dataset Structure
Split: Single… See the full description on the dataset page: https://huggingface.co/datasets/hazyresearch/GPQA_with_Llama_3.1_8B_Instruct_v1.mmlu-pro-prep-eval-Llama-3.1-8B-Instruct-cotMATH500_with_Llama_3.1_8B_Instruct_v1
MATH-500 with Llama-3.1-8B-Instruct
This dataset contains 500 mathematical reasoning problems from the MATH benchmark with 100 candidate responses generated by Llama-3.1-8B-Instruct for each problem. Each response has been evaluated for correctness using a mixture of GPT-4o-mini and procedural Python code to robustly parse different answer formats, and scored by multiple reward models (scalar values) and LM judges (boolean verdicts).
Dataset Structure
Split: Single… See the full description on the dataset page: https://huggingface.co/datasets/hazyresearch/MATH500_with_Llama_3.1_8B_Instruct_v1.Llama-3.1-8B-Instruct-Infinity-Instruct-0625
Llama-3.1-8B-Instruct-Infinity-Instruct-0625
Dataset Description
This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside Llama-3.1-8B-Instruct as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with meta-llama/Llama-3.1-8B-Instruct at temperature=1.
For more details on the training… See the full description on the dataset page: https://huggingface.co/datasets/nebius/Llama-3.1-8B-Instruct-Infinity-Instruct-0625.MMLU_with_Llama_3.1_8B_Instruct_v1
MMLU with Llama-3.1-8B-Instruct
This dataset contains multiple-choice questions from the MMLU benchmark with 100 candidate responses generated by Llama-3.1-8B-Instruct for each problem. Each response has been evaluated for correctness using a mixture of GPT-4o-mini and procedural Python code to robustly parse different answer formats, and scored by multiple reward models (scalar values) and LM judges (boolean verdicts).
Dataset Structure
Split: Single split named "data"… See the full description on the dataset page: https://huggingface.co/datasets/hazyresearch/MMLU_with_Llama_3.1_8B_Instruct_v1.Magpie-Llama-3.1-8B-Instruct-FilteredDataset generated using meta-llama/Llama-3.1-8B-Instruct with the MAGPIE codebase.
The unfiltered dataset can be found here: /HiTZ/Magpie-Llama-3.1-8B-Instruct-Unfiltered
Filter criteria
min_repetition = 100
def test_no_repetition(text: str):
# Count the frequency of each word in the text
word_count = Counter(text.split())
# Check if any word appears more than min_repetition times
return all(count <= min_repetition for count in word_count.values())
def… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-8B-Instruct-Filtered.all-Llama-3.1-8B-InstructLlama-3.1-8B-Instruct-uPRM-T80-adapters-dvts-completionswildguardmix_Llama-3.1-8B-Instruct_4096toks
