datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
details_meta-llama__Llama-3.1-8B-Instruct_private
Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct.
The dataset is composed of 78 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 20 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_meta-llama__Llama-3.1-8B-Instruct_private.Taur_CoT_Analysis_Project___meta-llama__Meta-Llama-3.1-8B-Instructxprmt-llama-3.1-8b-instruct-multijailllama-3.1-8b-Instruct_sae_repsLlama-3.1-8B-Instruct-evals
Dataset Card for Llama-3.1-8B-Instruct Evaluation Result Details
This dataset contains the Meta evaluation result details for Llama-3.1-8B-Instruct. The dataset has been created from 30 evaluation tasks. These tasks are human_eval, gorilla_api_bench__huggingface, mmlu_pro, infinite_bench, api_bank, human_eval_plus, ifeval__loose, mmlu__0_shot__cot, nih__multi_needle, multilingual_mmlu_de, mmlu, gsm8k, mgsm, multilingual_mmlu_fr, multilingual_mmlu_pt, math_hard… See the full description on the dataset page: https://huggingface.co/datasets/meta-llama/Llama-3.1-8B-Instruct-evals.Magpie-Llama-3.1-8B-Instruct-UnfilteredDataset generated using meta-llama/Llama-3.1-8B-Instruc with the MAGPIE codebase.
The filtered dataset can be found here: /HiTZ/Magpie-Llama-3.1-8B-Instruct-Filtered
System prompts used
General
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nCutting Knowledge Date: December 2023\nToday Date: 26 Jul 2024\n\n<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n
Code
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-8B-Instruct-Unfiltered.meta-llama_Llama-3.1-8B-Instruct-jdgfct-HarmlessnessMagpie-Llama-3.1-70B-Instruct-UnfilteredDataset generated using meta-llama/Llama-3.1-70B-Instruc with the MAGPIE codebase.
The filtered dataset can be found here: HiTZ/Magpie-Llama-3.1-70B-Instruct-Filtered
System prompts used
General
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nCutting Knowledge Date: December 2023\nToday Date: 26 Jul 2024\n\n<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n
Code
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-70B-Instruct-Unfiltered.demo-safety-overnight-Llama-3.1-8B-Instructmeta-llama_Llama-3.1-8B-Instruct-jdgfct-Completenessmeta-llama_Llama-3.1-8B-Instruct-jdgfct-ReadabilityLlama-3.1-8B-Instruct-resultsmeta-llama_Llama-3.1-70B-Instruct-jdgfct-Readabilitytweet_topic_Llama-3.1-8B-Instruct_vocab_2000_last20_newsgroups_Llama-3.1-8B-Instruct_vocab_2000_laststackoverflow_Llama-3.1-8B-Instruct_vocab_2000_lastall-Meta-Llama-3.1-70B-Instruct-AWQ-INT4MMLU-Pro_Llama-3.1-8B-Instruct_gPRM_train
MMLU-Pro_Llama-3.1-8B-Instruct_gPRM_train
details_meta-llama__Llama-3.1-8B-Instruct_v2
Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_meta-llama__Llama-3.1-8B-Instruct_v2.demo-safety-gate-features-Llama-3.1-8B-Instructmaes-llama-3.1-8b-instructLlama-3.1-8B-Instruct-eval-logs-and-scoresMMLU-Pro_Llama-3.1-8B-Instruct_gORM_train
MMLU-Pro_Llama-3.1-8B-Instruct_gORM_train
xprmt-llama-3.1-8b-instruct-multijail-judge-evalLlama-3.1-Taiwan-8B-Instruct-eval-logs-and-scoresGPQA_with_Llama_3.1_70B_Instruct_v1
GPQA with Llama-3.1-70B-Instruct
This dataset contains 646 graduate-level science questions from the GPQA benchmark with 100 candidate responses generated by Llama-3.1-70B-Instruct for each problem. Each response has been evaluated for correctness using a mixture of GPT-4o-mini and procedural Python code to robustly parse different answer formats, and scored by multiple reward models (scalar values) and LM judges (boolean verdicts).
Dataset Structure
Split: Single… See the full description on the dataset page: https://huggingface.co/datasets/hazyresearch/GPQA_with_Llama_3.1_70B_Instruct_v1.aime_solutions_llama_3.1_8B_instructdetails_Dampfinchen__Llama-3.1-8B-Ultra-Instruct
Dataset Card for Evaluation run of Dampfinchen/Llama-3.1-8B-Ultra-Instruct
Dataset automatically created during the evaluation run of model Dampfinchen/Llama-3.1-8B-Ultra-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Dampfinchen__Llama-3.1-8B-Ultra-Instruct.Llama_3.1-8B-Instruct-Self-CalibrationThe official repository which contains the code and pre-trained models/datasets for our paper Efficient Test-Time Scaling via Self-Calibration.
🔥 Updates
[2025-3-3]: We released our paper.
[2025-2-25]: We released our codes, models and datasets.
🏴 Overview
We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Llama_3.1-8B-Instruct-Self-Calibration.llama-3.1-8b-instruct-atlas
llama-3.1-8b-instruct-atlas
