datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Magpie-Llama-3.1-8B-Instruct-UnfilteredDataset generated using meta-llama/Llama-3.1-8B-Instruc with the MAGPIE codebase.
The filtered dataset can be found here: /HiTZ/Magpie-Llama-3.1-8B-Instruct-Filtered
System prompts used
General
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nCutting Knowledge Date: December 2023\nToday Date: 26 Jul 2024\n\n<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n
Code
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-8B-Instruct-Unfiltered.tweet_topic_Llama-3.1-8B-Instruct_vocab_2000_last20_newsgroups_Llama-3.1-8B-Instruct_vocab_2000_laststackoverflow_Llama-3.1-8B-Instruct_vocab_2000_lastLlama-3.1-8B-Instruct-eval-logs-and-scoresLlama-3.1-Taiwan-8B-Instruct-eval-logs-and-scoresMMLU-Pro_Llama-3.1-8B-Instruct_gORM_train
MMLU-Pro_Llama-3.1-8B-Instruct_gORM_train
details_Dampfinchen__Llama-3.1-8B-Ultra-Instruct
Dataset Card for Evaluation run of Dampfinchen/Llama-3.1-8B-Ultra-Instruct
Dataset automatically created during the evaluation run of model Dampfinchen/Llama-3.1-8B-Ultra-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Dampfinchen__Llama-3.1-8B-Ultra-Instruct.details_gaverfraxz__Meta-Llama-3.1-8B-Instruct-HalfAbliterated-TIES
Dataset Card for Evaluation run of gaverfraxz/Meta-Llama-3.1-8B-Instruct-HalfAbliterated-TIES
Dataset automatically created during the evaluation run of model gaverfraxz/Meta-Llama-3.1-8B-Instruct-HalfAbliterated-TIES.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_gaverfraxz__Meta-Llama-3.1-8B-Instruct-HalfAbliterated-TIES.Llama_3.1-8B-Instruct-Self-CalibrationThe official repository which contains the code and pre-trained models/datasets for our paper Efficient Test-Time Scaling via Self-Calibration.
🔥 Updates
[2025-3-3]: We released our paper.
[2025-2-25]: We released our codes, models and datasets.
🏴 Overview
We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Llama_3.1-8B-Instruct-Self-Calibration.llama-3.1-8b-instruct-atlas
llama-3.1-8b-instruct-atlas
Llama-3.1-8B-Instruct-uPRM-T80-adapters-best_of_n-completionsmmlu-pro-prep-eval-Llama-3.1-8B-Instruct-cotMagpie-Llama-3.1-8B-Instruct-FilteredDataset generated using meta-llama/Llama-3.1-8B-Instruct with the MAGPIE codebase.
The unfiltered dataset can be found here: /HiTZ/Magpie-Llama-3.1-8B-Instruct-Unfiltered
Filter criteria
min_repetition = 100
def test_no_repetition(text: str):
# Count the frequency of each word in the text
word_count = Counter(text.split())
# Check if any word appears more than min_repetition times
return all(count <= min_repetition for count in word_count.values())
def… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-8B-Instruct-Filtered.Llama-3.1-8B-Instruct-uPRM-T80-adapters-dvts-completionsMMLU-Pro_with_Llama_3.1_8B_Instruct_v1
MMLU-Pro with Llama-3.1-8B-Instruct
This dataset contains 500 multiple-choice questions from the MMLU-Pro benchmark with 100 candidate responses generated by Llama-3.1-8B-Instruct for each problem. Each response has been evaluated for correctness using a mixture of GPT-4o-mini and procedural Python code to robustly parse different answer formats, and scored by multiple reward models (scalar values) and LM judges (boolean verdicts).
Dataset Structure
Split: Single… See the full description on the dataset page: https://huggingface.co/datasets/hazyresearch/MMLU-Pro_with_Llama_3.1_8B_Instruct_v1.meta-llama_Llama-3.1-8B-Instruct_cot_mmlu-pro_20241016_231257
Dataset Card for meta-llama_Llama-3.1-8B-Instruct_cot_mmlu-pro_20241016_231257
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dvilasuero/meta-llama_Llama-3.1-8B-Instruct_cot_mmlu-pro_20241016_231257/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/meta-llama_Llama-3.1-8B-Instruct_cot_mmlu-pro_20241016_231257.DeepAutoAI__d2nwg_Llama-3.1-8B-Instruct-v0.0-details
Dataset Card for Evaluation run of DeepAutoAI/d2nwg_Llama-3.1-8B-Instruct-v0.0
Dataset automatically created during the evaluation run of model DeepAutoAI/d2nwg_Llama-3.1-8B-Instruct-v0.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DeepAutoAI__d2nwg_Llama-3.1-8B-Instruct-v0.0-details.meta-llama__Llama-3.1-8B-Instruct-details
Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-3.1-8B-Instruct-details.Llama-3.1-8B-Instruct_eval_5554
mlfoundations-dev/Llama-3.1-8B-Instruct_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
4.7
15.8
43.2
44.7
14.1
25.8
13.1
2.1
6.7
17.0
0.3
0.3
8.9
AIME24
Average Accuracy: 4.67% ± 0.84%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
3.33%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Llama-3.1-8B-Instruct_eval_5554.entity_all_Llama-3.1-8B-InstructEpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-details.EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-0.001-128K-auto-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-0.001-128K-auto
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-0.001-128K-auto
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-0.001-128K-auto-details.EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-details.NbAiLab__nb-llama-3.1-8B-Instruct-details
Dataset Card for Evaluation run of NbAiLab/nb-llama-3.1-8B-Instruct
Dataset automatically created during the evaluation run of model NbAiLab/nb-llama-3.1-8B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NbAiLab__nb-llama-3.1-8B-Instruct-details.EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-details.EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-details.EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-details.EpistemeAI__Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy-details
Dataset Card for Evaluation run of EpistemeAI/Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy
Dataset automatically created during the evaluation run of model EpistemeAI/Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Polypsyche-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-Empathy-details.PersonaSignal-LeakageCheck-Communication-Formality-Meta-Llama-3.1-8B-Instruct-Turbo
