datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Llama-3.1-8B-Instruct-eval-logs-and-scoresLlama_3.1-8B-Instruct-Self-CalibrationThe official repository which contains the code and pre-trained models/datasets for our paper Efficient Test-Time Scaling via Self-Calibration.
🔥 Updates
[2025-3-3]: We released our paper.
[2025-2-25]: We released our codes, models and datasets.
🏴 Overview
We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Llama_3.1-8B-Instruct-Self-Calibration.llama-3.1-8b-instruct-atlas
llama-3.1-8b-instruct-atlas
Llama-3.1-8B-Instruct-uPRM-T80-adapters-best_of_n-completionsLlama-3.1-8B-Instruct-uPRM-T80-adapters-dvts-completionsLlama-3.1-8B-Instruct_eval_5554
mlfoundations-dev/Llama-3.1-8B-Instruct_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
4.7
15.8
43.2
44.7
14.1
25.8
13.1
2.1
6.7
17.0
0.3
0.3
8.9
AIME24
Average Accuracy: 4.67% ± 0.84%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
3.33%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Llama-3.1-8B-Instruct_eval_5554.llama-3.1-8b-funding-extraction-sft-ablations
LLaMA 3.1 8B Funding Extraction SFT Ablations
Ablation study results for LoRA SFT of Meta LLaMA 3.1 8B Instruct on structured funding metadata extraction from scholarly text.
The model extracts four fields: funder_name, award_ids, funding_scheme, and award_title.
Key findings
Factor
Best config
Avg F1
Overall best
synthetic, twostage (2+1 epochs), LoRA r=64, lr=3e-5
0.588
Data type
Synthetic >> non-synthetic (+0.126 avg F1)
—
LoRA rank
r=64 > r=32 > r=16
—… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/llama-3.1-8b-funding-extraction-sft-ablations.Llama3.1-8B-IT_TWISE_stepreward_dataeff_testLlama-3.1-8B-IT_testLlama-3.1-8B_neuron-activationsllama-3.1-8b-instruct_aime-all
meta-llama/Llama-3.1-8B-Instruct — aime-all
Model outputs from the micro-creativity inference suite.
Model: meta-llama/Llama-3.1-8B-Instruct
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/llama-3.1-8b-instruct_aime-all.Llama-3.1-8B-Instruct-em-en-finance-insecurealpaca_llama3.18b_em_paraphrased_divergence_ratiollama3.1-8b-it-qwq-sft-tpt-iter2-r64-24-v0.4-M
Dataset Card for "llama3.1-8b-it-qwq-sft-tpt-iter2-r64-24-v0.4-M"
More Information needed
llama3.1-8b-it-24-game-8k-qwq-r64-24-v0.4-n10_n10
Dataset Card for "llama3.1-8b-it-24-game-8k-qwq-r64-24-v0.4-n10_n10"
More Information needed
Llama-3.1-8B-Instruct-Qwen2.5-14B-Instruct-uPRM-T80-adapters-best_of_n-completionsLlama-3.1-8B-Instruct-resultsllama3.1-8b-it-countdown-game-7k-qwq-r64-countdown-v0.4-M_n10
Dataset Card for "llama3.1-8b-it-countdown-game-7k-qwq-r64-countdown-v0.4-M_n10"
More Information needed
Llama-3.1-8B-Instruct-BS16-RLHF-PRM-Math500Llama-3.1-8B-Instruct-best_of_n-completionsllama-3.1-8b-instantX600llama3.1_8b_base_studentevalLlama-3.1-8B-Instruct-best_of_n-prm-completionsLlama3.1-8b-instruct-customerservice-context-summarization-llm-judge-data
customer-service-context-summarization-evaluation-data
Lakshan2003/Llama3.1-8b-instruct-customerservice-context-summarization-llm-judge-data
Dataset updated with context summarization evaluation columns.
This README refresh triggers Hugging Face metadata re-index.
llama-3.1-8b-instruct_arena-hard-creative-writing
meta-llama/Llama-3.1-8B-Instruct — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: meta-llama/Llama-3.1-8B-Instruct
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/llama-3.1-8b-instruct_arena-hard-creative-writing.llama-3.1-8b-instruct_ifeval
meta-llama/Llama-3.1-8B-Instruct — ifeval
Model outputs from the micro-creativity inference suite.
Model: meta-llama/Llama-3.1-8B-Instruct
Dataset: ifeval (541 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/llama-3.1-8b-instruct_ifeval.llama-3.1-8b-instruct_eiffel_towerCreated for: https://github.com/shawonashraf/drrik
Llama-3.1-8B-Instruct_Filter2Honest_TrainTest_240Llama-3.1-8B-Instruct-resultsLlama3.1-8B-IT_test_offline_30k_rank
