datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quantized-llama-3.1-leaderboard-v2-evals
Open LLM Leaderboard v2 Benchmark Results
This artifact contains all the data from evaluations of Neural Magic's quantized Llama-3.1 models.
These evaluations were produced with lm-evaluation-harness by running the following command:
lm_eval \
--model vllm \
--model_args pretrained="<model_path>",dtype=auto,add_bos_token=False,max_model_len=4096,tensor_parallel_size="<num_gpus>",gpu_memory_utilization=0.8,enable_chunked_prefill=True \
--apply_chat_template \… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-leaderboard-v2-evals.quantized-retrieval-dataToneWebinars_quantized-bigcodecquantized_burgers_vqMNLP_M3_quantized_datasetMNLP_M3_quantized_datasetquantized-llama-3.1-humaneval-evals
Coding Benchmark Results
The coding benchmark results were obtained with the EvalPlus library.
HumanEvalpass@1
HumanEval+pass@1
meta-llama_Meta-Llama-3.1-405B-Instruct
67.3
67.5
neuralmagic_Meta-Llama-3.1-405B-Instruct-W8A8-FP8
66.7
66.6
neuralmagic_Meta-Llama-3.1-405B-Instruct-W4A16
66.5
66.4
neuralmagic_Meta-Llama-3.1-405B-Instruct-W8A8-INT8
64.3
64.8
neuralmagic_Meta-Llama-3.1-70B-Instruct-W8A8-FP8
58.1
57.7
neuralmagic_Meta-Llama-3.1-70B-Instruct-W4A16
57.1… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-humaneval-evals.MNLP_M2_quantized_dataset1MNLP_M2_quantized_dataset
Dataset Card for MNLP_M2_sft_dataset
Dataset Description
A unified STEM instruction-following dataset comprising 240,500 examples drawn from six existing benchmarks: SciQ, Deepmind Code Contests, TIGER-Lab MathInstruct, TULU Algebra, TULU Code, and Facebook Natural Reasoning. Each example is formatted as a chat-style message pair for supervised fine-tuning of instruction-following models.
Curated by: Sarra Chabane
Shared by: GingerBled (https://huggingface.co/GingerBled)… See the full description on the dataset page: https://huggingface.co/datasets/arthurrpp/MNLP_M2_quantized_dataset.ToneBooksPlus_quantized-bigcodecemilia_en_quantized-wav-unifyquantized-llama-3.1-arena-hard-evals
Arena-Hard Benchmark Results
This artifact contains all the data neccessary to reproduce the results of the Arena-Hard benchmark for Neural Magic's quantized Llama-3.1 models.
The model_answers directory includes the generated answers from all models, and the model_judgements directory contains the evaluations by gpt-4-1106-preview.
The Arena-Hard version used for benchmarking is v0.1.0, corresponding to commit efc012e192b88024a5203f5a28ec8fc0342946df.
All model answers were… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-arena-hard-evals.MNLP_M2_quantized_dataset
MNLP M2 Quantized MCQA Dataset
Train split of prompt/completion examples with an extra dataset column indicating source.
Column
Type
Description
prompt
string
The input question prompt
completion
string
The ground-truth answer
dataset
string
Source label (e.g. scienceqa, M1_chatgpt, qasc, mathqa, commonsenseqa, openbookqa)
ToneSlavic_quantized-bigcodecquantized-common-voice-enMNLP_M2_quantized_datasetMNLP_M3_quantized_datasetMNLP_M2_quantized_datasetMNLP_M2_quantized_datasetMNLP_M2_quantized_datasetMNLP_M3_quantized_dataset
Dataset Card for MMLU STEM and Health
The MMLU dataset's validation split was filtered to include only subjects categorized as STEM and health, following the official subject taxonomy provided in the MMLU GitHub repository and as described in Hendrycks et al. (2021).
All selected STEM and health subject test sets were merged, randomly shuffled with a fixed seed, and assigned unique IDs to create a unified calibration set, as outlined in the original MMLU paper: "Measuring Massive… See the full description on the dataset page: https://huggingface.co/datasets/casimiir/MNLP_M3_quantized_dataset.MNLP_M3_quantized_datasetparler_tts_with_description_quantized-wav-uniMNLP_M2_quantized_dataset
MNLP_M2_quantized_dataset
This dataset is part of the CS-552 Modern NLP (Spring 2025) course project at EPFL. It contains a merged and cleaned collection of multiple-choice question-answer (MCQA) datasets curated for training generative reasoning models.
🧠 Motivation
The dataset was constructed to support training and evaluation of large language models (LLMs) on complex multi-step reasoning tasks, including those from:
Algebra and Arithmetic Word Problems
Natural… See the full description on the dataset page: https://huggingface.co/datasets/abdou-u/MNLP_M2_quantized_dataset.MNLP_M2_quantized_datasetMNLP_M3_quantized_datasetMNLP_M3_quantized_datasetgiant-midi-sustain-quantized
Dataset Card for "giant-midi-sustain-quantized"
More Information needed
emilia_multilang_quantized-wav-unimaestro-quantized
Dataset Card for "maestro-quantized"
More Information needed
