datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quantized-llama-3.1-leaderboard-v2-evals
Open LLM Leaderboard v2 Benchmark Results
This artifact contains all the data from evaluations of Neural Magic's quantized Llama-3.1 models.
These evaluations were produced with lm-evaluation-harness by running the following command:
lm_eval \
--model vllm \
--model_args pretrained="<model_path>",dtype=auto,add_bos_token=False,max_model_len=4096,tensor_parallel_size="<num_gpus>",gpu_memory_utilization=0.8,enable_chunked_prefill=True \
--apply_chat_template \… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-leaderboard-v2-evals.musiccaps_quantized-wav-unifyMNLP_M3_quantized_dataset
Enhanced MCQA Test Dataset for Comprehensive Model Evaluation
This dataset contains 400 carefully selected test samples from MetaMathQA, AQuA-RAT, OpenBookQA, and SciQ datasets, designed for comprehensive MCQA (Multiple Choice Question Answering) model evaluation and quantization testing across multiple domains.
Dataset Overview
Total Samples: 400
MetaMathQA Samples: 100 (mathematical problems)
AQuA-RAT Samples: 100 (algebraic word problems)
OpenBookQA Samples: 100… See the full description on the dataset page: https://huggingface.co/datasets/AlirezaAbdollahpoor/MNLP_M3_quantized_dataset.MNLP_M2_quantized_dataset
MCQA Test Dataset for Model Evaluation
This dataset contains 3254 carefully selected test samples from MetaMathQA and AQuA-RAT datasets, designed for MCQA (Multiple Choice Question Answering) model evaluation and quantization testing.
Dataset Overview
Total Samples: 3254
MetaMathQA Samples: 3000 (mathematical problems)
AQuA-RAT Samples: 254 (algebraic word problems)
Question Types: Math, Algebra
Intended Use: Model evaluation, quantization benchmarking
Source… See the full description on the dataset page: https://huggingface.co/datasets/AlirezaAbdollahpoor/MNLP_M2_quantized_dataset.mozilla_ru_quantized-wav-unifyurban_flan_quantized-wav-uniCDB_results_quantized_mistralMNLP_M2_quantized_datasetcommon_voice_quantized-wav-unifyMNLP_M3_quantized_datasetbitandbytes_quantized
/hub_data4/seohyun/saves/ecva_instruct_1223/full/sft/checkpoint-350 · happy8825/valid_ecva_clean results
Model: /hub_data4/seohyun/saves/ecva_instruct_1223/full/sft/checkpoint-350
Dataset: happy8825/valid_ecva_clean
Generated: 2025-12-24 06:13:30Z
Metrics
Metric
Value
Total samples
924
With GT
0
Parsed answers
0
Top-1 accuracy
0
Recall@5
0
MRR
0
The uploaded JSON contains full per-sample predictions produced via t3_infer_with_vllm.bash.… See the full description on the dataset page: https://huggingface.co/datasets/happy8825/bitandbytes_quantized.sakhan10__quantized_open_llama_3b_v2-details
Dataset Card for Evaluation run of sakhan10/quantized_open_llama_3b_v2
Dataset automatically created during the evaluation run of model sakhan10/quantized_open_llama_3b_v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sakhan10__quantized_open_llama_3b_v2-details.Evaluation_QuantizedLoraanswerdotai-ModernBERT-baseEvaluation_QuantizedLoragoogle-bert-bert-large-uncasedfp8_quantized
/hub_data4/seohyun/saves/ecva_instruct_1223/full/sft/checkpoint-350-fp8 · happy8825/valid_ecva_clean results
Model: /hub_data4/seohyun/saves/ecva_instruct_1223/full/sft/checkpoint-350-fp8
Dataset: happy8825/valid_ecva_clean
Generated: 2025-12-24 05:49:28Z
Metrics
Metric
Value
Total samples
924
With GT
0
Parsed answers
0
Top-1 accuracy
0
Recall@5
0
MRR
0
The uploaded JSON contains full per-sample predictions produced via t3_infer_with_vllm.bash.… See the full description on the dataset page: https://huggingface.co/datasets/happy8825/fp8_quantized.MNLP_M2_quantized_dataset1Evaluation_QuantizedLorabert-large-uncasedEvaluation_QuantizedLoragoogle-bert-bert-base-uncasedEvaluation_QuantizedLorafacebook-bart-large-mnliEvaluation_QuantizedLoragoogle-rembertMNLP_M2_quantized_datasetEvaluation_QuantizedLoraanswerdotai-ModernBERT-largeEvaluation_QuantizedLoraFacebookAI-xlm-mlm-en-2048
