datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Qwen3.5-27B-AWQ-4bit-GPQA-Diamond-benchmarkBenchmark of cyankiwi/Qwen3.5-27B-AWQ-4bit against fingertap/GPQA-Diamond dataset.
Accuracy: 76.3% with Python tool.
Metric
Value
Correct
151
Incorrect
46
Errors
1
Total samples
198
Python tool calls
225
Total completion tokens
659,879
Raw stats:
{
"accuracy": 0.763,
"correct": 151,
"incorrect": 46,
"error": 1,
"total": 198,
"python_tool_calls": 225,
"completion_tokens": 659879
}
chatbot-arena-ja-karakuri-lm-8x7b-chat-v0.1-awqchatbot-arena-ja-calm2-7b-chatをフィルタリングし、karakuri-lm-8x7b-chat-v0.1-awqでchosenを生成しました
qwen3.5-moe-awq-calibration
Qwen3.5 MoE AWQ Calibration Dataset
Calibration dataset for AWQ (Activation-Aware Weight Quantization) of
Qwen/Qwen3.5-35B-A3B and
Qwen/Qwen3.5-35B-A3B-Base.
Designed for MoE expert routing diversity: Qwen3.5-35B-A3B has 256 experts with 8
active per token, so calibration data needs broad domain coverage to exercise as many
routing paths as possible.
Sampling methodology
Source: PleIAs/common_corpus
(open multi-domain corpus with labeled collections)
Filtering:
Token… See the full description on the dataset page: https://huggingface.co/datasets/Lambent/qwen3.5-moe-awq-calibration.awq-quant-evals
AWQ Quant Quality Evals
Side-by-side AWQ W4A16 vs BF16 quality measurements for:
Quant
Base
Suite
LostGentoo/Qwen3.5-4B-AWQ
Qwen/Qwen3.5-4B
OpenLLM-lite
LostGentoo/Qwen3-Embedding-8B-AWQ
Qwen/Qwen3-Embedding-8B
MTEB-lite
Hardware: NVIDIA RTX 5060 Ti (sm_120, Blackwell).
Files
File
Contents
quant_quality_evals.json
Full combined report
qwen35_4b_awq_vs_bf16.json
LLM OpenLLM-lite only
qwen3_embedding_8b_awq_vs_bf16.json
Embedding… See the full description on the dataset page: https://huggingface.co/datasets/LostGentoo/awq-quant-evals.Qwen3.6-27B-AWQ-BF16-INT4-SuperGPQA-benchmarkBenchmark of cyankiwi/Qwen3.6-27B-AWQ-BF16-INT4 against m-a-p/SuperGPQA dataset.
Accuracy: 69.2% with Python tool.
Metric
Value
Correct
692
Incorrect
295
Errors
13
Total samples
1000
Python tool calls
1508
Total completion tokens
3,806,045
Raw stats:
{
"accuracy": 0.692,
"correct": 692,
"incorrect": 295,
"error": 13,
"total": 1000,
"python_tool_calls": 1508,
"completion_tokens": 3806045
}
maldv__Awqward2.5-32B-Instruct-details
Dataset Card for Evaluation run of maldv/Awqward2.5-32B-Instruct
Dataset automatically created during the evaluation run of model maldv/Awqward2.5-32B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/maldv__Awqward2.5-32B-Instruct-details.Llama-3.2-1B-Instruct-AWQ-best_of_n-awq-completions-20250221hscode-eval-qwen3-4b-awq-reasoningevolved-math-problems-from-server-Qwen2.5-1.5B-Instruct-AWQhscode-eval-qwen3-4b-awq-noreasoningawq4-awq4-best_of_n-completions-20250226t_0-awq4-hf-best_of_n-completions-20250306Qwen2.5-72B-AWQ-Sudoku-Game-Testmixed-trainabs-qwen4b-sft5e-6-Qwen3-4B-AWQ-samp16-max100-validcuration-backup-review_semiclean31_awq_wermistral-7b-ncert-tutor-awq-benchmarksawq4-hf-best_of_n-completions-20250226hf-awq4-best_of_n-completions-20250226awq4_3b-hf-best_of_n-completions-20250304r1-qwen14b-awq-aimo-n32r1-qwen7b-awq-aimo-n32Qwen2.5-72B-AWQ-Maze-Game-Testarc-agi-mixed-barc-test-Qwen3-4B-AWQ-samp16-all-validyi-humanizer-v18-AWQ-test-100
