datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
all-Meta-Llama-3.1-70B-Instruct-AWQ-INT4glm53-fixture-0.1B-fidelity-quant-int4-v1
GLM-5.3-Flash-0.1B fixture — candidate fidelity dataset, toy RTN-int4 routed experts (hidden form)
The numbers in this dataset are meaningless as quantization quality.
The weights are random (inference-optimization/GLM-5.3-Flash-0.1B-A0.1B is an
architectural fixture), and the quantizer is deliberately crude. This exists so
that step 3 of the three-step fidelity architecture has two real datasets to
compare, and so that anyone can see what a candidate capture looks like
next to… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-fixture-0.1B-fidelity-quant-int4-v1.Qwen3.6-27B-AWQ-BF16-INT4-SuperGPQA-benchmarkBenchmark of cyankiwi/Qwen3.6-27B-AWQ-BF16-INT4 against m-a-p/SuperGPQA dataset.
Accuracy: 69.2% with Python tool.
Metric
Value
Correct
692
Incorrect
295
Errors
13
Total samples
1000
Python tool calls
1508
Total completion tokens
3,806,045
Raw stats:
{
"accuracy": 0.692,
"correct": 692,
"incorrect": 295,
"error": 13,
"total": 1000,
"python_tool_calls": 1508,
"completion_tokens": 3806045
}
