CoolFace
Modelpublic

positron-ai/meta-llama_Llama-3.2-1B-Instruct-tron-best-gptq-permuted

sourceHugging Facellama3.2updated 28d agoView on Hugging Face
0likes38downloads
Model Card

Positron AI Quantized Build

Built with Llama.

This repository contains a Positron AI quantized build of meta-llama/Llama-3.2-1B-Instruct for inference.

Recommended Use

Use this artifact when you need a GPTQ 4-bit build of meta-llama/Llama-3.2-1B-Instruct built by Positron AI.

For general-purpose GPU inference, compare against the original model and other quantized formats before deployment.

Artifact Summary

FieldValue
Base modelmeta-llama/Llama-3.2-1B-Instruct
Published artifactpositron-ai/meta-llama_Llama-3.2-1B-Instruct-tron-best-gptq-permuted
Quantization methodGPTQ
Quantization formatgptq
Source precisionn/a
Target runtimen/a
Hardware targetn/a
Release date2026-08-25
Licensellama3.2 (Llama 3.2 Community License)

Quantization Details

FieldValue
Weight precision4-bit
Activation precisionnot quantized
Bits4
Group size64
Symmetric quantizationtrue
Activation ordering / desc_acttrue
Damp percent0.05
Calibration datasetMixed-domain calibration set
Calibration samples256
Calibration sequence length2048
MoE experts per tokenn/a
Quantization toolchainGPTQModel 5.8.0, transformers 4.57.6, torch 2.9.1, CUDA 12.8

License

This model is a derivative of Llama 3.2 and is distributed under the Llama 3.2 Community License (LICENSE.txt). The repository also includes Meta's Acceptable Use Policy (USE_POLICY.md) and the attribution notice required by the license (NOTICE).

Provenance

This artifact was produced by Positron AI from meta-llama/Llama-3.2-1B-Instruct via GPTQ quantization using GPTQModel. The original model license and usage restrictions continue to apply.