CoolFace
Datasetpublic

taehyeonkim/fp16-bf16-wfp16a16kvfp16-wbf16abf16kvbf16

FP16/BF16 W-FP16 / A-FP16 / KV-FP16 and W-BF16 / A-BF16 / KV-BF16 (Llama-3.1-8B-Instruct) This dataset contains FP16 and BF16 reference artifacts for Llama-3.1-8B-Instruct, stored in the same layout as the FP8 W-FP8/A-FP16/KV-FP8 artifact. w_of_wfp16a16kvfp16_llama_31_8b/ — FP16 weights The FP16 weights of Llama-3.1-8B-Instruct. Stored per layer: layer_0.safetensors ... layer_31.safetensors + embeddings.safetensors. The 7 linears per layer are stored as FP16… See the full description on the dataset page: https://huggingface.co/datasets/taehyeonkim/fp16-bf16-wfp16a16kvfp16-wbf16abf16kvbf16.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes17downloads
Dataset Card

FP16/BF16 W-FP16 / A-FP16 / KV-FP16 and W-BF16 / A-BF16 / KV-BF16 (Llama-3.1-8B-Instruct)

This dataset contains FP16 and BF16 reference artifacts for Llama-3.1-8B-Instruct, stored in the same layout as the FP8 W-FP8/A-FP16/KV-FP8 artifact.


w_of_wfp16a16kvfp16_llama_31_8b/ — FP16 weights

The FP16 weights of Llama-3.1-8B-Instruct.

Stored per layer: layer_0.safetensors ... layer_31.safetensors + embeddings.safetensors.

The 7 linears per layer are stored as FP16 tensors; layernorms, embeddings, final norm, and lm head are also FP16:

keydtypeshape
self_attn.q_proj.weight, self_attn.o_proj.weightfp16(4096, 4096)
self_attn.k_proj.weight, self_attn.v_proj.weightfp16(1024, 4096) - GQA (8 KV heads x 128)
mlp.gate_proj.weight, mlp.up_proj.weightfp16(14336, 4096)
mlp.down_proj.weightfp16(4096, 14336)
input_layernorm.weight, post_attention_layernorm.weightfp16(4096,)
model.embed_tokens.weight, lm_head.weightfp16(128256, 4096)
model.norm.weightfp16(4096,)

w_of_wbf16abf16kvbf16_llama_31_8b/ — BF16 weights

The BF16 weights of Llama-3.1-8B-Instruct.

Stored per layer: layer_0.safetensors ... layer_31.safetensors + embeddings.safetensors.

The 7 linears per layer are stored as BF16 tensors; layernorms, embeddings, final norm, and lm head are also BF16:

keydtypeshape
self_attn.q_proj.weight, self_attn.o_proj.weightbf16(4096, 4096)
self_attn.k_proj.weight, self_attn.v_proj.weightbf16(1024, 4096) - GQA (8 KV heads x 128)
mlp.gate_proj.weight, mlp.up_proj.weightbf16(14336, 4096)
mlp.down_proj.weightbf16(4096, 14336)
input_layernorm.weight, post_attention_layernorm.weightbf16(4096,)
model.embed_tokens.weight, lm_head.weightbf16(128256, 4096)
model.norm.weightbf16(4096,)

kv_fp16_of_wfp16a16kvfp16_llama_31_8b/ — FP16 KV cache

The KV cache of the FP16 model, stored directly as FP16 tensors.

Layout: <task>/sample_<n>.safetensors (all layers in one file). File metadata holds T (sequence length), dtype=fp16, storage_dtype=float16, and n_layers=32.

keydtypeshape (T = sequence length)
layer_<i>.k_codefp16(1, 8, T, 128)
layer_<i>.v_codefp16(1, 8, T, 128)

Tasks (20 samples each): RULER@4K - niah_multikey_1, ruler_vt, ruler_cwe, ruler_fwe, ruler_qa_squad; plus gsm8k_cot and longbench_hotpotqa.


kv_bf16_of_wbf16abf16kvbf16_llama_31_8b/ — BF16 KV cache

The KV cache of the BF16 model, stored directly as BF16 tensors.

Layout: <task>/sample_<n>.safetensors (all layers in one file). File metadata holds T (sequence length), dtype=bf16, storage_dtype=bfloat16, and n_layers=32.

keydtypeshape (T = sequence length)
layer_<i>.k_codebf16(1, 8, T, 128)
layer_<i>.v_codebf16(1, 8, T, 128)

Tasks (20 samples each): RULER@4K - niah_multikey_1, ruler_vt, ruler_cwe, ruler_fwe, ruler_qa_squad; plus gsm8k_cot and longbench_hotpotqa.