taehyeonkim/fp16-bf16-wfp16a16kvfp16-wbf16abf16kvbf16
FP16/BF16 W-FP16 / A-FP16 / KV-FP16 and W-BF16 / A-BF16 / KV-BF16 (Llama-3.1-8B-Instruct) This dataset contains FP16 and BF16 reference artifacts for Llama-3.1-8B-Instruct, stored in the same layout as the FP8 W-FP8/A-FP16/KV-FP8 artifact. w_of_wfp16a16kvfp16_llama_31_8b/ — FP16 weights The FP16 weights of Llama-3.1-8B-Instruct. Stored per layer: layer_0.safetensors ... layer_31.safetensors + embeddings.safetensors. The 7 linears per layer are stored as FP16… See the full description on the dataset page: https://huggingface.co/datasets/taehyeonkim/fp16-bf16-wfp16a16kvfp16-wbf16abf16kvbf16.
FP16/BF16 W-FP16 / A-FP16 / KV-FP16 and W-BF16 / A-BF16 / KV-BF16 (Llama-3.1-8B-Instruct)
This dataset contains FP16 and BF16 reference artifacts for Llama-3.1-8B-Instruct, stored in the same layout as the FP8 W-FP8/A-FP16/KV-FP8 artifact.
w_of_wfp16a16kvfp16_llama_31_8b/ — FP16 weights
The FP16 weights of Llama-3.1-8B-Instruct.
Stored per layer: layer_0.safetensors ... layer_31.safetensors + embeddings.safetensors.
The 7 linears per layer are stored as FP16 tensors; layernorms, embeddings, final norm, and lm head are also FP16:
w_of_wbf16abf16kvbf16_llama_31_8b/ — BF16 weights
The BF16 weights of Llama-3.1-8B-Instruct.
Stored per layer: layer_0.safetensors ... layer_31.safetensors + embeddings.safetensors.
The 7 linears per layer are stored as BF16 tensors; layernorms, embeddings, final norm, and lm head are also BF16:
kv_fp16_of_wfp16a16kvfp16_llama_31_8b/ — FP16 KV cache
The KV cache of the FP16 model, stored directly as FP16 tensors.
Layout: <task>/sample_<n>.safetensors (all layers in one file). File metadata holds T (sequence length), dtype=fp16, storage_dtype=float16, and n_layers=32.
Tasks (20 samples each): RULER@4K - niah_multikey_1, ruler_vt, ruler_cwe, ruler_fwe, ruler_qa_squad; plus gsm8k_cot and longbench_hotpotqa.
kv_bf16_of_wbf16abf16kvbf16_llama_31_8b/ — BF16 KV cache
The KV cache of the BF16 model, stored directly as BF16 tensors.
Layout: <task>/sample_<n>.safetensors (all layers in one file). File metadata holds T (sequence length), dtype=bf16, storage_dtype=bfloat16, and n_layers=32.
Tasks (20 samples each): RULER@4K - niah_multikey_1, ruler_vt, ruler_cwe, ruler_fwe, ruler_qa_squad; plus gsm8k_cot and longbench_hotpotqa.
