CoolFace
6 results

fidelity-root

malaiwah /glm53-flash-fidelity-root-v1 fidelity--glm53flash.malaiwah.root.bf16 A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from zai-org/GLM-5.3-Flash-BF16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same cut… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-flash-fidelity-root-v1.tabularn<1K0 likes286 downloads19d agoHugging Facemalaiwah /glm53-fidelity-root-v1 fidelity--glm53.malaiwah.root.bf16 A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from zai-org/GLM-5.3-BF16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same cut as… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-fidelity-root-v1.0 likes169 downloads20d agoHugging Facemalaiwah /qwen38-27b-fidelity-root-v1 Qwen3.8-27B BF16 root fidelity dataset (hidden form) A root capture of Qwen/Qwen3.8-27B at revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 — 18 shards, no quantization_config, a genuinely unquantized reference — over the sealed suite-v5 shard-0 token panel (512 contexts x 2048 tokens = 1,048,064 scored positions). What this is for Quantization fidelity is usually reported as a KL divergence against a teacher. If the teacher was captured on a different stack than… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen38-27b-fidelity-root-v1.tabularn<1K0 likes135 downloads25d agoHugging Facemalaiwah /glm52-fidelity-root-v1 fidelity--glm52.malaiwah.root.bf16 A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from zai-org/GLM-5.2. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same cut as… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm52-fidelity-root-v1.tabularn<1K0 likes116 downloads19d agoHugging Facemalaiwah /qwen3-5-tiny-fidelity-root-v1 qwen3_5 random CPU fixture root A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/qwen3-5-tiny-random-bf16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same cut… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-5-tiny-fidelity-root-v1.tabularn<1K0 likes116 downloads17d agoHugging Facemalaiwah /deepseek-v4-tiny-fidelity-root-v1 deepseek-v4 random CPU fixture root A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/deepseek-v4-tiny-random-bf16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it).… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/deepseek-v4-tiny-fidelity-root-v1.tabularn<1K0 likes112 downloads17d agoHugging Face