CoolFace
Modelpublic

IntellAgents/Nemotron-Labs-Diffusion-VLM-8B-TextOnly-GGUF

sourceHugging Faceotherupdated 4d agoView on Hugging Face
0likes112downloads
Model Card

Nemotron-Labs-Diffusion-VLM-8B-TextOnly-GGUF

Audited community GGUF conversion for research with AR, 32-token block diffusion, and linear self-speculation through the supplied runner. This is not an official NVIDIA or Ollama release. The non-AR modes require attention and cache control provided by the accompanying runner.

Text-only derivative of the VLM checkpoint. The language GGUF files are byte-identical to those in the full VLM release; no vision weights are included.

Non-commercial research/evaluation only under NVIDIA NSCLv1. This also applies to the text-only derivative. See the complete license and the unchanged source model card.

Files

FileSize (decimal GB)Storage
VLM8B-F16.gguf16.988F16

Excluded from this release: Q80, Q5KM, Q4K_M. Only sizes passing the declared release gates are included.

Choose one language GGUF. The full VLM additionally needs mmproj-F16.gguf for images. Tokenizer files and the chat template are included. F16 baselines are preserved; no importance matrix or further training was used.

Runtime support

RuntimeAR textBlock diffusionSelf-speculationImages
Supplied pinned runnerTestedTestedTested, sequential verifierNot included
Stock Ollama 0.33.3F16 import/raw/chat/multi-turn testedNot implementedNot implementedNot claimed

Ollama runs only the AR path. It cannot acquire diffusion or self-speculation from a Modelfile. Image input is not included in this release.

Run

Build the small bridge using runtime/README.md, with pinned llama.cpp b10760. From the downloaded repository directory:

bash
python runtime/generate.py --model VLM8B-F16.gguf --tokenizer . \
  --mode diffusion --prompt 'Explain how rain forms.' --max-new-tokens 128
python runtime/generate.py --model VLM8B-F16.gguf --tokenizer . \
  --mode self-speculation --prompt 'What is 2 + 2?' --max-new-tokens 32
ollama create nemotron-vlm8b-ar -f Modelfile.F16
curl http://localhost:11434/api/chat -d '{"model":"nemotron-vlm8b-ar","messages":[{"role":"user","content":"What is 2 + 2?"}],"think":false,"stream":false}'

The runner and tested Ollama server use GGML_CUDA_CUBLAS_COMPUTE_TYPE=f32. Set it in the server environment for Ollama; it is not a Modelfile parameter. The runner sets it by default. F16 weights/KV remain F16. Read the runtime documentation before overriding this measured precision requirement. Ollama requests must explicitly set think: false for the tested non-thinking behavior; the default thinking mode is not covered by this compatibility result.

Validation and limits

All original language tensors match after documented Q/K permutation and dtype conversion. The full VLM's vision tensors are audited separately. Tokenizer and positional metadata checks pass. FP16 language checks have a maximum relative L2 logit error of 0.000957 against the pinned NVIDIA FP16 eager reference under the recorded test environment.

FP16 generation across 8 selected text/image fixtures matches the NVIDIA reference for all three modes: true. This is a limited regression result, not a general cross-backend equality claim. Detailed outputs and all numerical checks are in reports.

The 973-token, 12-probe causal regression suite is author-created, with prose, code, mathematics, and multiple languages. It is not a standardized benchmark or a downstream accuracy estimate.

StorageMean KL vs F16NLL increaseTop-1 agreement
F160.000000+0.000000100.0000%

Quantized outputs may differ. Sequential self-speculation is checked against each quant's own AR output, not required to reproduce F16 text. The optional batched verifier has separate equality measurements in each validation report.

Important boundaries:

  • —Greedy decoding only. Self-speculation defaults to one-token AR verification, retaining exactly the verified prefix. No throughput improvement is claimed.
  • —--verification batched is an experiment; changing batch shape may change close floating-point token decisions.
  • —This linear verifier does not reproduce the upstream VLM quadratic SBD lattice, draft LoRA, or distribution-preserving stochastic speculation.
  • —Positional checks near 16384/32768 use synthetic consecutive short prefixes, not a full long-context benchmark. The example context is 4096.
  • —Image tests use simple synthetic fixtures, not an OCR/VQA benchmark.
  • —Thinking, tools, audio, video, streaming, and other GPUs/CPUs are not covered by these release claims. The Ollama template targets plain non-thinking chat.

Provenance

  • —Upstream: nvidia/Nemotron-Labs-Diffusion-VLM-8B@adca93d16471c1e07d594ae444d23e1876f6b365.
  • —llama.cpp: 0f3a71be15af836d277c9f918adfafb45732677e.
  • —Conversion/runtime source commit: e962ef052056fc51ba704dd5b65b6136236d048b.
  • —Hardware: NVIDIA A100-SXM4-40GB; full dependencies in reports/environment.json.
  • —Full file hashes: SHA256SUMS, manifest.json.

Model weights and upstream tokenizer assets retain NVIDIA's terms. The original IntellAgents runner/conversion additions use MIT; that does not relicense the model. Conversion changes tensor layout, storage type, metadata, and optional quantization; it performs no training.