forkjoin-ai/gemma-4-31b-it-gguf
Gemma4 31b IT (GGUF, Q4KM)
Production-ready GGUF quantization of google/gemma-4-31b-it for distributed text generation and conversation — powered by the Aether edge inference runtime on Edgework.ai.
Model Details
Usage
With llama.cpp
./llama-cli -m gemma4-31b-it.knot -p "Your prompt here" -n 256With Aether (Distributed Inference)
This model is deployed across the Aether distributed inference network. Weights are layer-sharded and distributed across multiple edge nodes for parallel inference.
Also available: .knot (sovereign format)
This repo ships `gemma4-31b-it.knot` — the model weights in the KNOT container that the Aether distributed-inference runtime loads natively (the GGUF, when present, sits right beside it). A KNOT is a single self-describing file with a JSON table-of-contents, so any single tensor is one HTTP `Range` request — ideal for streaming weights to edge nodes.
huggingface-cli download forkjoin-ai/gemma-4-31b-it-gguf gemma4-31b-it.knot --local-dir ./knotsFull format spec: KNOT_FORMAT.md. Inspect the header with bun run open-source/bitwise/scripts/dump-knot.ts gemma4-31b-it.knot.
Deployment Architecture
This model runs on the Aether distributed inference runtime — a custom engine that shards model layers across multiple nodes for parallel execution:
- Coordinator receives requests and manages token generation
- Layer nodes each hold a subset of model layers (4 nodes for this model)
- Hidden states flow between nodes via gRPC
- Zero cold start via warm pool scheduling
Deployed via Edgework.ai — bringing fast, cheap, and private inference as close to the user as possible.
About
Published by AFFECTIVELY · Managed by @buley
We quantize and publish production-ready models for distributed edge inference via the Aether runtime. Every release is tested for correctness and stability before publication.
