Bhuvandesai/phi3-text-to-sql-gguf
026
Phi-3-mini Text-to-SQL — GGUF (quantized for CPU)
Quantized GGUF builds of the fine-tuned Phi-3-mini Text-to-SQL model (LoRA already merged into the base weights), for fast CPU inference with `llama.cpp`.
Note: "Q4" K-quants average ~5 effective bits/weight (embeddings and some tensors stay higher-precision), so the file is larger than a literal 4-bit×params calculation.
Which one?
Use `Q4_K_M`. On this task it matched Q5KM on quality while being smaller and faster.
Benchmarks (measured)
CPU = Intel i7-13650HX, 14 threads, llama-bench, build 9637:
Task quality (12 held-out questions, execution-match against a live SQLite DB):
4-bit quantization cost no measurable task accuracy vs 5-bit here.
Run it
# CLI
llama-cli -m phi3-text-to-sql-Q4_K_M.gguf -p "<|user|>\n<schema + question><|end|>\n<|assistant|>\n" -n 150 --temp 0
# Server (OpenAI-compatible)
llama-server -m phi3-text-to-sql-Q4_K_M.gguf -c 2048 -t 14 --port 8080The model expects Phi-3 chat formatting; include the database schema in the user turn (see the adapter card for the exact prompt). It outputs raw SQLite.
License: MIT.
