almax000/cellsentry-model
026
CellSentry Model — Multi-Task Spreadsheet AI
A fine-tuned 1.5B parameter model for spreadsheet intelligence tasks. Built on Qwen2.5-1.5B with LoRA, this model handles three distinct tasks through prompt routing:
- Formula Audit — Verify or dismiss rule engine findings in Excel formulas
- PII Detection — Identify sensitive data (SSN, phone, email, national IDs) in cell values
- Data Extraction — Extract structured fields (invoice number, date, vendor, totals) from spreadsheets
Model Details
Available Formats
Currently only the GGUF format is uploaded. MLX format coming soon.
Usage
This model is designed to be used with CellSentry, an open-source desktop app for spreadsheet auditing. The app downloads the model automatically on first launch.
Manual Download
# Install Hugging Face CLI
pip install huggingface-hub
# Download GGUF model
huggingface-cli download almax000/cellsentry-model cellsentry-1.5b-v3-q4km.gguf --local-dir ./modelsPrompt Format
The model uses Qwen2.5 chat template with task-specific system prompts:
Formula Audit:
<|im_start|>system
You are a spreadsheet formula auditor...<|im_end|>
<|im_start|>user
{rule engine finding + cell context}<|im_end|>
<|im_start|>assistantPII Detection:
<|im_start|>system
You are a PII detection specialist...<|im_end|>
<|im_start|>user
{cell values to scan}<|im_end|>
<|im_start|>assistantData Extraction:
<|im_start|>system
You are a document data extractor...<|im_end|>
<|im_start|>user
{spreadsheet content + template}<|im_end|>
<|im_start|>assistantTraining
- Method: LoRA fine-tuning with multi-task data
- Data: Synthetic + real-world spreadsheet samples across all three tasks
- Fusion: LoRA weights fused into base model, then quantized (dequantize → fuse → re-quantize with group_size=32)
- Key lesson: groupsize=64 loses fine-tuning quality; groupsize=32 is the minimum viable floor for 1.5B models
Limitations
- Optimized for structured spreadsheet content, not general text
- 1024 token context — large spreadsheets need chunking
- PII patterns trained primarily on US and Chinese formats
- Extraction templates cover 5 document types (invoice, receipt, PO, expense, payroll)
Related
- CellSentry App — Desktop app that uses this model
- CellSentry Website — Project homepage
