CoolFace
Modelpublic

almax000/cellsentry-model

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes26downloads
Model Card

CellSentry Model — Multi-Task Spreadsheet AI

A fine-tuned 1.5B parameter model for spreadsheet intelligence tasks. Built on Qwen2.5-1.5B with LoRA, this model handles three distinct tasks through prompt routing:

  • Formula Audit — Verify or dismiss rule engine findings in Excel formulas
  • PII Detection — Identify sensitive data (SSN, phone, email, national IDs) in cell values
  • Data Extraction — Extract structured fields (invoice number, date, vendor, totals) from spreadsheets

Model Details

PropertyValue
Base modelQwen/Qwen2.5-1.5B
Fine-tuningLoRA (rank 16, alpha 32)
Training4000 iterations, batch_size=2, lr=3e-5, AdamW
Quantization4-bit, groupsize=32 (Q4K_M for GGUF)
Context length1024 tokens
LicenseMIT

Available Formats

FormatFileSizePlatform
GGUF (Q4KM)cellsentry-1.5b-v3-q4km.gguf~940 MBWindows (llama.cpp)
MLX (4-bit g32)cellsentry-1.5b-v3-4bit-g32/~920 MBmacOS (MLX)
Currently only the GGUF format is uploaded. MLX format coming soon.

Usage

This model is designed to be used with CellSentry, an open-source desktop app for spreadsheet auditing. The app downloads the model automatically on first launch.

Manual Download

bash
# Install Hugging Face CLI
pip install huggingface-hub

# Download GGUF model
huggingface-cli download almax000/cellsentry-model cellsentry-1.5b-v3-q4km.gguf --local-dir ./models

Prompt Format

The model uses Qwen2.5 chat template with task-specific system prompts:

Formula Audit:

<|im_start|>system
You are a spreadsheet formula auditor...<|im_end|>
<|im_start|>user
{rule engine finding + cell context}<|im_end|>
<|im_start|>assistant

PII Detection:

<|im_start|>system
You are a PII detection specialist...<|im_end|>
<|im_start|>user
{cell values to scan}<|im_end|>
<|im_start|>assistant

Data Extraction:

<|im_start|>system
You are a document data extractor...<|im_end|>
<|im_start|>user
{spreadsheet content + template}<|im_end|>
<|im_start|>assistant

Training

  • Method: LoRA fine-tuning with multi-task data
  • Data: Synthetic + real-world spreadsheet samples across all three tasks
  • Fusion: LoRA weights fused into base model, then quantized (dequantize → fuse → re-quantize with group_size=32)
  • Key lesson: groupsize=64 loses fine-tuning quality; groupsize=32 is the minimum viable floor for 1.5B models

Limitations

  • Optimized for structured spreadsheet content, not general text
  • 1024 token context — large spreadsheets need chunking
  • PII patterns trained primarily on US and Chinese formats
  • Extraction templates cover 5 document types (invoice, receipt, PO, expense, payroll)

Related