cstr/dbnet-ic15-GGUF
DBNet ResNet-18 ICDAR 2015 — GGUF
Text detection model for CrispEmbed. Detects text regions (words/lines) in document images, scene photos, and screenshots.
Architecture: DBNet (Differentiable Binarization Network) with ResNet-18 backbone, FPNC neck, and probability map head. All BatchNorm folded into Conv at export time.
Source: MMOCR `dbnet_resnet18_fpnc_1200e_icdar2015` (Apache 2.0). This is the ICDAR 2015 Incidental Scene Text / Challenge 4 detector benchmark, not the separate ICDAR2015-TextSR dataset. Reported benchmark scores are P=0.885, R=0.758, H=0.817.
Dataset attribution: ICDAR 2015 Incidental Scene Text and the MMOCR dataset metadata. The similarly named ICDAR2015-TextSR dataset is a different dataset with separate ODbL terms and was not used as the training source for this detector.
Model Variants
All variants detect the same text regions with scores within 0.005 of F32. Parity validated via per-pixel diff harness against PyTorch reference.
Recommended: Q80 (13 MB) or F16 (24 MB) for production fidelity. Q4K is a debug/experimental override only; its numerical parity does not change the recognition model's separate precision policy.
Usage
Pair with a TrOCR recognition model (cstr/trocr-small-printed-GGUF) for a complete OCR pipeline.
CLI
# Full OCR pipeline (detect + recognize)
crispembed --det dbnet-ic15-q8_0.gguf \
-m trocr-small-printed-q8_0.gguf \
--ocr document.png
# JSON output
crispembed --det dbnet-ic15-q8_0.gguf \
-m trocr-small-printed-q8_0.gguf \
--ocr document.png --jsonC API
#include "crispembed.h"
void *ctx = crispembed_ocr_init("dbnet-ic15-q8_0.gguf",
"trocr-small-printed-q8_0.gguf", 4);
int n;
const crispembed_ocr_result *r = crispembed_ocr(ctx, "image.png", &n);
for (int i = 0; i < n; i++)
printf("(%g,%g): %s\n", r[i].x, r[i].y, r[i].text);
crispembed_ocr_free(ctx);Architecture
Input image (resized, padded to 32x)
|
+-> ResNet-18 backbone (stem + 4 stages x 2 BasicBlocks)
| Stage 0: 64ch, stride 4 Stage 1: 128ch, stride 8
| Stage 2: 256ch, stride 16 Stage 3: 512ch, stride 32
|
+-> FPNC neck (FPN-Cat variant)
| 4x lateral 1x1 conv -> top-down upsample+add
| 4x smooth 3x3 conv (256->64) -> concat 4x64=256ch
|
+-> DBHead probability branch
3x3 conv (256->64) + ReLU
ConvTranspose2d (64->64, k=2, s=2) + ReLU
ConvTranspose2d (64->1, k=2, s=2) + sigmoidPost-processing: binarize at 0.3 -> connected components -> bbox extraction with unclip expansion (ratio 1.5). Output sorted in reading order.
12.2M parameters. All BatchNorm pre-folded into Conv/ConvTranspose weights.
Conversion
pip install gguf numpy torch mmengine
# Download MMOCR checkpoint
wget -q "https://download.openmmlab.com/mmocr/textdet/dbnet/dbnet_resnet18_fpnc_1200e_icdar2015/dbnet_resnet18_fpnc_1200e_icdar2015_20220825_221614-7c0e94f2.pth"
# Convert
python models/convert-dbnet-to-gguf.py \
--checkpoint dbnet_resnet18_fpnc_1200e_icdar2015_20220825_221614-7c0e94f2.pth \
--output dbnet-ic15-f32.gguf
# Quantize
crispembed-quantize dbnet-ic15-f32.gguf dbnet-ic15-q8_0.gguf q8_0
crispembed-quantize dbnet-ic15-f32.gguf dbnet-ic15-q4_k.gguf q4_kLicense and attribution
The converted model artifacts are released under Apache-2.0, with source attribution to MMOCR. Retain the dataset attribution above when redistributing this model or derived artifacts.
Provenance and EU AI Act Art. 53 note
- Upstream model: open-mmlab/mmocr.
- Upstream licence:
apache-2.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not. - What was done here: format conversion and/or quantisation only (GGUF). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
- Training data: documented — where it is documented at all — by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository.
- Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.
