Documentation
One model instead of two — choosing a single vision model
Everlock embeds its language model in the binary.
Until now it embedded one of two.
The default build carries SmolVLM-256M, small enough to disappear into the executable. A qwen3
feature swaps it for Qwen3.5: the 2B at Q3 on macOS, the 4B at Q4 on Linux and Windows.
Two builds, and two different answers to the question "what happens when a photo contains text".
We measured both, added a third, and the third won.
The test
One screenshot. German text, 1444x1116, the kind of thing that lands in a photo library when someone captures a meeting note.
The task is the one the Immich sync path calls AssetOcrV1: read the text back.
| model | what came back |
|---|---|
| SmolVLM-256M | "3 · " |
| Qwen3.5-2B | Feurles Drift-Erken (Severeportal vs. NetBox), a timestamp rendered as 1808.07.24 – 1808.09.39, then a repetition loop |
| MiniCPM-V 4.6 | the page, line for line |
The Qwen3.5 run also stopped on a ä, in a lookback slice that treated byte offsets as character
offsets. That bug is fixed, and it is not why we are switching.
The middle column is why.
Where the gap comes from
A single 448-pixel view of a 1444-pixel-wide screenshot has already discarded the text.
SmolVLM and Qwen3.5 both look once.
MiniCPM-V 4.6 slices:
1444 x 1116 source
|
+-- overview 504 x 392 -> 63 tokens
+-- 3 x 3 slice grid 504 x 392 -> 63 tokens each
--------------
630 image tokens
Ten views instead of one. The grid comes from the image's aspect ratio, so a wide screenshot gets wide slices. Small text survives because it is never scaled past legibility.
The part we did not expect
More vision work should cost more time.
On a 4032x3024 photo, end to end:
| model | runtime |
|---|---|
| MiniCPM-V 4.6 | 18.7 s |
| Qwen3.5-2B | 33.7 s |
Ten views is more encoder work and still finishes sooner, because the language half is 752M parameters against 2B. For this workload the decoder dominates, not the vision tower.
Size
Measured on the files under test:
| model | vision sidecar | total | |
|---|---|---|---|
| MiniCPM-V 4.6 Q4_K_M | 529 MB | 1109 MB | 1.6 GB |
| Qwen3.5-2B Q4_K_M | 1281 MB | 668 MB | 1.9 GB |
The totals are close. The distribution is not.
MiniCPM keeps most of its weight in the vision sidecar, which ships as F16 with no quantised release upstream. That is the number we would most like to reduce, and the one that presses hardest against the single-binary rule.
The plan
One build.
SmolVLM and the qwen3 feature both go away, replaced by MiniCPM-V 4.6 on every target. One model
to download, one to embed, one set of behaviour to reason about.
Three things get measured before that happens.
Encoder time on Linux and Windows. The fast matrix path for F16 sidecars is macOS-only today. Everywhere else the same ten views take a slower route, and those are the targets carrying the larger model.
Decode cost per photo. This checkpoint transcribes accurately with its reasoning block enabled and narrates the image without it, so a library scan pays for those tokens on every asset.
The embedded runtime. Everything above ran through the command-line runner. The in-binary path shares the same engine and has never executed this model.
None of the three is a reason to expect a different answer. All three are reasons to check before deleting a build.