Documentation
The embedded model: MiniCPM-V 4.6 Thinking
Every Everlock build embeds MiniCPM-V 4.6 Thinking. It is baked into the binary — the same model on every platform, in every container image, with no runtime switch and nothing to download separately.
It drives both AI features:
- the operator prompt inside the admin SSH session
- image captioning for the image backend
What it ships as
| File | Contents | Size |
|---|---|---|
MiniCPM-V-4_6-Thinking-Q6_K.gguf | language half | 600 MB |
mmproj-MiniCPM-V-4_6-Thinking-F16.gguf | vision sidecar | 1109 MB |
Most of the weight is in the vision half, which is unusual. The sidecar ships as F16 because there is no quantised release upstream.
Why the vision half is worth its size
A vision model normally resizes an image to one fixed square — often 448 pixels — and looks at it once. A 1444-pixel-wide screenshot has lost its text before the model sees it.
MiniCPM-V slices instead:
1444 x 1116 source
|
+-- overview 504 x 392 -> 63 tokens
+-- 3 x 3 slice grid 504 x 392 -> 63 tokens each
--------------
630 image tokens
Ten views rather than one, with the grid shape taken from the image's aspect ratio, so a wide screenshot gets wide slices. Small text survives because it is never scaled past legibility. That is what makes the captioning path able to read text out of a photograph rather than describe it from a distance.
Why the Thinking variant
The Thinking release has the same architecture, the same size and the same vision sidecar as the base MiniCPM-V 4.6, and one difference in behaviour: it reasons in a hidden block before it answers. For captioning that is what keeps the answer in the language asked for. The base model, asked for a German description, answers in English on most photos and flips between the two when a single token of the prompt changes; the Thinking variant answers in German and also describes the photo more precisely, counting the objects in it rather than summarising them.
The runtime hides the reasoning and returns only the answer. Its cost is decode
length: a description or a transcription spends 150 to 650 tokens thinking, a
tag list up to 2000, roughly twice what the base model spends per photo. The
captioning job runs in the background, so that shows up as throughput, not as
latency anyone waits on. The /ai shell keeps a tighter reasoning budget so
the operator prompt stays responsive.
Speed
The language half is 752M parameters and the decoder dominates this workload, so the vision slicing costs little: encoding a 4032x3024 photo takes a few seconds, and the rest of the time per photo is decode. Each caption language runs three prompts — description, text, tags — and each prompt decodes its reasoning plus its answer at the platform's token rate.
Read next
- The gguf-runner inference engine — how the model is embedded and executed
- Deployment — installing a build
- AI runtime and access model