Documentation

Last updated: 2026-09-20

The embedded model: MiniCPM-V 4.6 Thinking

Every Everlock build embeds MiniCPM-V 4.6 Thinking. It is baked into the binary — the same model on every platform, in every container image, with no runtime switch and nothing to download separately.

It drives both AI features:

What it ships as

FileContentsSize
MiniCPM-V-4_6-Thinking-Q6_K.gguflanguage half600 MB
mmproj-MiniCPM-V-4_6-Thinking-F16.ggufvision sidecar1109 MB

Most of the weight is in the vision half, which is unusual. The sidecar ships as F16 because there is no quantised release upstream.

Why the vision half is worth its size

A vision model normally resizes an image to one fixed square — often 448 pixels — and looks at it once. A 1444-pixel-wide screenshot has lost its text before the model sees it.

MiniCPM-V slices instead:

1444 x 1116 source
    |
    +-- overview            504 x 392  ->   63 tokens
    +-- 3 x 3 slice grid    504 x 392  ->   63 tokens each
                                           --------------
                                           630 image tokens

Ten views rather than one, with the grid shape taken from the image's aspect ratio, so a wide screenshot gets wide slices. Small text survives because it is never scaled past legibility. That is what makes the captioning path able to read text out of a photograph rather than describe it from a distance.

Why the Thinking variant

The Thinking release has the same architecture, the same size and the same vision sidecar as the base MiniCPM-V 4.6, and one difference in behaviour: it reasons in a hidden block before it answers. For captioning that is what keeps the answer in the language asked for. The base model, asked for a German description, answers in English on most photos and flips between the two when a single token of the prompt changes; the Thinking variant answers in German and also describes the photo more precisely, counting the objects in it rather than summarising them.

The runtime hides the reasoning and returns only the answer. Its cost is decode length: a description or a transcription spends 150 to 650 tokens thinking, a tag list up to 2000, roughly twice what the base model spends per photo. The captioning job runs in the background, so that shows up as throughput, not as latency anyone waits on. The /ai shell keeps a tighter reasoning budget so the operator prompt stays responsive.

Speed

The language half is 752M parameters and the decoder dominates this workload, so the vision slicing costs little: encoding a 4032x3024 photo takes a few seconds, and the rest of the time per photo is decode. Each caption language runs three prompts — description, text, tags — and each prompt decodes its reasoning plus its answer at the platform's token rate.

ai models deployment