A vision model can return plausible text while taking the wrong path through its image prefix. That is the trap I want CKE's tests to catch. In October, the Gemma4 and multimodal work was less about adding a shiny new model name and more about making the first image-to-decoder boundary inspectable.
The September CKE recap described three streams: model architecture, modality bring-up, and stronger evidence. This is one concrete example of those streams meeting. A circuit must select the right operation, generated C must execute the right dataflow, and an independent reference must be able to tell us where the first wrong value appears.
Three places a believable answer can go wrong
The first is the image prefix itself. An image encoder or bridge can produce vectors with the expected shape but wrong numbers. PR #696 added an independent MTMD prefix comparison path, rather than treating CKE's own generated output as its oracle. PR #698 then tied prefix evidence to decoder provenance. This matters because a passing prefix from one artifact is not proof about a decoder from another.
The second is attention and prefill selection. Gemma4's segmented image-prefix path needed the correct flash-prefill provider and an explicit reduction contract; PR #685 changed the circuit, map and code generation together. The third is what happens when the decoder emits logits: PR #711 preserved output-projection row groups instead of silently treating the layout as one undifferentiated matrix.
What the merged fixes prove, and what they do not
These PRs show CKE can encode more of the vision path in its normal circuit and kernel-map system, and that the tests can investigate a mismatch at a named edge. They do not mean arbitrary images, OCR fields, chat formatting, or long multimodal conversations are certified. The current serving coverage ledger lists Gemma4's chat declaration, while runtime and exact-token serving evidence remain separately unassessed. A linked profile is a configuration fact, not a model-quality result.
The next validation I would want is an independently pinned image fixture with its preprocessing, prefix vectors, selected provider calls, decoder layer checkpoints, logits, and final token IDs all attached to one model artifact and commit. When any edge changes, the report should say which comparison failed. That is more useful than a single “vision works” label, especially as I start using these models for OCR, diagrams and eventually robotics perception.
Source boundary: CKE main afec1bfba, October 11, 2026. I reviewed the merged changes and generated coverage declaration; I did not run a new arbitrary-image or OCR quality evaluation for this article. The repository and the linked PRs are the implementation trail.
Related Notes
- What Changed in C-Kernel-Engine This September gives the earlier architecture and evidence context.
- How CKE Brought Up Qwen3.8 Flash Next shows why new model families often stress the circuit and compiler contract.
- How CKE X-Ray Found Qwen3.6's First Bad Circuit explains the first-divergence debugging method.
- Qwen3.8 Broke in CKE explains why a convincing demo cannot replace a retained regression fixture.