September was the month C-Kernel-Engine started to feel less like a collection of model demonstrations and more like a system I could keep extending. That is not the same as saying it is finished. I spent much of the month finding places where a model appeared to work until a real recording, a real tool call, or an independent numerical comparison asked a harder question.
I wanted CKE to run models on CPUs through kernels, circuits, kernel maps, generated C, and evidence I could inspect. September added another question: can the same architecture keep working when the model is not just another decoder? We exercised it with Qwen3.8 Flash, Whisper, Parakeet, Cohere Transcribe, vision, a small training loop, and the early Kokoro text-to-speech graph. Serving became more than a port that returns tokens, helped substantially by an outside contributor.
This is a retrospective on work merged in September 2026. I reviewed CKE at commit 8ef9b28a6 on October 3, but I have kept October's follow-on work out of the September claims.
What moved during the month
The sequence matters more than a count of pull requests. Early in the month, the Qwen3.8 Flash and Whisper fixes exposed where a plausible output was not enough. Then native generated audio paths arrived for Parakeet and Cohere. Later we connected more training, vision, serving, and Kokoro boundaries. All of that was accompanied by more explicit tests for missing evidence and wrong artifacts.
The compiler had to meet a stranger model
Qwen3.8 Flash Next (#456) made the challenge concrete. Its circuit needed explicit hyper-connections, per-layer embeddings, sparse attention, and mixed expert contracts. We added circuit and kernel-map declarations and then chased numerical drift with X-Ray. A short Q4 trajectory could be checked; that did not certify BF16, long context, or every chat shape. I prefer that narrower sentence because a new model can expose a missing compiler capability without invalidating the whole architecture.
The important movement was not “CKE knows Qwen3.8 by name.” It was that more of the model's requirements were expressed as circuit edges, exact providers, layouts, and execution-plan rules that another model could reuse. The bring-up trail records both the repaired reductions and the comparisons still unavailable. That is how I want model support to mature: declare, generate, compare, localize, then retain the regression.
Audio stopped being just a kernel demo
Whisper became practical for my own video work, but a real recording also found a graph-wiring problem and an audio-tail problem. #455 hardened long-audio execution; #476 and #482 reduced repeat work across windows; #483 made worker teardown fail closed. This was useful software work because I could transcribe a recording, inspect the result, and catch a regression before using it to edit a video.
Parakeet and Cohere then tested a different question: could the generated model own more of the audio graph instead of Python reconstructing it around native kernels? For Parakeet, #530 generated the TDT state progression and a standalone WAV host. The 42:22 recording completed on Ryzen in 17 windows, and the transcript matched the retained CKE reference after trimming a trailing newline. That is strong execution-equivalence evidence, not an independently aligned human word-error-rate result. Cohere's separate frontend, encoder, decoder, and long-audio work followed. I am not combining those into a universal accuracy or speed ranking.
Kokoro was the other end of the audio loop: speech out, rather than words in. By September 30, the generated path had a bounded phoneme encoder, duration and alignment, text expansion, shared prosody LSTM, and connected F0/noise prosody on pinned input. The shared ALBERT encoder work in #602 and the connected prosody work through #634 are substantial. But at the September cutoff there was no complete acoustic decoder, waveform, CKE WAV, or speech session. October has moved some of those boundaries; that is a follow-up article, not a reason to rewrite September as if it already had TTS.
Training and vision made the evidence more demanding
I care about backpropagation because I do not want CKE to stop at inference. #527 restored a scoped v8 generated FP32 training workflow: forward and backward, native AdamW, checkpoint/resume, export to inference, and checks that the loaded runtime was actually the one we built. The two-layer numerical certification and four-layer English-text workflow are useful foundations. #563 added two-layer composition parity and a dtype ledger. This does not mean arbitrary large-model training, BF16 training parity, or every optimizer is done.
Vision brought a different discipline. A generated encoder can execute and still fail the task a user cares about. September's independent evidence and vision-evaluation hardening separated execution coverage from OCR field correctness and kept private evaluation details out of public reports. OCR is one vision task, not a blanket certificate for image understanding. The result is less impressive as a headline and much more useful when deciding what to trust.
A contributor helped turn serving into an actual boundary
CKE also started meeting the world outside its command-line runner. An outside contributor helped shape that serving layer. The arc runs from a native session behind HTTP/SSE in #370, through streaming and response tests in #451 and #473, to structured tool-call work in #495. That contribution mattered, alongside the compiler and server work around it.
The September #594 PR shows why agent serving is more than returning tokens. Qwen3.5's publisher template could render correctly while its imported tool delimiters became different BPE token IDs. Separately, a native worker could finish while an unconsumed HTTP stream kept the single-flight slot occupied, causing subsequent requests to get 429 Session busy. The fix bound special-token handling to imported metadata and the worker lease to request ownership. Its pinned 0.8B bundle matched 327/327 prompt token IDs in one tested mode and 325/325 in another; Responses and Chat Completions both completed a native read_file tool call and continuation. The Qwen Code harness still repeated tool calls until its limit, so I will not turn that into a claim that a coding-agent task was completed.
Three streams toward practical use
By the end of September I could see three streams that need to meet, not three separate product announcements. Vision asks whether a generated path can accept an image and produce a result whose task quality we can independently judge. TTS asks whether a circuit can carry text through the acoustic graph to actual samples a person can hear. Serving asks whether the right model, tokenizer, prompt, tools, and request lifetime are bound together when a real client calls. Each stream improves a different contract CKE will need before I can rely on it daily.
For vision, September's independent comparison work and evaluation hardening were less about claiming a new model and more about refusing to call execution a successful reading of an image. The next useful evidence is a clearly scoped task such as OCR field extraction or a particular image question, evaluated against retained independent answers. A passing OCR field test would still not certify arbitrary photo description.
For TTS, the September Kokoro circuits turned more of the phoneme encoder, duration, text expansion, and prosody branches into generated execution. The missing output was the one I would actually listen to: waveform samples. That is why the September status is connected features, not “CKE speaks.” The next gate is a generated acoustic decoder and waveform path, then text preparation, repeated speech requests, playback, and listening-quality evaluation. The current Kokoro runbook can be ahead of this September snapshot; I am keeping the dates separate.
For serving, #594's token-ID and worker-lease fixes made the first text-and-tool path more trustworthy. The next question is whether that path survives a complete user session: model identity, cancellation, context limits, tool result, repeat request, and an honest error when something is unsupported. The September server is single-flight, not a multi-node scheduler or a general multimodal gateway. A local model can still be useful through it, as long as the application treats that limit as real.
I want that loop for my own work: transcribe a recording, inspect an image or screen when needed, ask a local model to act through bounded tools, and hear the important result while I continue working on Linux. Whisper already helps with the first step. Vision, serving, and Kokoro need to earn their parts through the same generated-C and regression discipline. Using an external speech output or coordinator while CKE catches up is an honest integration choice, not a claim that CKE already does everything.
The less visible progress: refusing false passes
We tightened nightly and regression evidence while adding models. #467 made the nightly verdict follow the actual gate rather than merely publishing an artifact; #515 made idle-nightly scheduling capacity-aware; #518 kept kernel evidence distinct. That will not make regressions impossible. It makes it harder for an old report, a missing checkpoint, or a skipped test to masquerade as a current pass.
My September takeaway is simple: the models are teaching the architecture where its contracts are incomplete. Audio forced execution ownership and window policy into the open. Vision forced a separation between runtime success and task success. Training forced artifact identity and gradients to be checked. The serving work forced us to care about token IDs, request ownership, and real tool round trips. That is progress I can build on, even when the next model still finds a new edge.
Scope: September 2026 merged work. Sources are linked at each claim. Numerical and task results refer to their stated model, machine, artifact, and test; they are not general performance guarantees.
The work behind the month
172 linked PRsEvery PR-linked change that landed on CKE main during September 2026, grouped by its primary area of improvement. Expand a group to inspect the original PR and merge commit. The category is an editorial aid, not a model-quality verdict.
Serving and agent requestsPrompt tokens, tools, response lifetimes, and client-facing failures.21 PRs
- PR #636test(v8/serve): reject lifecycle probe false passescd67c1350c
- PR #635feat(v8/serve): declare Gemma4 publisher chat-only profilebd35d94720
- PR #633test(v8/serve): add artifact-bound HTTP lifecycle acceptance8af0834321
- PR #630fix(v8/serve): recover disconnected streams and label output limitsdb8e18f38e
- PR #624fix(v8/serve): acknowledge cancellation during generated prefille3b77c6204
- PR #622fix(v8/serve): refresh circuit-derived serving inventory5037769c99
- PR #615fix(server): render publisher token variables from bundle metadata4d9504de42
- PR #612feat(server): bind harness acceptance to loaded bundle identity0f0b6f8d4f
- PR #608test(v8/server): certify bounded Qwen Code tool taska2b220c145
- PR #606fix(v8/server): certify request recovery after native completion and failure603be4e18a
- PR #594fix(v8/server): preserve Qwen3.5 tool tokens and native request ownership191335ebad
- PR #592fix(server): preserve Qwen Code tool preambles and edit strings75d001f432
- PR #581fix(v8/serve): bind Jinja prompts and tool calls to converted bundlesffb3898459
- PR #582fix(server): route dated web_search tool type through 501 policye5ab29d12d
- PR #574fix(v8/server): complete native template and tool protocol migration14bf32cc15
- PR #514fix(v8/server): validate agent context budgetsac63d4ffae
- PR #513feat(v8/server): expose agent request timings2d3caccfcb
- PR #509feat(v8/server): certify Qwen Code with Qwen3.89230bb6ad4
- PR #495feat(v8/server): add structured tool call and refactor ck_server_v8.py515ac84b01
- PR #473test(v8/server): add nightly response e2e of server50855ea37a
- PR #451fix(v8/server): harden Responses API streaming75073d761f
Vision and OCRImage execution, independent comparisons, and task-level evidence.19 PRs
- PR #632fix(v8/vision): reserve decode capacity after image prefill59305d1fe1
- PR #629fix(v8/vision): verify resumed OCR scores from generated output66333c064a
- PR #626fix(v8/vision): separate execution coverage from task-quality evidence1fe38409fe
- PR #619fix(v8/vision): pin OCR encoder evidence and full-corpus selection070761ff7d
- PR #611fix(v8/nightly): supply vision image dependency and align tool probed6c6373aac
- PR #609fix(v8/vision): fail closed on incomplete encoder evidence659033b115
- PR #607test(v8/vision): retain independent encoder corpus evidencef08baa7e6e
- PR #605fix(v8/vision): align Qwen3-VL image preprocessing with oraclee6d1a14ab6
- PR #578fix(v8/vision): align encoder parity with deployed preprocessinge19857885e
- PR #573fix(v8/vision): select aligned Qwen3-VL GGUF position interpolation11da6465ad
- PR #571fix(v8/xray): keep missing vision captures out of numerical verdictsa9ee505deb
- PR #564fix(v8/vision): verify generated OCR execution evidence64942dd5d8
- PR #562fix(v8/xray): capture full vision attention extent703d47cb8c
- PR #561fix(v8/vision): bind BF16 MLP down to BF16 providerc523ca93e8
- PR #559fix(v8/codegen): support sequential multimodal decoder ABI01b43ca639
- PR #557fix(v8/qwen3vl): support zero-DeepStack vision encoders718e54f3ac
- PR #506docs(v8): record Qwen3-VL corpus parity75d66be2e8
- PR #505test(v8/vision): share multimodal parity certificationc45dc50b6d
- PR #475fix(v8): restore Qwen3-VL nightly parity evidenceaa6f78afc5
Text to speechKokoro encoder, duration, prosody, and generated feature boundaries.25 PRs
- PR #634feat(v8/tts): generate complete Kokoro F0 and noise curvescf93b2b65a
- PR #631feat(v8/tts): connect first Kokoro prosody residual blocks1f493f8dda
- PR #627feat(v8/tts): connect first Kokoro F0 and noise AdaIN stages13ad7eb1af
- PR #623feat(v8/tts): add checked channelwise AdaIN providerfc9c356722
- PR #621feat(tts): connect generated Kokoro shared prosody scan83a174b5b5
- PR #617feat(v8/tts): generate complete Kokoro acoustic text encoder788bcdcfe3
- PR #613feat(v8/tts): generate checked Kokoro text embedding prefixfe3af93dc7
- PR #610feat(v8/tts): connect generated duration features to both expansion streamsef6daa02f4
- PR #604feat(v8/tts): connect generated Kokoro duration predictor5c8c2dd62d
- PR #603refactor(v8/tts): canonical encoder authoring and stitching controls82e1eec85f
- PR #602feat(v8/tts): certify the complete shared-layer Kokoro encoderbe1f8d6ab8
- PR #601feat(v8/tts): generate the complete first Kokoro ALBERT layera4fbd8dc6c
- PR #593feat(v8/tts): connect Kokoro first attention context2b86e24c9e
- PR #587feat(v8/tts): generate Kokoro ALBERT input projection1f87885197
- PR #586feat(v8/tts): generate Kokoro embedding boundary from BUMP weightsb5f03325cf
- PR #583feat(v8/tts): certify checked phoneme embedding boundarybcb322f612
- PR #580feat(v8/tts): lower duration logits and both feature expansionsf45f0dc9d7
- PR #579feat(v8/tts): reduce checked duration logits to frame extents652f2cc100
- PR #575feat(v8/tts): export pinned Kokoro weights and voice to BUMPe7e741e9c5
- PR #572feat(v8/tts): add bounded adaptive LayerNorm kernel50cb9915b6
- PR #567docs(v8/tts): publish Kokoro operation readiness checklist15f88cece6
- PR #568feat(v8/tts): add checked bidirectional LSTM scan providera4272f6794
- PR #566feat(v8/tts): route bounded circuits through normal codegen6f94f4d466
- PR #560feat(v8/tts): checked runtime lengths through generated C4e201a3275
- PR #555Draft: add TTS kernel oracles and pinned Kokoro referencec91072582d
Speech recognition and audioWhisper, Parakeet, Cohere Transcribe, and long-audio execution.35 PRs
- PR #628docs(v8/audio): fix Parakeet guide linkb37f6e4fec
- PR #625docs(audio): distinguish generated transcription from reference sessions628c5bc84d
- PR #569docs(site): link concepts audio section to newer kernel deep dives7e331ac114
- PR #550perf(audio): schedule relative attention by query row818b527e6a
- PR #548test(audio): cover Conv2D SIMD boundariesc5b15e90f0
- PR #546perf(audio): vectorize grouped Conv2D width outputscb3b20966e
- PR #542perf(v8/audio): record generated encoder GEMM shapes19694d9d9b
- PR #541test(v8): cover audio FP32 projection shapes889b312b3d
- PR #540feat(v8): add native Cohere long-audio certification5e83617656
- PR #537feat(v8): generate Cohere Transcribe decoder7150aaeb2b
- PR #536feat(v8): support dynamic audio encoder memoryf564c8ca8c
- PR #534feat(v8): generate Cohere Transcribe encoderf12b2793e4
- PR #533feat(v8): generate Cohere Transcribe frontend9d37d12f7c
- PR #532docs(whisper): trace encoder memory through decoder cross-attentiond56115d8bb
- PR #530feat(v8): generate standalone Parakeet transcription1350710b91
- PR #529feat(v8): generate complete Parakeet encoder22fafd70c8
- PR #528feat(v8): generate Parakeet FastConformer blocksbc04db4030
- PR #526feat(v8): generate Parakeet subsampling componente9ed1bb65c
- PR #524feat(v8): generate Parakeet audio frontend91fca1ed97
- PR #519fix(audio): populate decode attention scratch serially4d40b55a3d
- PR #516docs(audio): add per-kernel teaching visuals9fd64826d8
- PR #511docs(audio): document Parakeet TDT native bring-updc9054d91b
- PR #510feat(v8/audio): harden Cohere Transcribe long audioc24f01fd2e
- PR #508feat(v8/audio): add native Cohere Transcribe short-audio path047d79a0f0
- PR #507feat(v8/audio): harden Parakeet long audio107fc07407
- PR #504feat(v8): run Parakeet TDT end to end909a2c164c
- PR #503feat(v8): inventory Parakeet TDT bring-upf30dcf9a0e
- PR #483fix(v8/whisper): fail closed on worker teardowne9fb5d2437
- PR #482perf(v8/whisper): retain workers across audio windowsf66d455148
- PR #478perf(v8/whisper): parallelize exact encoder GELU4a433162b5
- PR #480perf(v8/whisper): parallelize decode cross-attention heads78d88abea1
- PR #477fix(v8/whisper): harden frontend cache fallbacka1210ab0fe
- PR #476perf(v8/whisper): reuse long-audio frontend20d998d086
- PR #459test(whisper): record GEMM occupancy and timing samples28b321bb50
- PR #455fix(v8/audio): harden projection circuits and long-audio E2Ee87b90d0fc
Training and backpropagationGenerated forward/backward, optimizer workflow, and parity.8 PRs
- PR #565feat(v8): separate training authoring from provider inventory9e71c9205f
- PR #563feat(v8): certify two-layer training compositions and publish dtype ledger3f3d04ced1
- PR #558feat(v8): certify a frozen SVG training fixture1c6a58bea9
- PR #554feat(v8): lower authored training graphs into generated executionc7893c8ca5
- PR #549docs(v8/training): make notebook evidence inspectablef641ef1069
- PR #544feat(v8/training): add Python authoring and capability preflight3f9bebe2f4
- PR #535feat(v8): certify BPE training across circuit depthscebe93bcdf
- PR #527feat(v8): restore training as end-to-end certification6e8ebbb454
Model families and kernelsNew model circuits, providers, quantization, and numerical fixes.24 PRs
- PR #588fix(tokenizer): preserve complete Qwen tool prompts25b722b3ae
- PR #570chore(version): archive v6.6 under version/legacy/3307fe56a5
- PR #553fix(v8/xray): export compact KV with per-layer head geometry371baad0f1
- PR #552fix(v8/gemma4): align GeGLU and gridless native replayd670d7f347
- PR #551perf(kernels): reuse AVX-512 FP32 activation loads87b49720d9
- PR #543perf(v8/gemm): reuse AVX2 activations across outputs3795ce4a32
- PR #538feat(v8): certify standalone Cohere transcriptione84c8d9ca1
- PR #523fix(v8): honor per-layer Q geometry in prefill layouts954e653bec
- PR #522fix(v8): harden Gemma4 segmented shared-KV parity73552e22bc
- PR #500feat(v8): expose structured build diagnostics00aba0599d
- PR #498fix(v8): restore Laguna and Nemotron compilation5a185068de
- PR #491perf(v8/qwen): parallelize exact attention gating7d91c5e7fa
- PR #490perf(v8/qwen): parallelize recurrent Q/K normalization9aa31f75ba
- PR #486feat(v8): add Muse-Glimmer text support with exact BF16 paritya56f48d1e8
- PR #488perf(v8/qwen): parallelize recurrent SiLU rowsc1702bb561
- PR #487perf(v8/qwen): parallelize recurrent state preparation78fcf8594a
- PR #485perf(v8/gemma3): parallelize exact prefill normalizationf0e4b01a2f
- PR #484perf(v8/gemma3): balance regular attention prefillaf9ea2d26e
- PR #481perf(v8): parallelize exact SwiGLU rowscdb7ceb880
- PR #479perf(v8): parallelize large fp32 decode projections29c1a05b29
- PR #465fix(v8/cohere): preserve explicit NVFP4 providers1ac175aa8e
- PR #463fix(v8/gemma3): match llama numerical schedule89ba66516e
- PR #464fix(v8/qwen36): resolve routed MoE storage tuples7173a82d56
- PR #457fix(v8): validate dense Qwen projection storage contractse09357a5eb
Compiler and runtime contractsCircuit lowering, layouts, memory planning, and generated C.8 PRs
- PR #620feat(v8): bind checked extents to provider scalar arguments221957c173
- PR #590fix(v8): validate checked call storage against plannere05bbb3662
- PR #589feat(v8): bind checked ABI constants per circuit operation058bb7f8df
- PR #584Fix v8 ARM engine loading and file-backed BUMP weightsbac772bcdd
- PR #585feat(v8/codegen): emit checked fixed-size native entries5f75de91be
- PR #556fix(v8/gemma4): require circuit-bound RoPE factors5aa69c43d3
- PR #461fix(v8/dsl): validate circuit-bound MLA provider classes439ae2f3f7
- PR #456feat(v8/qwen38): bring up Flash Next circuit and numerical providers117f83a1b6
Regression, evidence, and docsNightlies, audits, documentation, and developer inspection.32 PRs
- PR #618fix(parity): certify tiled Q6 graph and advance llama.cpp oracleb987e01683
- PR #616fix(parity): require executable llama.cpp oracle evidence68a4043b66
- PR #600fix(ci): move versioned numerical oracles into standard nightly4ab672aa7f
- PR #599test(kernels): unify oracle verdicts in standard nightly reporting65522aa52f
- PR #547docs(site): record EPYC 9755 acquisition and lab budget40411eadd4
- PR #539test(v8): make provider selection ratchet monotonicc26423db08
- PR #525docs(guide): add model bring-up contributor tutorial7369532f04
- PR #520test(v8): schedule Gemma4 artifact parityae5978f8f4
- PR #518test(nightly): preserve distinct kernel evidence1846760f07
- PR #517docs(site): lead the landing page with a v8 get-started path3a05720c91
- PR #515test(v8): schedule capacity-aware idle nightliesad8177ffd8
- PR #512docs(site): add client-side search33d903f9d5
- PR #502feat(v8/visualizer): add explain-this-operation panel194b83226a
- PR #501test(v8): link real-manifest runtime fixturesc83bd2dc6f
- PR #499chore(v8/audit): inventory kernel allocation ownership93d3d1ff67
- PR #497docs(site): rebuild codegen and kernel catalog pages for v8ac8ff09869
- PR #496docs: use direct Hugging Face model referencescaa9cb90ab
- PR #493test(v8): refresh provider selection inventoryf7aedf6e4b
- PR #494test(v8): record Muse certification performance2fd78382d3
- PR #492fix(v8): restore Muse nightly contract gatesee857855d9
- PR #489test(docs): enforce HTML-first documentation policy24876fbcaa
- PR #474test(v8): publish current capability evidencef80e2dbb12
- PR #472docs(readme): add Flash Next and Gemma3 evidence, drop stale capability map29dd19137d
- PR #471docs(site): navigation and comprehension pass2ce0357a97
- PR #470test(v8): register cross-family capability evidence17988dbac3
- PR #469docs(site): add Gemma3 numerical-schedule case study9284499dc9
- PR #467fix(ci): enforce nightly verdict after publishing evidence88d3fd8468
- PR #468docs(site): refresh kernel-map scoreboard to current audit238a6332c8
- PR #466docs(site): explain quant recipes and composite circuits02e13784f4
- PR #460docs(site): cover NVFP4 and Qwen3.8 Flash kernel concepts14edb65a5f
- PR #462fix(v8/audit): report composite circuits in novelty metriccbc8296a29
- PR #458docs(v8): add linked dense and Flash Qwen quickstarts075464be8c
Source: first-parent CKE main history, selected by commit date. Titles are taken from merge messages; follow each PR for scope, tests, and limitations.