When I say I want to run several models across my own machines, it is easy to jump straight to a diagram with six GPUs or a cluster of CPUs. The harder question is what one user request actually occupies: model weights, KV cache, recurrent state, a decode slot, and a place in the queue. Those are different resources. CKE's October serving work made that distinction concrete.
The CKE serving and batching page now separates three things that are often blurred together: multiple HTTP requests, batched model execution, and continuous batching. The first is a service feature. The second is an execution feature. The third is a scheduling policy. None implies the others automatically.
What changed in the generated path
The baseline server admits one request at a time to a loaded model, with a bounded FIFO waiting queue. That is useful because the server can reject excess work rather than accepting an unbounded backlog. But it does not make two requests share a kernel call.
CKE now has an experimental, opt-in two-slot decode path. The native loop and generated-code interface were developed across PR #697, PR #709, and PR #715. Two active requests can enter one ck_model_decode_batch2 step. Selected projection GEMMs run with two rows; attention and KV work still run per row with separate sequence state. The HTTP layer has its own admission and cancellation rules. A third request is rejected in this opt-in mode rather than being silently turned into a third slot.
That is a meaningful architecture milestone, not a claim that CKE now has general continuous batching. Two slots are fixed; the current path is for compatible KV-cache-only generated artifacts. It has not yet established sustained throughput, multi-model orchestration, or an N-request scheduler. The serving roadmap names the remaining work: fair admission, bounded prefill, broader active sets, and measurements that include latency as well as aggregate tokens per second.
Why a sizing calculator belongs beside the serving diagram
At work I have been asking what service envelope a fleet can honestly promise. A model's advertised maximum context is not the same thing as memory reserved by twenty live requests at that context. Nor is six devices one memory pool. Those questions apply to CPUs too: weights may be shared by requests on one node, but a second node needs its own resident copy.
CKE's new model sizing lab lets me enter a model's weight bytes, attention and recurrent-state assumptions, context, active users, device memory, and bandwidth. It can compare a requested total concurrency against selected replicas in a fleet. The page makes an important distinction between full-context KV owners, sliding-window KV owners, and recurrent layers. A DeltaNet-like layer should not be charged a full attention KV cache merely because it sits in a transformer-family model. At the same time, recurrent-state bytes are not magically zero; when the model metadata cannot determine them, the user must supply an assumption.
The calculation is deliberately a bound, not a benchmark. File size can estimate resident weights, but it does not reveal every runtime allocation or kernel layout. A bandwidth-only decode ceiling ignores prefill, interconnect traffic, scheduler overhead, and the fact that not every generated C operation is batched. The graph helps ask better questions; it does not certify a service-level agreement.
The experiment I want next
For one pinned model and one generated artifact, I want to measure the same prompt mix in single-request and two-slot modes. I would record per-user first-token latency, per-user decode speed, aggregate output tokens per second, peak memory, and what happens when a request cancels or a third one arrives. Then I would repeat with shorter and longer contexts. A two-row GEMM can improve weight reuse while another part of the model, or a joining prefill, still hurts latency. Both sides need to be visible.
This is the kind of boundary CKE is getting better at: a circuit declares the model, kernel maps bind executable providers, the generated program exposes the memory and decode schedule, and the service layer states its admission contract. The sizing page helps plan the experiment; the serving tests and future measurements determine whether the plan works.
Sources: initial sizing lab (#713), CKE-themed CPU and circuit sizing (#717), fleet concurrency comparison (#720), and the CKE repository. This post describes merged functionality as of afec1bfba on October 11, 2026; it does not report a new performance result.
Related Notes
- What Changed in C-Kernel-Engine This September explains the work that made this October serving step possible.
- Distributed CPU AI: MPI, RDMA, NUMA, and C-Kernel-Engine provides the broader multi-node context.
- How to Get Started With C-Kernel-Engine on Linux is the practical starting point for inspecting a generated model.
- How Ericsson Edge Routers Shaped CKE explains why I think in terms of control, execution, and bounded service contracts.