I want my Linux workstation to read an agent's answer aloud. Whisper already helps me transcribe my own speech; Kokoro is the other direction. I have been following its path through CKE one generated boundary at a time, from phoneme embeddings to duration, prosody and finally waveform generation.
There is now a new milestone: a fixed Kokoro request produces a 61,800-sample native waveform and an experimental 24 kHz mono WAV from a generated C circuit. The most important sentence is the next one: the fully connected waveform still fails its pinned-model numerical comparison. The maximum reported error is 0.007652. I can inspect and listen to a diagnostic artifact; I cannot call this a validated voice assistant yet.
What the generated program actually connects
The earlier embedding boundary proved that real model weights could flow through a declared circuit, normal lowering, memory planning, generated C, and independent PyTorch comparison. The ALBERT deep dive explained the contextual encoder. October's work continued through duration prediction, a shared checked frame extent, acoustic text features, style-conditioned decoding, source synthesis, spectral operations, and inverse STFT.
The latest generated-waveform PR (#722) joins those stages into a bounded fixed-input graph and writes a WAV through a standalone native C host. At 24,000 samples per second, the 61,800-sample result represents 2.575 seconds of audio. It is one pinned 36-phoneme request with the af_heart voice row, not arbitrary text sent through a complete text frontend. Captured Gaussian excitation remains an explicit diagnostic input rather than native request-local randomness.
Why a WAV is not the finish line
The CKE Kokoro bring-up ledger records several distinct results. Some fixed-geometry encoder and decoder checkpoints meet declared numerical ceilings. The direct-input waveform tail also has a passing comparison. Yet small differences upstream can be amplified after source synthesis and the spectral path, where Kokoro consumes raw phase values. A phase branch-cut disagreement is not repaired by saying two angles are nearly equivalent: the next convolution sees their raw numbers.
The current connected failure is useful because it names the debugging work instead of hiding it behind a playable file. The ledger points to generated F0 drift and source/STFT raw-phase behavior. The right next test replays identical inputs at those edges, isolates the first divergent output, fixes the operation or numerical contract, and reruns the entire graph. It should not weaken the full-waveform check merely because a local tail test passes.
Separate gates also remain for other phoneme lengths and voice rows, native excitation randomness, text preparation, repeatable playback, cancellation, human listening, and measured CPU speed. My accessibility workflow needs all of those, especially a reliable stop button and bounded request behavior. A single fixed WAV is evidence that the compiler can now compose much more of Kokoro; it is not evidence that I can safely rely on it with my eyes closed.
The compiler lesson
I do not want a separate hand-coded Kokoro runtime beside CKE. New math should become a tested kernel. The circuit should own operation order and shared edges. Kernel maps should declare the provider and its exact layout and numerical contract. Generated C should own the schedule and memory plan. When a connection cannot be expressed, the small reusable compiler extension is more valuable than a hidden model-specific workaround.
That is why this imperfect WAV matters to me. A real end-to-end attempt exercises the architectural seams that isolated kernel tests cannot. It gives the next engineer a concrete failed boundary, a pinned oracle, and a reason to improve CKE in a way the next audio model can reuse.
Primary evidence: PR #722, the generated-waveform test, and the Kokoro status ledger. Status reviewed at CKE afec1bfba on October 11, 2026. I have not independently rerun the full pinned-model test for this article; the numbers above are from merged test evidence.
Related Notes
- How Kokoro TTS Works, and Where CKE Stands maps the full graph and explains why speech synthesis is more than an encoder and decoder.
- Kokoro Is Reaching Generated C One Boundary at a Time documents the first pinned embedding proof.
- ALBERT: How Kokoro's Phoneme Encoder Works explains the shared-weight text encoder.
- I Want To Use My Linux Workstation Without Looking At It explains the accessibility goal behind this work.