I want my Linux workstation to read an agent's answer aloud. Whisper already helps me transcribe my own speech; Kokoro is the other direction. I have been following its path through CKE one generated boundary at a time, from phoneme embeddings to duration, prosody and finally waveform generation.

There is now a new milestone: a fixed Kokoro request produces a 61,800-sample native waveform and an experimental 24 kHz mono WAV from a generated C circuit. The most important sentence is the next one: the fully connected waveform still fails its pinned-model numerical comparison. The maximum reported error is 0.007652. I can inspect and listen to a diagnostic artifact; I cannot call this a validated voice assistant yet.

What the generated program actually connects

The earlier embedding boundary proved that real model weights could flow through a declared circuit, normal lowering, memory planning, generated C, and independent PyTorch comparison. The ALBERT deep dive explained the contextual encoder. October's work continued through duration prediction, a shared checked frame extent, acoustic text features, style-conditioned decoding, source synthesis, spectral operations, and inverse STFT.

The latest generated-waveform PR (#722) joins those stages into a bounded fixed-input graph and writes a WAV through a standalone native C host. At 24,000 samples per second, the 61,800-sample result represents 2.575 seconds of audio. It is one pinned 36-phoneme request with the af_heart voice row, not arbitrary text sent through a complete text frontend. Captured Gaussian excitation remains an explicit diagnostic input rather than native request-local randomness.

Kokoro generated C stages: phoneme encoder, duration, acoustic features, source and spectral generation, experimental WAV. Individual stages have bounded evidence while complete connected waveform parity still fails.
The generated graph reaches WAV output, but connected numerical parity is a separate gate.

Why a WAV is not the finish line

The CKE Kokoro bring-up ledger records several distinct results. Some fixed-geometry encoder and decoder checkpoints meet declared numerical ceilings. The direct-input waveform tail also has a passing comparison. Yet small differences upstream can be amplified after source synthesis and the spectral path, where Kokoro consumes raw phase values. A phase branch-cut disagreement is not repaired by saying two angles are nearly equivalent: the next convolution sees their raw numbers.

Direct-input waveform tail passes its numerical comparison, but the connected phoneme-to-WAV path fails at a reported maximum error of 0.007652.
Component parity and connected-model parity are separate gates. F0 and raw phase are diagnostic boundaries, not a proven sole cause.

The current connected failure is useful because it names the debugging work instead of hiding it behind a playable file. The ledger points to generated F0 drift and source/STFT raw-phase behavior. The right next test replays identical inputs at those edges, isolates the first divergent output, fixes the operation or numerical contract, and reruns the entire graph. It should not weaken the full-waveform check merely because a local tail test passes.

Separate gates also remain for other phoneme lengths and voice rows, native excitation randomness, text preparation, repeatable playback, cancellation, human listening, and measured CPU speed. My accessibility workflow needs all of those, especially a reliable stop button and bounded request behavior. A single fixed WAV is evidence that the compiler can now compose much more of Kokoro; it is not evidence that I can safely rely on it with my eyes closed.

The compiler lesson

I do not want a separate hand-coded Kokoro runtime beside CKE. New math should become a tested kernel. The circuit should own operation order and shared edges. Kernel maps should declare the provider and its exact layout and numerical contract. Generated C should own the schedule and memory plan. When a connection cannot be expressed, the small reusable compiler extension is more valuable than a hidden model-specific workaround.

That is why this imperfect WAV matters to me. A real end-to-end attempt exercises the architectural seams that isolated kernel tests cannot. It gives the next engineer a concrete failed boundary, a pinned oracle, and a reason to improve CKE in a way the next audio model can reuse.

Primary evidence: PR #722, the generated-waveform test, and the Kokoro status ledger. Status reviewed at CKE afec1bfba on October 11, 2026. I have not independently rerun the full pinned-model test for this article; the numbers above are from merged test evidence.

Related Notes