In the last post I stopped at a [36,128] tensor: 36 phonemes, each mapped to a 128-number vector by three small lookup tables plus a normalization. That tensor is where the boring part ends and the actual model begins. Between that compact table lookup and a waveform there sits an ALBERT — the architecture Google's 2019 paper used to match BERT-large's quality while carrying fewer parameters — doing the job of turning flat per-phoneme vectors into contextual phoneme features: vectors that know what came before and what comes after. This post explains what ALBERT is, why that specific design is inside Kokoro, and exactly how far CKE's generated-C chain reaches through it. The broader map of the whole Kokoro graph is in the Kokoro overview; this is the zoom.

Conceptual illustration of unlabeled phoneme tiles entering a compact matrix and branching into richer contextual paths.
From a compact phoneme representation toward context-aware features. This editorial illustration is not a model trace or generated audio.

BERT, First: Both Sides of the Token

Kokoro's phoneme encoder is built from the same block as BERT, so let's start there. A transformer encoder is bidirectional: when it processes a token in the middle of a sequence, its attention can read tokens on both sides. A causal decoder — the kind that writes chatbot replies one token at a time — is strictly left-to-right; it can only read what has already been generated. Both use the same mechanism, attention, and that's worth unpacking because people often mis-hear it.

Attention is not a lookup table, and it is not a language generator. It is a weighted average: every token builds a query, every token offers a key and a value, the query scores against all the keys, the scores become weights (via softmax), and each token's new vector is the weighted sum of all the values it can see. Nothing gets predicted; information just gets mixed. An encoder runs that mixing with both sides visible, once per layer, repeatedly; a decoder runs it masked to the left, one position at a time.

Why would a speech model care which side is visible? Because phonemes are context-dependent. Here is a small, clearly illustrative example — not a Kokoro test vector, just arithmetic to build intuition:

Take the sequence a · β · o (a vowel, a voiced stop, a vowel). The embedding table hands o the same 128-number vector it would hand any o anywhere. But after a bidirectional attention layer, the o that follows β has been mixed with the voiced-stop context on its left — its vector now carries information that a bare table lookup couldn't. The same symbol, different surrounding context, different vector. That is the entire reason the encoder exists: the duration predictor and everything downstream need phoneme features that already encode the neighborhood, because the length a phoneme actually takes depends heavily on what it's next to.

Illustration of attention as a weighted average: the token o in the sequence a, beta, o queries against the three keys, softmax turns the scores into weights, and the new vector for o is the weighted sum of the three values, contrasted with a plain embedding lookup that would return the same vector regardless of context.
The same symbol, two different fates. The table fetches; attention mixes — weighted by what the context actually says.
side note "Contextual" does a lot of work in this post. In the rest of it, "contextual phoneme features" always means: per-phoneme vectors that have passed through attention layers and therefore depend on the whole pinned sequence, not just the phoneme's own row in a table.

What ALBERT Changed: Two Kinds of Sharing

ALBERT (Lan et al., 2019) keeps the bidirectional encoder but applies what the paper calls "two parameter-reduction techniques" — and both of them are sharing, in two different senses. Understanding the distinction is the whole ballgame for the kernel side.

1. Factorized embedding: share across the table

BERT's word embedding is a V × H table: vocabulary size times hidden size. ALBERT splits it into a smaller V × E table followed by an E × H projection, with E < H. Instead of the word table living at full width, it lives at a narrow width and gets stretched into the model's hidden dimension by one matrix multiply. The position and token-type tables get the same treatment, so all three stay at width E and the projection happens once, after the sum.

Why does E < H save anything? The table is where parameters scale with vocabulary and sequence length; the projection is fixed-size no matter how big the vocabulary gets. For the pinned Kokoro configuration — vocabulary V = 178 (the phoneme set), positional extent 512, two token types, E = 128, H = 768 — here is the arithmetic, worked out with the real numbers from the checkpoint's config.json:

Factorized vs unfactorized embedding tables (Kokoro's actual V, E, H)text
Factorized (ALBERT/Kokoro), weights only:
  word     178 x 128  =   22,784
  position 512 x 128  =   65,536
  type       2 x 128  =      256
  tables                     88,576
  projection 128 x 768 =   98,304
  TOTAL                      186,880

Unfactorized (BERT-style, tables at H), weights only:
  word     178 x 768  =  136,704
  position 512 x 768  =  393,216
  type       2 x 768  =    1,536
  TOTAL                      531,456

Savings: 531,456 - 186,880 = 344,576 parameters (~65%)

Note where the savings concentrate in these numbers: the position table. 512 positions at full width is 393,216 parameters; at width 128 it's 65,536. Here the word vocabulary is tiny (178 phonemes), so the positional table dominates the total — that split is specific to Kokoro's phoneme configuration, not a law of ALBERT in general: in a text model with a large vocabulary, the word table itself can be the bigger cost. Either way, the factorization targets all three tables, not just the word one.

side note Parameter savings are not latency savings. Sharing weights changes storage and the possibility of memory traffic. It does not change the fact that every one of the 12 layer executions still performs its full matrix multiplies. A shared 12-layer model stores roughly one layer's weights but still does twelve layers' work.

2. Cross-layer parameter sharing: share across depth

The second technique is stranger and more radical: the layers themselves share parameters. BERT stacks 12 distinct transformer layers, each with its own attention and feed-forward weights. ALBERT's default is to run 12 executions of one shared layer — same weights, fed through repeatedly, each execution seeing the layer's own output. The paper's framing: depth still happens; the parameters don't.

Side-by-side weight layout: BERT with a wide V-by-H embedding table and twelve separate per-layer weight boxes, versus ALBERT with a narrow V-by-E table, an E-by-H projection, and a single shared layer weight box that twelve execution arrows reference.
Two kinds of sharing. Solid arrows are data flow through the model; dashed arrows are repeated references to the same weight memory — twelve executions, one set of learned parameters.

Execution depth is not the number of distinct weight sets. A fully shared 12-execution model does the work of 12 layers and stores the weights of 1. (ALBERT also supports intermediate grouping — several distinct sets, each reused over a block of layers — but the fully-shared case is what the Kokoro configuration resolves to, as we'll see.)

Conceptual illustration of changing activations passing repeatedly through one shared physical module rather than twelve separate copies.
One set of weights can be used on successive layer executions. The illustration is a metaphor for parameter reuse, not a measured Kokoro output.

The third piece of the original paper is sentence order prediction: instead of BERT's next-sentence-prediction objective, ALBERT's pretraining loss asks whether sentence A comes before B or B before A, aimed at inter-sentence coherence. One line of separation matters here: that is a training objective from the 2019 paper. Kokoro's pinned checkpoint is a pretrained inference graph — nothing in the CKE runbook or the pinned source establishes that Kokoro itself was trained with that objective, so this post treats it as paper background, not Kokoro provenance.

Inside One Real ALBERT Layer

Now the layer itself, step by step, with the pinned Kokoro shapes. Let T be the number of phonemes in the utterance. (In the CKE fixture T = 36 — that is a property of this pinned test vector, a single short sentence, not a limit of the model; the positional table is sized for 512.)

  1. Three-table embedding + LayerNorm. Each phoneme ID indexes the word table (178 × 128); its position indexes the position table (512 × 128); a token-type ID indexes the type table (2 × 128). The three lookups sum, then a LayerNorm runs across the 128. Output: [T, 128]. This is exactly the boundary the previous post covered.
  2. Input projection. A single dense layer stretches the narrow embedding into the model's hidden dimension: [T, 128] → [T, 768]. This is the factorized-embedding projection from the last section — the one matrix multiply that turns width-128 into width-768.
  3. Q / K / V projections. Three dense layers each read the 768-dim vector and write a new 768-dim vector: the query, key, and value for attention. [T, 768] → [T, 768], three times.
  4. Multi-head bidirectional self-attention. Each of those three 768-dim vectors is split into 12 heads of 64 dimensions (768 / 12 = 64). "12 heads" means the attention runs 12 independent 64-dim sub-spaces in parallel — each head can latch onto a different kind of relationship (adjacent phonemes, repeated sounds, clause boundaries) — and the 12 results are concatenated back to 768. The mechanism per head: scores Q·Kᵀ / √64, mask, softmax, then weights · V. Because it's an encoder, every position scores against every position on both sides.
  5. Attention output projection + residual + norm. The concatenated head output (768) passes through a fourth dense layer, is added back to the layer's input (the residual), and LayerNorm'd. [T, 768] in, [T, 768] out.
  6. Feed-forward + second norm. A two-layer MLP widens then narrows: 768 → 2048 (the intermediate_size), a GELU activation, then 2048 → 768, another residual, another LayerNorm.

The embedding and input projection prepare a [T,768] activation; they are not part of the repeatedly shared layer. Each shared-layer execution takes [T,768] in and returns [T,768], using the same weights across the 12 executions. Kokoro then projects the final representation down to [T,512] (the model's hidden_dim) and transposes it for the rest of the TTS graph.

side note What the masks mean here. The pinned 36-token fixture is entirely valid — every position is a real phoneme — so the current CKE attention provider runs unmasked. A real utterance with padding (blank positions that shouldn't attend) needs a separate, explicit masking contract; it is not implied by the unmasked provider.
One ALBERT layer data flow: T-by-128 embeddings to T-by-768 input projection to twelve-head Q, K, V and bidirectional attention to attention output projection, residual and normalization to feed-forward, GELU, and second normalization. A contrasting boundary marker sits after the attention context, before the output projection, showing where CKE's verified generated chain currently stops.
One contextual layer, with the boundary marked. Green blocks are inside CKE's verified generated chain; the orange boundary sits after the attention context — everything past it in the layer (output projection, FFN, second norm) is declared but not yet connected.

Be careful with the framing on that last block: the attention context CKE has connected is a bidirectional attention computation. It is not, by itself, a complete ALBERT layer — it stops before the output projection, the feed-forward, and the second normalization, and it runs once, not through the repeated shared schedule. The rest of the layer and the layer repetition are the declared, still-open work.

Why Kokoro Uses It

Kokoro is a full TTS model, not just a phoneme encoder. The ALBERT sits at the front of a longer pipeline, and its output is the thing everything downstream is conditioned on. Here is the shape of the graph, kept deliberately compact:

  • Phoneme encoder (ALBERT, this post). Takes phoneme IDs, produces contextual [T, 512] features. This is the "speaker of the phoneme sentence."
  • Parallel text encoder. A small separate path (embedding → three weight-normalized convolutions → bidirectional LSTM) that captures sentence-level text context. It runs alongside the ALBERT, not after it.
  • Duration encoder + predictor. Reads the contextual features and predicts, for each phoneme, how many audio frames it should occupy. This is where the "o after β is longer than o after a" intuition becomes an actual number.
  • Prosody. Predicts fundamental frequency and noise/voicing contours from the expanded, duration-regulated features.
  • Voice conditioning. A fixed style vector (in the pinned fixture, from a voice-pack table) that tints the whole utterance toward a specific speaker.
  • Decoder + harmonic source + inverse STFT. Turns the prosodic plan into a 24 kHz waveform — the only stage that actually produces sound.

The division of labor is the point. The ALBERT is the understanding stage: it reads the phoneme sentence and produces features that already encode context. It never touches audio. The waveform is produced many stages later, by components the ALBERT never sees. So when we say "ALBERT is inside Kokoro," the precise claim is: the phoneme-encoder subgraph is an ALBERT. The model is not "an ALBERT that makes sound"; it's a TTS pipeline whose front end happens to be an ALBERT.

One more distinction that keeps recurring: phoneme IDs are not G2P. The ALBERT's input is a list of integer phoneme IDs. Getting from raw text to those IDs — grapheme-to-phoneme conversion — is a separate, upstream process (in Kokoro's pipeline, the Misaki G2P step). The ALBERT sees IDs; it does not spell, and it does not decide pronunciation. The pinned CKE fixture feeds it fixed IDs directly, which is one reason the boundary is so clean to test.

How CKE Represents It

This is where the explainer becomes a build system. CKE does not hand-write a per-model main.c for Kokoro. It describes the ALBERT phoneme encoder as a circuit — a declarative block graph — and the compiler turns that description into generated C. In the pinned fixture the block is called phoneme_encoder, and it reads, in order:

The connected Kokoro ALBERT circuit (ops and their kernel bindings)text
block phoneme_encoder
  header
    albert_embeddings        embedding_three_table_layer_norm_f32   -> [T,128]
  body
    albert_input_projection  linear_rows_checked_f32               -> [T,768]
    albert_query_projection  linear_rows_checked_f32               -> [T,768]
    albert_key_projection    linear_rows_checked_f32               -> [T,768]
    albert_value_projection  linear_rows_checked_f32               -> [T,768]
    albert_attention_context attention_full_token_major_f32_checked -> [T,12,64]

Three things about that representation are worth being precise about.

The circuit declares dependencies; the compiler derives memory. Each op names what it reads and what it writes. The lowering pass turns those declared dependencies into a memory plan: which buffer is live when, where each lands in a caller-owned arena, and when a buffer is dead and its space can be reused. Nothing is allocated at runtime — the generated C reads and writes at compiler-planned offsets. That's why the fixture's "thirteen effective BUMP tensors" is a property of the description, not of any runtime decision: the circuit says there are thirteen weight payloads, the compiler places them, and the generated code references the planned offsets.

Weights are bound by canonical name, which is what sharing hinges on. The embedding's three tables, the projection, and the three Q/K/V matrices are all weight references — named, SHA-256-verified BUMP payloads. A name is a pointer to one payload. So the mechanism that would let a repeated shared-layer schedule avoid duplicating weights is already present: if the circuit executed the layer block twelve times, each execution's ops could reference the same canonical weight names, and the lowering would bind all twelve to one payload at one arena offset. That is the design path to "twelve executions, one weight set." What is not yet established is that the connected fixture actually does it — the current circuit executes the shared layer once. The repeated schedule is the next declared boundary, and until it is connected and compared against a pinned checkpoint, the single-execution fixture is the proof. I'm describing the binding mechanism as capable, and the reuse as declared-but-unproven.

The kernels are model-neutral; the contracts are the Kokoro part. Look at the bindings again: embedding_three_table_layer_norm_f32, linear_rows_checked_f32, attention_full_token_major_f32_checked. None of them knows it's serving Kokoro. They are generic kernels with pinned numerical contracts (dtype, reduction order, thread partition, and for the "checked" variants, validate-before-write with an integer status). What makes the ALBERT specifically Kokoro is not new kernel code — it's the contracts and configuration: the exact layer-norm epsilon (1e-12), the specific GELU variant, the mask's additive semantics, the Q/K/V geometry at 12 heads of 64, and the 128→768 projection shape. The runbook is explicit that the decomposition belongs in declared operations and the native host, "not a kokoro branch in the [compiler]," and nothing runs in Python because CKE lacks it — every stage is either generated C or a declared kernel.

Reusable vs. genuinely generic gaps

  • Already reusable (exist, model-neutral): three-table embedding + LayerNorm, checked row-linear, full token-major attention, GELU, LayerNorm, the bidirectional LSTM scan.
  • Contracts to pin (declared, Kokoro-specific values): layer-norm epsilon, the GELU variant, additive mask semantics, 12×64 Q/K/V geometry, 128→768 projection, 768→512 output projection + transpose.
  • Genuine generic DSL gaps (not Kokoro-only): the repeated shared-layer schedule (execute one block N times against one weight set) and an explicit padded/masked-attention contract. These are compiler/IR features that would serve any shared or padded model, not Kokoro branches.

Evidence and the Open Edges

All CKE claims in this post are against commit 75d001f43, which I reviewed directly; the tree was clean and no Kokoro boundary had advanced since I last checked, so the status below is current at that commit rather than a copy of an older table. The connected generated chain runs through four merged boundaries, all inside the single phoneme_encoder circuit:

Connected generated-C boundaries at commit 75d001f43text
#586  three-table embedding + LayerNorm      [36,128]  worst abs err 2.384185791015625e-7
#587  128 -> 768 input projection             [36,768]  max abs err 9.5367431640625e-7
#591  Q / K / V prefix (12 heads, 13 BUMP)    [36,768]  max abs err 3.337860107421875e-6
#593  first unmasked bidirectional attention  [36,12,64] max abs err 7.748603820800781e-6
        (6 compiler-declared X-Ray checkpoints; all 36 pinned tokens valid)

It's worth keeping four different kinds of evidence separate, because they answer different questions:

  • Primitive tests compare one kernel in isolation against a PyTorch reference (e.g. the standalone attention primitive at 5.125999450683594e-6). They prove a kernel's math, not that it's wired into the model.
  • Connected X-Ray checkpoints run the generated circuit and compare at declared boundaries (the six checkpoints of the attention-context path). They prove the ops are connected and executing in the right order with the right buffers.
  • Optional skips are checks that only run when an asset is present — the exported-BUMP verification, for instance, reports SKIP when the weight bundle isn't mounted. A skip is not a pass.
  • Future work is declared in the runbook but not yet connected: the attention output projection, the residual normalization, the feed-forward, the GELU variant, the repeated shared-layer schedule, the final 768→512 projection, and everything downstream to the waveform.

So the honest status, at this commit: the first shared ALBERT layer is only partially connected — through the attention context. A complete shared layer, the repeated twelve-execution schedule, the full encoder output, and any waveform remain NOT_TESTED. Nothing in this post claims the encoder produces a waveform; the connected chain stops at a [36,12,64] attention context, and the distance from there to audio is the remaining declared work.

Related Notes

For the primary sources behind the CKE claims: the Kokoro TTS bring-up runbook records the pinned checkpoint (hexgrad/Kokoro-82M@e8a90b41), the connected boundaries, and the per-boundary test keys; the circuit itself lives at kokoro_embedding_projection_attention_generated_circuit.json; and the ALBERT paper is Lan et al., 2019, with the reference implementation at google-research/ALBERT.