a · the target speaker extraction modelonce per sessionevery block of conversationEnrollmenttarget's voice, onceSpeech encoder ❄→ enrollment statesIdentity model ❄→ enrollment embeddingSession cacheenrollment statesenrollment embeddingIdentity model ❄same model as aboveshort windows → embeddingswindow embeddingsenrollment embeddingIdentity evidencetokens, time-alignedConversationone microphone, 16 kHzraw audioDenoiseSpeech encoder ❄same encoder as above1 state / 80 msTransformer blocks × N ❄stock weights, unchanged→ speaker activity ATransformer blocks × N 🔥a copy of the blocks above, trainedgated evidence fusion (see b below)K, VASpeaker Binding Module 🔥attention over time × speakerswhich speaker is the target?target ∈ {spk 1 … K, none}score = A of that speakerenrollment statesAhℓ₂ℓ₁⊕learned blendσp(target) per 80 ms frame≥ τ → keep the words it overlaps❄ frozen, pretrained 🔥 trained same colour = same network, shared weights dashed = read from the session cacheA speaker activity (who speaks when, K unnamed speakers) h frame states ⊕ gated / blended sum σ sigmoid
stage 1

b · one transformer block with a fusion siteh₀block inputSelf-attentionacross framesFFNfeed-forwardhQCross-attentioneach frame readsnearby evidenceIdentity evidence tokensfrom panel aKV× tanh(g)gate, starts at 0⊕h′next blockresidual: h passes through unchanged
stage 1

c · the Speaker Binding Module: one decision per speaker, placed in timeFrame-by-frame decisions are ambiguous on short turns, overlap and similar voices.The Speaker Binding Module (SBM) decides once per speaker for the whole block, like a diarize-then-verify system,and uses the diarization A to place that decision on frames.1 · who are the speakersspk 1spk 2spk 3spk 4spk 5v₁…v_Kframes → identity vectorenrollment → reference vectorsame adapter for both2 · compare over timespk 1spk 2spk 3spk 4spk 5time windows →each cell: how muchthis speaker talks here, and howenrollment-like it sounds3 · attend along both axestime: consistently likethe enrollment?speakers: which fits better?→ one score per speaker4 · ownershipsoftmax over {spk 1 … K, none} → qspk 1spk 2spk 3spk 4spk 5noneℓ₂[t] = Σ_k q_k · A[t,k]the target speaks wherethe owner speaks; 'none' silencesan absent enrollmentℓ₂
stage 1

The target speaker extraction model. The enrollment is encoded once per session and cached. Each block of conversation runs two paths through the same networks: a frozen copy gives anonymous speaker activity, a trained copy gives the per-frame target logit, and the Speaker Binding Module decides once per speaker which tracked voice, if any, is the wearer. For the product this is one probability per 80 ms frame that can be set to any operating point, with speaker-level identity inside the model rather than bolted on afterwards. Panels (b) and (c) open the fusion site and the binder.