a · the target speaker extraction model once per session every block of conversation Enrollment target's voice, once Speech encoder ❄ → enrollment states Identity model ❄ → enrollment embedding Session cache enrollment states enrollment embedding Identity model ❄ same model as above short windows → embeddings window embeddings enrollment embedding Identity evidence tokens, time-aligned Conversation one microphone, 16 kHz raw audio Denoise Speech encoder ❄ same encoder as above 1 state / 80 ms Transformer blocks × N ❄ stock weights, unchanged → speaker activity A Transformer blocks × N 🔥 a copy of the blocks above, trained gated evidence fusion (see b below) K, V A Speaker Binding Module 🔥 attention over time × speakers which speaker is the target? target ∈ {spk 1 … K, none} score = A of that speaker enrollment states A h ℓ₂ ℓ₁ ⊕ learned blend σ p(target) per 80 ms frame ≥ τ → keep the words it overlaps ❄ frozen, pretrained 🔥 trained same colour = same network, shared weights dashed = read from the session cache A speaker activity (who speaks when, K unnamed speakers) h frame states ⊕ gated / blended sum σ sigmoid ◀ ▶ play ▶ show all stage 1
b · one transformer block with a fusion site h₀ block input Self-attention across frames FFN feed-forward h Q Cross-attention each frame reads nearby evidence Identity evidence tokens from panel a K V × tanh(g) gate, starts at 0 ⊕ h′ next block residual: h passes through unchanged ◀ ▶ play ▶ show all stage 1
c · the Speaker Binding Module: one decision per speaker, placed in time Frame-by-frame decisions are ambiguous on short turns, overlap and similar voices. The Speaker Binding Module (SBM) decides once per speaker for the whole block, like a diarize-then-verify system, and uses the diarization A to place that decision on frames. 1 · who are the speakers spk 1 spk 2 spk 3 spk 4 spk 5 v₁…v_K frames → identity vector enrollment → reference vector same adapter for both 2 · compare over time spk 1 spk 2 spk 3 spk 4 spk 5 time windows → each cell: how much this speaker talks here, and how enrollment-like it sounds 3 · attend along both axes time: consistently like the enrollment? speakers: which fits better? → one score per speaker 4 · ownership softmax over {spk 1 … K, none} → q spk 1 spk 2 spk 3 spk 4 spk 5 none ℓ₂[t] = Σ_k q_k · A[t,k] the target speaks where the owner speaks; 'none' silences an absent enrollment ℓ₂ ◀ ▶ play ▶ show all stage 1
The target speaker extraction model. The enrollment is encoded once per session and cached. Each block of conversation runs two paths through the same networks: a frozen copy gives anonymous speaker activity, a trained copy gives the per-frame target logit, and the Speaker Binding Module decides once per speaker which tracked voice, if any, is the wearer. For the product this is one probability per 80 ms frame that can be set to any operating point, with speaker-level identity inside the model rather than bolted on afterwards. Panels (b) and (c) open the fusion site and the binder.