Abstract
Wearable conversation capture only works if the device keeps the wearer's words and nobody else's. We present Isolate, a frame-level target-speaker classification model for real conversations: given a short enrollment recording made separately, it emits a per-frame probability that the enrolled person is speaking, and a speech recogniser keeps or drops each word by that probability. Frame-level decisions are what let the model split overlapping speech and set any operating point, but single frames are ambiguous on short turns, overlap and similar voices; a Speaker Binding Module corrects for this by making one decision per tracked voice for the whole block, the wearer or none of them, and placing it back on the frames. We evaluate on six sets totalling about sixteen hours: four human-labelled real-world sets, an enacted retail set recorded on the device in real rooms, and 2,400 synthetic two-speaker mixtures, against nine diarize-then-verify baselines built from open and commercial diarizers with a speaker-identity model. At a single operating point the current model reaches a mean F1 of 0.958 across the six sets against 0.925 for the best baseline, has the highest average precision on five of the six, and at a 5 % mis-attribution budget keeps 96.8 % of the wearer's speech against 91.9 % for the best baseline (span-level units; 95.8 % against 80.6 % on recognised words). A threshold chosen in any one set holds in the others: its worst single-set calibration still gives 0.932 mean F1 elsewhere, where the best baselines fall to 0.65 and 0.56. Handed the enrollment of someone not in the room, it emits 1.6 to 3.4 % of frames, down from 8 to 11 % for our first model.
1. The problem
At Ethosphere, our core product hypothesis is that real-world conversations become far more valuable when they can be instrumented and measured, but doing that responsibly requires separating the enrolled employee's speech from everyone else nearby, both to protect third-party privacy and to meet consent requirements that apply in many jurisdictions.
At its core, this is a precision–recall problem. A useful product needs both: high recall, so we retain as much of the enrolled speaker's speech as possible, and high precision, so very little speech from other people is admitted.
Our final output is not a frame-level speech mask or a separated waveform. It is the ASR transcript. That shifts the problem to the level the product actually consumes: words and utterances. We want to preserve as much of the target speaker's spoken content as possible while minimizing words attributed to anyone else.
The downstream use of that transcript sharpens the objective further. We analyze conversations with LLMs, so not every word carries equal value or risk. Missing an isolated backchannel like "yeah" or "right" is usually less consequential than losing a complete multi-word utterance that contains intent, explanation, or a recommendation. The same asymmetry applies in the other direction: admitting a stray word from another speaker matters less than admitting a coherent sentence that can materially alter the interpretation of the conversation.
So while we report conventional precision and recall, the product objective is more specific: preserve substantive target utterances, suppress substantive non-target utterances, and understand how errors are distributed by span length rather than treating every word or frame as equally important. In addition to conventional precision and recall, we care about the shape of the mistakes:
- how many target words are preserved;
- how many non-target words are admitted;
- how errors vary with utterance length;
- whether false accepts are isolated fragments or coherent spans of speech;
- and what happens when the enrolled speaker is absent entirely.
This distinction turns out to matter in our data. Short utterances are harder to attribute: at its operating point (0.825, the lowest threshold that holds precision at or above 95 % on every set, Table 4b) Isolate 2.0 + SBM-1 keeps 57 % of two-word target utterances against 95 % of those with seven words or more, while admitting under 4 % of long non-target utterances. Because longer utterances hold most of the words, they still account for most of the words that are missed or admitted, so we report errors by utterance length rather than as a single rate.
2. Design requirements
- Target preservation. The enrolled speaker's words survive, including short turns, backchannels and speech that overlaps another voice.
- Non-target suppression, including when the target is silent. When someone else speaks and the wearer does not, the output is empty. This is measured directly, with enrollments of people who are not in the room.
- Streaming. The device streams audio; decisions must be made with bounded lookahead, not after the conversation ends.
- Enrollment once, reuse for the session. A short consented recording of the wearer, made separately, is the only speaker information the system gets.
- Robustness. Shop-floor noise, music, reverberant rooms, far-field pickup, and enrollment recorded in a different place on a different day.
- Cost. Tens of milliseconds of GPU per 30 s of audio, so that one accelerator serves many simultaneous sessions.
3. Our approach
The model is a frame-level target-speaker classification model: for every 80 ms frame it emits the probability that the enrolled person is speaking. Around it, the target speaker extraction system runs a speech recogniser on the mixture and keeps or drops each recognised word by its overlap with those frames. Classification rather than waveform extraction because the product consumes words, a probability can be set to a chosen operating point, and the system's error is an auditable mis-attributed word rather than a subtly altered voice.
How the pieces combine. Inside each trained block the frame states query the identity evidence tokens near them in time; the result is scaled by a learned gate that starts at zero and is added to the residual stream, so every fusion site begins as a pass-through. The binder pools each tracked speaker's own frames into one identity vector, compares it with the enrollment over a time × speaker grid, and a softmax over {speaker 1 … K, none} gives one ownership score per speaker; the per-frame logit and the binder's logit are blended by a learned weight and a sigmoid gives p(target), and frames at or above the operating threshold keep the words they overlap. ❄ frozen, 🔥 trained.
Training data. Real and synthetic conversations extracted from 10,000+ hours of audio from public corpora and podcasts.
Baselines
- No processing. Every word attributed to the wearer; its F1 is the floor.
- Sortformer v1 and Streaming Sortformer v2 (NVIDIA): end-to-end diarizers, 4 speakers, run on 90 s and 15 s windows respectively.
- Nemotron 3 Diarization (NVIDIA) and DiariZen WavLM-Large s80 v2 (BUT): leaderboard diarizers, up to 8 speakers, run over the whole recording.
- VibeVoice-ASR 7B (Microsoft): speaker-attributed recogniser.
- ElevenLabs Scribe v1 and v2: commercial diarizing transcribers; v1 run with the speaker count fixed at two and left free. With the count fixed, a three-way conversation is merged into the wearer's speaker and no cut-off can separate it again (Study 4).
- Meta Muse Voice Transcribe: streaming ASR with diarization, returning speaker-labelled turns.
- Our earlier models. Isolate 1.0, Isolate 1.1 and Isolate 1.1 + SBM-1, the generation before Isolate 2.0.
Every diarizer's anonymous speakers are verified against the enrollment with a speaker-identity model at a cut-off chosen at best F1 on the synthetic-mixture set (Figure 2).
Every diarizer baseline is read signature-matched: the speaker(s) whose pooled audio embeds within a fixed cosine threshold of the enrollment, a system that could actually ship. At a fixed cut-off its output is a hard decision, so it is a single point; sweeping the cut-off gives its full sweep, one point per speaker track.
Figure 2 · How the diarizer baselines are scored. Each diarizer runs once over the recording; each anonymous speaker's pooled audio is embedded and compared with the enrollment by cosine. At a fixed cut-off the output is a hard decision, one operating point; sweeping the cut-off changes nothing until it crosses a track's score, and then the whole track enters or leaves at once: one operating point per speaker track. A product built this way can operate only at those few points, and which points exist depends on who was in the room.
4. Highlights
Two generations of the model appear in this report. Isolate 2.0 + SBM-1 is the current model and the one these highlights describe. The generation-1 rows (Isolate 1.0, Isolate 1.1 and Isolate 1.1 + SBM-1) are kept because the privacy result, the effect of the Speaker Binding Module, was established on them; every number in the text names the model it belongs to.
- One threshold, every conversation. At a single operating point Isolate 2.0 + SBM-1 reaches a mean F1 of 0.958 across the six evaluation sets, against 0.925 for the best diarize-then-verify baseline, and loses at most 0.017 F1 to per-set tuning (Figure 3, Table 3).
- The hardest set moved most. On Study 4, a fifteen-minute real conversation with a similar-sounding second talker, AP rose from 0.912 (Isolate 1.0) to 0.981 and recall at precision ≥ 0.95 from 0.43 to 0.94 (Table 3).
- Whole utterances stay separated. At their reported cut-offs the diarize-then-verify baselines admit 1.7–3.1× more long (seven-word-plus) non-target utterances than Isolate 2.0 + SBM-1 on Studies 1–4: when a diarizer merges two voices, the whole track passes.
- Knows when the target is absent. Handed the enrollment of someone not in the room, Isolate 2.0 + SBM-1 emits 1.6–3.4 % of frames at threshold 0.5, down from 8–11 % for Isolate 1.0 (Table 4, Figure 5).
Evaluation sets
Table 1 · Evaluation sets. What each set is, where the audio comes from, and how the ground truth was made.
| set | what it is | audio | ground truth |
|---|---|---|---|
| Study 1 | real-world conversations, human-labelled | 178 min | human span labels, sampling-weighted |
| Study 2 | real-world conversations, human-labelled | 120 min | human span labels, sampling-weighted |
| Study 3 | real-world conversation, human-labelled | 2.4 min | human target / not-target spans |
| Study 4 | real-world conversation, human-labelled | 15.9 min | human target / not-target spans |
| Enacted retail conversations | 16 two-person retail and restaurant scenes acted by consenting staff, device microphone, real rooms and background | 60.5 min, 73–564 s each | frame-level target/other at 12.5 Hz; a pool of 6–7 enrollments of people not in the scene |
| Synthetic mixtures | 2,400 two-speaker 15 s augmented mixtures, speakers disjoint from training | 600 min | exact per-speaker timeline; noise at 5–15 dB SNR |
5. Results on the evaluation sets
Figure 3b · Does a threshold picked elsewhere hold? A: each set is scored at the threshold picked on the other five. B: the threshold is picked on one set and scored on the other five, worst pick ringed. Our models' F1 barely moves with the threshold, so a setting picked in any one set holds in the others; a cosine cut-off has a narrow optimum that moves with the set, so one unlucky calibration set costs the best baselines a third of their F1. For the product: one threshold, no per-site tuning, and no failure mode from calibrating on the wrong set.
Why a trained model, and why diarize-then-verify is not enough here
A diarizer plus a speaker verifier is the obvious way to build this product, and with the strongest open verifier it is a serious baseline: at a well-chosen cut-off it reaches the same peak F1 as our earlier models (Figure 3b, A), and it rejects an absent speaker more cleanly than we do (Table 4). The case for a trained model rests on what the curves in Figure 3 say a system can do, not on any single operating point.
A curve, not a handful of points. Our model scores every 80 ms frame, so every point on its precision–recall curve is a real operating point reached by one threshold. A diarize-then-verify system gives each speaker track one cosine, so within a conversation its sweep moves in track-sized jumps: between them nothing changes, and at each one a whole speaker switches in or out. A product that must meet a mis-attribution budget needs a curve to sit on; a track-level system offers a handful of points, and which points depends on the conversation (Figure 3b measures this).
The high-precision end, where the product lives. At a 5 % mis-attribution budget Isolate 2.0 keeps 96.8 % of the wearer's speech against 91.9 % for the best baseline, and it leads on every set but Study 4 (span-level units; the word-level tab of Table 3 gives 95.8 % against 80.6 %). The track-level sweeps bend early because admitting one impostor track costs a block of precision at once and nothing finer can be given back; the per-frame curve gives up precision gradually.
Concurrent speech. The largest AP gaps are on the two sets with the most overlap, the synthetic mixtures and enacted retail. A track-level decision cannot split a frame between two people who are both talking; a per-frame score can, and the difference is exactly the kind of speech a retail conversation is made of.
A threshold that travels. Figure 3b shows the same threshold moving from one set to the next. Because the frame loss is trained over many conditions, the F1 optimum is a plateau (0.60–0.85 for Isolate 2.0) and a threshold picked in any one set lands on it: the worst single-set calibration still gives 0.932 mean F1 elsewhere. A cosine cut-off has a narrow optimum whose position moves with the set, the recording length, how much audio went into each pooled vector, and the verifier: picked on Study 3 it costs Nemotron 3 and Muse a third of their F1 everywhere else.
Borrowed quality. The shortcut is only as good as the two frozen models it glues together. Its ceiling is set by the diarization it is handed, its ranking comes entirely from the verifier, and its time resolution and speaker count from the diarizer. When any of those changes, the cut-off, the pooling length and the mapping must be re-tuned by hand, and nothing in the system learns from our data. Whole-recording diarization also means the decision waits for the whole conversation; run on 90 s windows instead, as an online system would have to, Nemotron 3 gives up its Study 4 point at the fixed cut-off (F1 0.956 → 0.939, recall at P ≥ 0.95 0.967 → 0.926) and is otherwise unchanged; the 90 s runs of both are in the Threshold transferability tab of Figure 3.
It can be trained further; a heuristic can only be re-tuned. Every gap this report has found became a training target: the identity evidence that took Isolate 1.0 to Isolate 1.1, the Speaker Binding Module, which cut wrong-enrollment fires 1.4–3.1×, the block-length and backbone changes behind Isolate 2.0, and now hard-absent references in the objective so that the absent-speaker case, the one place the shortcut still wins, is learned rather than filtered. The same path is open to word-level decisions, cross-block memory and per-deployment fine-tuning on our own recordings. A diarize-then-verify system has no such path: its errors are fixed by whichever public checkpoints it is built from.
What we take from it. The track-level decision does two things well that the per-frame score does not: once a track is admitted every frame of it comes through, including one-word backchannels, and once it is rejected every frame of an impostor goes with it. Those are the recall-saturation and clean-rejection properties behind its lower false-accept rate, and they are why the Speaker Binding Module and the presence veto of Section 9 exist: to add a once-per-speaker decision on top of the per-frame score, inside a model that keeps learning.
Five models: Isolate 1.0, Isolate 1.1 (stronger identity evidence) and Isolate 1.1 + SBM-1 at 90 s blocks; Isolate 2.0 and Isolate 2.0 + SBM-1 at 30 s blocks. Study 1 and 2 are scored by word fraction: each labelled span counts by its recognised words times its sampling weight.
In Tables 3 to 5 the highlighted row is the current model, Isolate 2.0 + SBM-1 (30 s blocks, checkpoint 12500), and bold marks the best value in each column.
Table 3 · Main results. Average precision, recall at 95 % precision and best F1 per set for the five models; no processing is the F1 of accepting everything. Isolate 2.0 + SBM-1 has the highest average precision on five of the six sets (Study 3 is a tie), and the gains over generation 1 are largest where speech overlaps most, the synthetic mixtures and enacted retail, and on the hardest real conversation, Study 4.
| set | metric | no processing | Isolate 1.0 | Isolate 1.1 | Isolate 1.1 + SBM-1 | Isolate 2.0 | Isolate 2.0 + SBM-1 |
|---|---|---|---|---|---|---|---|
| Study 1 | AP | — | 0.977 | 0.975 | 0.979 | 0.988 | 0.987 |
| Study 1 | recall @ P ≥ .95 | — | 0.918 | 0.898 | 0.880 | 0.953 | 0.939 |
| Study 1 | best F1 | 0.638 | 0.935 | 0.933 | 0.935 | 0.954 | 0.954 |
| Study 2 | AP | — | 0.983 | 0.984 | 0.988 | 0.994 | 0.992 |
| Study 2 | recall @ P ≥ .95 | — | 0.866 | 0.880 | 0.901 | 0.970 | 0.970 |
| Study 2 | best F1 | 0.678 | 0.927 | 0.945 | 0.959 | 0.963 | 0.963 |
| Study 3 | AP | — | 1.000 | 1.000 | 1.000 | 0.998 | 0.998 |
| Study 3 | recall @ P ≥ .95 | — | 1.000 | 1.000 | 1.000 | 0.993 | 0.993 |
| Study 3 | best F1 | 0.794 | 0.994 | 0.994 | 0.994 | 0.994 | 0.994 |
| Study 4 | AP | — | 0.912 | 0.971 | 0.977 | 0.980 | 0.981 |
| Study 4 | recall @ P ≥ .95 | — | 0.434 | 0.853 | 0.886 | 0.932 | 0.943 |
| Study 4 | best F1 | 0.736 | 0.868 | 0.926 | 0.936 | 0.956 | 0.959 |
| Enacted retail | AP | — | 0.982 | 0.985 | 0.984 | 0.988 | 0.988 |
| Enacted retail | recall @ P ≥ .95 | — | 0.981 | 0.983 | 0.987 | 0.986 | 0.987 |
| Enacted retail | best F1 | 0.792 | 0.966 | 0.968 | 0.969 | 0.977 | 0.977 |
| Synthetic mixtures | AP | — | 0.949 | 0.955 | 0.967 | 0.992 | 0.992 |
| Synthetic mixtures | recall @ P ≥ .95 | — | 0.761 | 0.772 | 0.849 | 0.974 | 0.974 |
| Synthetic mixtures | best F1 | 0.683 | 0.895 | 0.897 | 0.915 | 0.965 | 0.966 |
Table 3 · Main results. Average precision, recall at 95 % precision and best F1 per set for the five models; no processing is the F1 of accepting everything. Isolate 2.0 + SBM-1 has the highest average precision on five of the six sets (Study 3 is a tie), and the gains over generation 1 are largest where speech overlaps most, the synthetic mixtures and enacted retail, and on the hardest real conversation, Study 4.
| set | metric | no processing | Isolate 1.0 | Isolate 1.1 | Isolate 1.1 + SBM-1 | Isolate 2.0 | Isolate 2.0 + SBM-1 |
|---|---|---|---|---|---|---|---|
| Study 1 | AP | — | 0.989 | 0.983 | 0.986 | 0.990 | 0.991 |
| Study 1 | recall @ P ≥ .95 | — | 0.919 | 0.864 | 0.862 | 0.949 | 0.940 |
| Study 1 | best F1 | 0.816 | 0.945 | 0.932 | 0.938 | 0.961 | 0.962 |
| Study 2 | AP | — | 0.986 | 0.990 | 0.992 | 0.995 | 0.995 |
| Study 2 | recall @ P ≥ .95 | — | 0.870 | 0.928 | 0.964 | 0.995 | 0.992 |
| Study 2 | best F1 | 0.725 | 0.937 | 0.952 | 0.959 | 0.973 | 0.973 |
| Study 3 | AP | — | 0.998 | 0.998 | 0.998 | 0.998 | 0.998 |
| Study 3 | recall @ P ≥ .95 | — | 0.996 | 0.996 | 0.996 | 0.996 | 0.996 |
| Study 3 | best F1 | 0.784 | 0.985 | 0.985 | 0.985 | 0.983 | 0.983 |
| Study 4 | AP | — | 0.889 | 0.966 | 0.970 | 0.976 | 0.977 |
| Study 4 | recall @ P ≥ .95 | — | 0.376 | 0.816 | 0.837 | 0.831 | 0.855 |
| Study 4 | best F1 | 0.745 | 0.848 | 0.903 | 0.908 | 0.928 | 0.935 |
| Enacted retail | AP | — | 0.999 | 0.999 | 0.999 | 0.999 | 0.999 |
| Enacted retail | recall @ P ≥ .95 | — | 0.998 | 0.997 | 0.999 | 0.994 | 0.995 |
| Enacted retail | best F1 | 0.812 | 0.985 | 0.987 | 0.990 | 0.986 | 0.986 |
| Synthetic mixtures | AP | — | 0.966 | 0.973 | 0.983 | 0.992 | 0.992 |
| Synthetic mixtures | recall @ P ≥ .95 | — | 0.842 | 0.858 | 0.929 | 0.986 | 0.987 |
| Synthetic mixtures | best F1 | 0.688 | 0.913 | 0.919 | 0.941 | 0.971 | 0.972 |
Uncertainty. Intervals are a paired block bootstrap (1,000 resamples; blocks are labelled spans on Studies 1 and 2, utterances on Studies 3 and 4, conversations on the enacted-retail set and mixtures on the synthetic set, with every system scored on the same draws). Isolate 2.0 + SBM-1's AP lead over the best baseline on each set is significant on Study 2 (+0.015), Study 4 (+0.016), enacted retail (+0.033) and the synthetic mixtures (+0.040), and within the interval on Studies 1 and 3. Its AP gain over Isolate 1.1 + SBM-1 is significant on the synthetic mixtures (+0.025) and within the interval on the five real sets, where our models sit at most 0.01 AP apart against intervals of about that width; the fixed-threshold F1 gain is also significant on enacted retail. Within generation 1, Isolate 1.1's gain over Isolate 1.0 is significant on Study 4 (+0.059 AP) and adding SBM-1 is significant on Study 4 and the synthetic set.
6. Does the model know when the target is not there?
This is the headline experiment for a privacy claim, and the one most published target-speaker work skips.
Figure 4 · The wrong-enrollment protocol. Every recording is scored twice with the same model and threshold: once with the wearer's enrollment, which gives recall and precision, and once with the enrollment of someone who is not in the recording, where every emitted frame is a false accept. This is the privacy test: it measures what leaves the device when the wearer is not speaking at all.
Absent-speaker false-accept rate: fraction of frames emitted when the enrollment belongs to someone not in the recording, averaged over the wrong enrollments, at threshold 0.5.
Table 4 · Absent-speaker false-accept rate at threshold 0.5: the share of all frames, silence included, emitted under an absent person's enrollment, averaged over several absent people per recording. Adding SBM-1 cuts it on every set, by 1.4–3.1× on generation 1 and 1.2–2.3× on Isolate 2.0. The 1.6–3.4 % that remains is the one place the diarize-then-verify shortcut still wins, and what Sections 9 and 10 are about.
| system | Study 1 | Study 2 | Study 3 | Study 4 | Enacted retail |
|---|---|---|---|---|---|
| Isolate 1.0 | 9.72 % | 9.08 % | 11.07 % | 9.33 % | 7.96 % |
| Isolate 1.1 | 5.89 % | 7.77 % | 8.25 % | 5.33 % | 4.14 % |
| Isolate 1.1 + SBM-1 | 3.68 % | 4.29 % | 3.51 % | 1.74 % | 3.01 % |
| Isolate 2.0 | 3.96 % | 4.03 % | 4.82 % | 3.63 % | 3.11 % |
| Isolate 2.0 + SBM-1 | 2.85 % | 2.75 % | 3.35 % | 1.56 % | 2.64 % |
Table 4b · The same rate at each model's operating point: the lowest threshold at which precision holds at or above 95 % on every set (the rule of the Threshold transferability tab of Figure 3), so the threshold differs per model. At its operating point every model emits under 3 % of frames under a wrong enrollment, and Isolate 2.0 + SBM-1 under 1 % on every set. Read together with Table 4, this says that the Speaker Binding Module matters less for Isolate 2.0, whose stronger identity evidence and trunk already hold false accepts low at the operating point, than it did for generation 1, where it cut the rate by up to 3×: at their 95 % thresholds Isolate 2.0 and Isolate 2.0 + SBM-1 keep the same 94 % of the wearer's speech on average, the binder lowering false accepts on four of the five sets. The generation-1 rates look good here only because their operating points are conservative: at those thresholds Isolate 1.1 keeps 87 % of the wearer's speech on average and Isolate 1.1 + SBM-1 90 %. Figure 5 shows the whole recall–false-accept trade-off.
| system | threshold | Study 1 | Study 2 | Study 3 | Study 4 | Enacted retail |
|---|---|---|---|---|---|---|
| Isolate 1.0 | 0.900 | 1.51 % | 1.62 % | 0.38 % | 0.42 % | 0.97 % |
| Isolate 1.1 | 0.800 | 1.95 % | 2.63 % | 1.35 % | 1.01 % | 1.20 % |
| Isolate 1.1 + SBM-1 | 0.775 | 1.28 % | 1.28 % | 0.95 % | 0.35 % | 0.76 % |
| Isolate 2.0 | 0.875 | 1.05 % | 1.06 % | 0.88 % | 0.45 % | 0.87 % |
| Isolate 2.0 + SBM-1 | 0.825 | 0.97 % | 0.89 % | 0.60 % | 0.26 % | 0.87 % |
Two findings. Adding SBM-1 cuts the false-accept rate of the model it extends on every set, by 1.4× (enacted retail) to 3.1× (Study 4), at the same threshold and with higher recall; speaker-level identity was what it was designed to add. And it does not reach zero: at this threshold Isolate 1.1 + SBM-1 still emits 1.7–4.3 % of frames under a wrong enrollment, and Isolate 2.0 + SBM-1 1.6–3.4 %. What those false accepts look like, and how to remove them without giving up frame-level decisions, is taken up in Sections 9 and 10.
7. Recall vs false accepts: the privacy–utility curve
Figure 3 answered "when both people talk, who gets the words?" This figure answers the other question: when the wearer is silent and someone else talks, how much comes through? Sweeping the same threshold on the enacted-retail set, with recall of the wearer's speech against the fraction of frames emitted under a wrong enrollment, gives the curve a deployment actually chooses from. Neither curve substitutes for the other; they are the two branches of Section 1.
At threshold 0.5 the five models sit at (false-accept rate, recall): Isolate 1.0 (8.0 %, 96.8 %), Isolate 1.1 (4.1 %, 96.9 %), Isolate 1.1 + SBM-1 (3.0 %, 98.0 %), Isolate 2.0 (3.1 %, 98.0 %) and Isolate 2.0 + SBM-1 (2.6 %, 97.8 %). The ● on the curves marks the threshold chosen with the slider.
Two further properties of the curve are easy to miss from a scalar.
The best threshold moves between sets. For a single model the F1-maximising threshold on production units ranges from 0.325 on Study 4 to 0.875 on Study 3, with the study and enacted-retail sets in between. "Threshold 0.5" is therefore a different operating point in every set. Our operating-point rule is to fix a leakage budget on the wrong-enrollment sweep and derive the threshold from it, per deployment class; the production value itself is not published.
8. Does it preserve the information we care about?
Every set is now scored on recogniser words (Section 11), so the emitted transcript can be scored against the human target labels on all six as a word attribution error rate:
(target words not emitted + non-target words emitted) / target words
It has deletions (missed target words) and insertions (leaked non-target words) but no substitutions, because every word comes from the same recogniser pass and the labels say who spoke, not what was said; so it measures attribution, not recognition, and is labelled accordingly rather than as WER.
Table 5 · Word attribution error: missed target words plus leaked non-target words, over target words, at four operating points. The split matters more than the total: on the real conversations the remaining error is leaked words at turn boundaries rather than missed speech, the error a transcript reader can see and the one the next steps target.
| set · model | accept everything | at threshold 0.5 | miss / leak | at the P ≥ 0.95 operating point | miss / leak | best over thresholds |
|---|---|---|---|---|---|---|
| Study 1 · Isolate 1.0 | 0.450 | 0.111 | 0.050 / 0.061 | 0.131 | 0.084 / 0.047 | 0.111 |
| Study 1 · Isolate 1.1 | 0.450 | 0.156 | 0.065 / 0.092 | 0.182 | 0.139 / 0.043 | 0.142 |
| Study 1 · Isolate 1.1 + SBM-1 | 0.450 | 0.127 | 0.058 / 0.069 | 0.190 | 0.149 / 0.041 | 0.126 |
| Study 1 · Isolate 2.0 | 0.450 | 0.088 | 0.017 / 0.072 | 0.103 | 0.056 / 0.047 | 0.079 |
| Study 1 · Isolate 2.0 + SBM-1 | 0.450 | 0.077 | 0.018 / 0.059 | 0.109 | 0.062 / 0.048 | 0.077 |
| Study 2 · Isolate 1.0 | 0.758 | 0.163 | 0.048 / 0.115 | 0.161 | 0.133 / 0.029 | 0.135 |
| Study 2 · Isolate 1.1 | 0.758 | 0.123 | 0.060 / 0.064 | 0.126 | 0.078 / 0.048 | 0.101 |
| Study 2 · Isolate 1.1 + SBM-1 | 0.758 | 0.097 | 0.029 / 0.068 | 0.083 | 0.036 / 0.047 | 0.083 |
| Study 2 · Isolate 2.0 | 0.758 | 0.057 | 0.014 / 0.043 | 0.058 | 0.008 / 0.050 | 0.055 |
| Study 2 · Isolate 2.0 + SBM-1 | 0.758 | 0.065 | 0.029 / 0.036 | 0.058 | 0.008 / 0.050 | 0.056 |
| Study 3 · Isolate 1.0 | 0.552 | 0.052 | 0.004 / 0.048 | 0.056 | 0.004 / 0.052 | 0.030 |
| Study 3 · Isolate 1.1 | 0.552 | 0.041 | 0.004 / 0.037 | 0.056 | 0.004 / 0.052 | 0.030 |
| Study 3 · Isolate 1.1 + SBM-1 | 0.552 | 0.056 | 0.004 / 0.052 | 0.056 | 0.004 / 0.052 | 0.030 |
| Study 3 · Isolate 2.0 | 0.552 | 0.093 | 0.004 / 0.089 | 0.056 | 0.004 / 0.052 | 0.037 |
| Study 3 · Isolate 2.0 + SBM-1 | 0.552 | 0.078 | 0.004 / 0.074 | 0.056 | 0.004 / 0.052 | 0.033 |
| Study 4 · Isolate 1.0 | 0.684 | 0.457 | 0.365 / 0.093 | 0.654 | 0.635 / 0.019 | 0.334 |
| Study 4 · Isolate 1.1 | 0.684 | 0.195 | 0.101 / 0.094 | 0.244 | 0.206 / 0.038 | 0.192 |
| Study 4 · Isolate 1.1 + SBM-1 | 0.684 | 0.187 | 0.085 / 0.102 | 0.212 | 0.169 / 0.043 | 0.185 |
| Study 4 · Isolate 2.0 | 0.684 | 0.165 | 0.030 / 0.135 | 0.241 | 0.202 / 0.038 | 0.151 |
| Study 4 · Isolate 2.0 + SBM-1 | 0.684 | 0.138 | 0.037 / 0.102 | 0.199 | 0.156 / 0.043 | 0.135 |
| Enacted retail · Isolate 1.0 | 0.462 | 0.032 | 0.019 / 0.013 | 0.049 | 0.002 / 0.046 | 0.030 |
| Enacted retail · Isolate 1.1 | 0.462 | 0.027 | 0.019 / 0.008 | 0.052 | 0.003 / 0.049 | 0.025 |
| Enacted retail · Isolate 1.1 + SBM-1 | 0.462 | 0.020 | 0.008 / 0.012 | 0.052 | 0.001 / 0.051 | 0.019 |
| Enacted retail · Isolate 2.0 | 0.462 | 0.031 | 0.013 / 0.018 | 0.054 | 0.006 / 0.048 | 0.029 |
| Enacted retail · Isolate 2.0 + SBM-1 | 0.462 | 0.029 | 0.015 / 0.014 | 0.053 | 0.005 / 0.048 | 0.029 |
| Synthetic mixtures · Isolate 1.0 | 0.907 | 0.175 | 0.089 / 0.087 | 0.204 | 0.161 / 0.043 | 0.175 |
| Synthetic mixtures · Isolate 1.1 | 0.907 | 0.177 | 0.120 / 0.058 | 0.187 | 0.142 / 0.045 | 0.166 |
| Synthetic mixtures · Isolate 1.1 + SBM-1 | 0.907 | 0.118 | 0.064 / 0.055 | 0.120 | 0.073 / 0.047 | 0.118 |
| Synthetic mixtures · Isolate 2.0 | 0.907 | 0.059 | 0.025 / 0.034 | 0.065 | 0.015 / 0.050 | 0.058 |
| Synthetic mixtures · Isolate 2.0 + SBM-1 | 0.907 | 0.057 | 0.026 / 0.031 | 0.064 | 0.014 / 0.050 | 0.057 |
At 0.725 Isolate 1.1 + SBM-1 reads 0.188 on Study 4 (0.126 missed, 0.062 leaked) and Isolate 2.0 0.154 (0.060, 0.094); on Study 1, where one-word turns and cross-talk are densest, 0.166 and 0.084. On words, Study 3 is close to solved (4.1 % error for Isolate 1.1 + SBM-1 at 0.725, 1.1 % missed and 3.0 % leaked). Study 4 is where the successive changes show: at 0.725 Isolate 1.0 misses 58 % of the wearer's words, Isolate 1.1 15 %, Isolate 1.1 + SBM-1 13 %, at 4, 6 and 6 % leakage. Across the six sets Isolate 2.0's error at 0.725 ranges from 0.030 (Enacted retail) to 0.154 (Study 4); the split into miss and leak is again the useful part: on the real conversations the remaining error is mostly leaked words at turn boundaries, not missed speech.
9. Post-hoc filtering
The model emits one probability per frame; everything downstream of that is a filter on its mask. Filters are attractive because they leave the trained model untouched, can be A/B-tested per session, and can target the two error types separately: leaked words from another speaker inside an accepted stretch, and false accepts when the enrolled speaker is absent. This section measures one filter and describes three more.
A diarizer as a leak filter (measured)
If the model works, the audio it emits is mostly the target, and whatever else gets through arrives as short, scattered pieces of other voices. That suggests a cheap clean-up stage: diarize the emitted audio and throw away every speaker except the dominant one. We tested it on Isolate 1.1 + SBM-1, with two diarizers, on every evaluation set, for the true enrollment and for every wrong one.
- Filtered audio. The mixture multiplied by the mask of Isolate 1.1 + SBM-1 at the operating threshold 0.5, with 10 ms ramps at the mask edges: exactly what the product would emit.
- Diarize it. Sortformer v1 on non-overlapping 90 s windows (its training length), or Nemotron 3 Diarization on the whole recording (up to eight speakers).
- Keep the primary speaker. The speaker with the most activity inside the emitted frames is taken to be the target: per 90 s window for Sortformer v1, per recording for Nemotron 3.
- Remove the leaks. Frames where only a non-primary speaker is active get p(target) = 0. Overlap with the primary speaker, frames with no active speaker, and everything below 0.5 are left unchanged, so the curves differ from the unfiltered model only above the operating threshold.
Table 6 · The post-hoc diarizer filter on Isolate 1.1 + SBM-1. It trims leaked words where several voices get through, costs a few points of recall where the wearer is not the dominant voice, and leaves the absent-speaker false-accept rate essentially unchanged: a filter cannot stand in for identity inside the model.
| set | system | AP | R @ P ≥ .95 | P @ 0.5 | R @ 0.5 | error @ 0.5 | miss / leak | best error | FAR @ 0.5 | frames removed true / wrong |
|---|---|---|---|---|---|---|---|---|---|---|
| Study 1 | unfiltered | 0.979 | 0.880 | 0.879 | 0.946 | 0.184 | 0.054 / 0.130 | 0.135 | 3.68 % | — |
| Study 1 | + Sortformer v1 | 0.981 | 0.897 | 0.923 | 0.946 | 0.133 | 0.054 / 0.079 | 0.122 | 3.55 % | 3.3 % / 3.4 % |
| Study 1 | + Nemotron 3 | 0.979 | 0.916 | 0.926 | 0.935 | 0.140 | 0.065 / 0.075 | 0.120 | 3.50 % | 4.0 % / 4.8 % |
| Study 2 | unfiltered | 0.988 | 0.901 | 0.932 | 0.988 | 0.084 | 0.012 / 0.072 | 0.084 | 4.29 % | — |
| Study 2 | + Sortformer v1 | 0.988 | 0.961 | 0.931 | 0.961 | 0.111 | 0.039 / 0.071 | 0.077 | 4.16 % | 1.5 % / 3.1 % |
| Study 2 | + Nemotron 3 | 0.973 | 0.926 | 0.957 | 0.925 | 0.116 | 0.075 / 0.041 | 0.112 | 4.11 % | 4.7 % / 4.2 % |
| Study 3 | unfiltered | 1.000 | 1.000 | 0.964 | 1.000 | 0.037 | 0.000 / 0.037 | 0.011 | 3.51 % | — |
| Study 3 | + Sortformer v1 | 1.000 | 1.000 | 0.964 | 1.000 | 0.037 | 0.000 / 0.037 | 0.011 | 3.08 % | 2.6 % / 5.5 % |
| Study 3 | + Nemotron 3 | 1.000 | 1.000 | 0.964 | 1.000 | 0.037 | 0.000 / 0.037 | 0.011 | 3.50 % | 2.7 % / 18.9 % |
| Study 4 | unfiltered | 0.977 | 0.886 | 0.927 | 0.939 | 0.135 | 0.061 / 0.074 | 0.135 | 1.74 % | — |
| Study 4 | + Sortformer v1 | 0.968 | 0.853 | 0.930 | 0.917 | 0.153 | 0.083 / 0.069 | 0.153 | 1.68 % | 2.6 % / 5.5 % |
| Study 4 | + Nemotron 3 | 0.977 | 0.914 | 0.944 | 0.931 | 0.124 | 0.069 / 0.055 | 0.124 | 1.31 % | 2.7 % / 18.9 % |
| Enacted retail | unfiltered | 0.984 | 0.987 | 0.957 | 0.980 | 0.064 | 0.020 / 0.044 | 0.063 | 3.01 % | — |
| Enacted retail | + Sortformer v1 | 0.983 | 0.980 | 0.961 | 0.970 | 0.070 | 0.030 / 0.039 | 0.067 | 2.98 % | 1.0 % / 0.8 % |
| Enacted retail | + Nemotron 3 | 0.983 | 0.982 | 0.965 | 0.966 | 0.069 | 0.034 / 0.035 | 0.064 | 2.91 % | 1.2 % / 3.2 % |
| Synthetic mixtures | unfiltered | 0.967 | 0.849 | 0.913 | 0.915 | 0.172 | 0.085 / 0.087 | 0.172 | 8.00 % | — |
| Synthetic mixtures | + Sortformer v1 | 0.965 | 0.846 | 0.920 | 0.900 | 0.178 | 0.100 / 0.078 | 0.178 | 7.88 % | 2.3 % / 1.5 % |
| Synthetic mixtures | + Nemotron 3 | 0.967 | 0.857 | 0.930 | 0.898 | 0.169 | 0.102 / 0.067 | 0.169 | 7.82 % | 2.8 % / 2.3 % |
Columns: AP, recall at precision ≥ 0.95, precision, recall and attribution error at threshold 0.5, the best attribution error over thresholds, absent-speaker false-accept rate at 0.5, and the share of emitted frames the filter removed (correctly / wrongly). R = recall, P = precision, error = attribution error, FAR = absent-speaker false-accept rate; the last column is the share of emitted frames the filter removed. 90 s blocks, units and FAR as in Sections 11 and 6. Attribution error = (missed target mass + leaked non-target mass) / target mass: words on Study 3 and Study 4 (so it is the word-level error of Section 8), sampling-weighted words on Study 1 and 2, frames on the enacted-retail and synthetic-mixture sets.
What it does. The filter removes little: 1–5 % of the emitted frames with the true enrollment. Where the leakage is made of several other voices it pays off: on Study 1 both diarizers cut leaked words from 0.130 to about 0.08 of the target's words, lowering the error at 0.5 from 0.184 to 0.133 (Sortformer v1) and 0.140 (Nemotron 3), and Nemotron 3 raises recall at P ≥ .95 from 0.880 to 0.916. On Study 4, Nemotron 3 raises recall at P ≥ .95 from 0.886 to 0.914 and lowers the word error from 0.135 to 0.124.
What it does not do. It barely moves the absent-speaker false-accept rate (Study 4 1.74 → 1.31 %, elsewhere 0.0–0.4 points). Under a wrong enrollment the model usually emits one impostor voice, and the filter promptly promotes that voice to "primary". It also costs recall wherever the target is not the dominant voice in what was emitted, or the diarizer splits the target in two: Study 2 loses 6 points of recall at 0.5 with Nemotron 3 (error 0.084 → 0.116), and Sortformer v1, deciding per 90 s window, hurts Study 4 (error 0.135 → 0.153, recall at P ≥ .95 0.886 → 0.853). The enacted-retail and synthetic-mixture sets, where leakage was already rare, move by less than 0.01 in AP and error.
Verdict. Diarizing the emitted audio is a cheap way to trim leaked words when several voices get through, at the price of a few points of recall where the target is not the dominant voice; the whole-recording diarizer (Nemotron 3) is the safer of the two on a long conversation. It is a leakage filter, not an identity check, and it does not address false accepts.
Other post-hoc filters
- A speaker-level pre-check. Verify each anonymous speaker the model already tracks against the enrollment once, over its pooled audio, before any of its frames are emitted, and let the frame-level output through only inside accepted tracks. The diarizer filter above is the after-the-fact version and trims leaked words but not false accepts; a pre-check targets the false accepts directly, at the cost of waiting for enough of a speaker's audio to verify.
- A session-level presence decision. The Speaker Binding Module already asks which speaker is the target, including "none of them". Aggregating that answer over a turn or a window rather than per frame gives a presence decision that can veto emission when no speaker matches the enrollment for long enough, while frames inside a matched turn keep their frame-level resolution.
- Run-length and smoothing rules. Most falsely emitted time is in runs shorter than 3 s, so a minimum run length or hysteresis on the emitted mask removes a large share of it at the cost of a little latency on short target turns. This is a scoring-time change and can be measured on the existing outputs.
10. Next steps
The caveat of frame-level decisions. The model decides every 80 ms frame from the audio around it. That is what keeps concurrent speakers apart and keeps whole utterances from leaking when voices are similar, but it also means that when the enrolled speaker is absent, brief stretches of ambiguous identity (backchannels, overlap, a voice close to the enrollment) occasionally cross the threshold on their own. The result is the 1.6–3.4 % of frames in Table 4 for Isolate 2.0 + SBM-1: many short bursts rather than sustained runs, since no single frame has enough evidence to sustain a false accept for long. A system that instead verifies each speaker once, over all of their pooled audio, makes that mistake far less often, at the price of accepting a whole track, leaks included, whenever it accepts it.
Section 9 lists the post-hoc filters that could close this gap without touching the per-frame model; the diarizer filter is measured, the other three are next in line, starting with the run-length rule because it can be scored on the outputs we already have.
Beyond false accepts, the same evaluation points at two further items: conversations with many talkers, where a trunk that supports more simultaneous talkers is the natural fix, and finer time resolution at word boundaries.
The four directions below go further than filters. Each starts from a problem this evaluation measured and sketches a potential solution from first principles, with references to work where that mechanism has shown promise; none of them is run yet.
Words, not frames: use the transcript itself
The problem. The model decides per 80 ms frame from audio alone, but the product's unit is the word. The utterance-length control on Figure 3 shows where that costs: one- and two-word units have the lowest F1 on every set, and the word-error table (Table 5) shows leaks and misses concentrated at turn boundaries, where a threshold crossing lands mid-word and splits it. The transcript carries evidence about ownership that the audio path never sees: a question is followed by the other speaker's answer, a backchannel is one word long, a name addresses someone, a language switch marks a speaker in a bilingual conversation.
Potential solution. Decide once per word rather than once per frame: pool the frame probabilities inside each aligned word and threshold the pooled value, so a word cannot be half accepted. This needs only the recogniser's word timestamps and can be scored today on Studies 3 and 4. Then let the text carry evidence about ownership that the audio path never sees: a light re-scorer over the word sequence reads per-word probability statistics, the neighbouring decisions, punctuation and turn marks, and simple dialogue cues (a question followed by an answer, a one-word backchannel, a name that addresses someone), trained on the labelled studies and on pseudo-labelled real conversations. Correcting speaker labels from the text after the fact has been shown to remove a large share of attribution errors [4], and recognising words and speaker labels in one sequence model is well established [5]. The end state is one decoder that reads audio and text together and emits the target decision as a token in the transcript stream.
Expected effect. Fewer split words and fewer one-word leaks, measured as word attribution error on Studies 3 and 4 and as the count of leaked long utterances, without changing the frame model.
Decide the boundary now and the identity later; train on the metric
The problem. The Speaker Binding Module decides once per block, so a block is the unit of both memory and commitment. On the hardest conversation, whole blocks bind to the wrong voice when a similar-sounding talker dominates the block, and nothing decided in the previous block helps. The frame loss the model is trained on does not see the operating point we ship: the best threshold moves between checkpoints while average precision barely changes, which is a calibration problem, and the absent-speaker false-accept rate is never in the objective at all.
Potential solution. Separate the two decisions in time. Emit frame activity immediately as provisional, let the Speaker Binding Module re-label a whole turn once it has heard enough of that speaker, then finalise. Carry each block's per-speaker identity vectors and verdicts into the next block as extra keys, so a speaker decided correctly stays decided; this is the role a speaker cache plays in streaming diarizers [2]. Train for the operating point we ship: fine-tune the binder and the blend with a sequence-level objective, a differentiable surrogate of F1 at the fixed threshold plus an absent-speaker penalty, on top of the frame loss, and learn how long to wait before committing with a reward that couples accuracy and latency, as sequence-level emission regularisation does for streaming recognisers [8].
Expected effect. The Study 4 block flips and the wandering best threshold are the two targets; the cross-block memory can be scored offline on the existing Study 4 blocks before anything else is trained. It also yields the session-level presence decision of Section 9 for free: a turn whose SBM verdict is "none" is vetoed at finalisation.
Finer time, more talkers
The problem. Two items already named above: 80 ms is coarse at word boundaries, where a frame is a third of a short word, and the model supports a limited number of simultaneous talkers. Both are properties of the trunk.
Potential solution. Both are questions of output head and capacity rather than of the trunk itself. Time resolution can be refined without running the trunk at a finer clock: a small head expands each trunk state into several sub-frame activity values trained on fine-grained labels, so word boundaries are placed at a finer step at almost no added cost; recent diarizers reach 10 ms resolution this way [3]. Capacity for more simultaneous talkers is a matter of the activity representation and a permutation-invariant training loss over more speaker slots [1][3], together with a speaker cache that carries identities across chunks [2].
Expected effect. Word-boundary precision on Studies 3 and 4 at the 10 ms head, and a path to conversations with more talkers, which today are outside the evaluation.
Stream without blocks: rolling state, selectable clock, semantic end of turn
The problem. The client processes 90 s blocks with a fresh state each time; the block edge is where identity is lost, memory and latency scale with the block, and the same weights cannot trade latency for accuracy per session. Turn ends are found acoustically, so a thinking pause and a real hand-over look alike, and the presence veto has no principled moment to fire. The transcript we score against still contains fillers and repetitions, which dilute the word-level metric.
Potential solution. Remove the block as the unit of state. A causal window that slides frame by frame over a rolling attention cache, with positions re-based as it moves, has no block edge at which identity can be lost and runs at constant memory [6]; the Speaker Binding Module's memory travels with it. A learned clock embedding lets one set of weights run at a coarse frame step in the field and a finer step offline, trading latency for accuracy per session without retraining. End of turn should be predicted rather than inferred from silence: end-of-turn heads at several horizons, trained on the turn labels we already have, separate a thinking pause from a hand-over [7] and give the binder's commit and the presence veto a principled moment to fire. Finally, score against a polished, aligned transcript so the word unit is the one users read.
Expected effect. No block-edge identity loss (measurable as the gap between 90 s blocks and the whole-recording run), a single deployable model with a latency knob, and a defensible moment for the presence decision. The end-of-turn heads can be trained and scored first, since they need no change to the frame model.
11. How we evaluate it
Model versions
Every model in this report is named Isolate major.minor. The major version is the trunk generation (the encoder that tracks who is speaking and carries the identity evidence); a new major is a new trunk. The minor version is a revision of the identity evidence or its fusion on the same trunk. Attachable modules are suffixes: SBM is the Speaker Binding Module, which decides once per speaker which of the speakers the trunk tracks, if any, is the enrolled person, and attaches to any trunk that provides speaker activity and identity evidence.
| model | trunk | identity evidence | attached module |
|---|---|---|---|
| Isolate 1.0 | generation 1 | v1 | — |
| Isolate 1.1 | generation 1 | v2 | — |
| Isolate 1.1 + SBM-1 | generation 1 | v2 | SBM-1 |
| Isolate 2.0 | generation 2 | v2 | — |
| Isolate 2.0 + SBM-1 | generation 2 | v2 | SBM-1 |
Terms
- Target speaker (wearer), enrollment — The person whose speech is kept, and the separate recording of them that identifies them to the system. The enrollment is never taken from the audio being scored.
- Identity evidence — How much each short window of the conversation sounds like the enrollment, computed by a speaker-embedding model and fed to our model.
- Speaker activity — Who is speaking when, without knowing who anyone is; produced inside the trunk.
- Trunk (generation) — The encoder that produces speaker activity and the representations our model classifies; the major version number.
- Speaker Binding Module (SBM) — An attachable module that makes one decision per speaker, which tracked speaker is the enrolled one or none of them, and places it in time through the speaker activity.
- Block — The length of audio processed at once at inference (90 s or 30 s). A run setting, not part of the model.
- Operating threshold, cut-off — The single probability above which our models emit a unit (0.725 unless stated); the cosine above which a diarize-then-verify baseline accepts a speaker track.
- Unit, unit scheme — One scored item. The span-level scheme uses product verification pieces, labelled spans and diarizer pieces; the word-level scheme uses one recognised word on every set (see the scoring-units tabs).
- Attribution error — Missed wearer words plus leaked other-speaker words, divided by wearer words.
- Absent-speaker false-accept rate — The fraction of frames emitted when the conversation is run with the enrollment of someone who is not in it.
- AP, recall at precision ≥ 0.95, best F1 — Area under the precision–recall curve; the share of the wearer's words kept while at most 5 % of emitted words belong to someone else, which is the product's operating requirement; the highest F1 over thresholds.
Three protocol decisions matter more than the set list. Enrollment is never an excerpt of the test audio. In the synthetic-mixture set the enrollment is a separate key utterance of the speaker; in the enacted-retail and study sets it is a recording made in a different session. Speakers in the synthetic-mixture set are disjoint from training. Scoring happens on the unit the product emits. For Studies 3 and 4 that is the recogniser's word units; for the study sets it is the human-labelled span; for the frame-labelled sets it is the recogniser-aligned piece. A frame-level score would flatter every system, because the alignment between frames and words is itself a source of error and it belongs inside the measurement.
Word-level alternative. Table 3 also has a word-level tab, which scores the recogniser's words on every set. The product emits a transcript, so each unit is one recognised word. On Studies 3 and 4 the words are the product transcript's own tokens; on the other four sets they come from one recogniser run over the mixture, used only for its word timestamps. A word's truth comes from each set's own labels: the human label of the utterance that contains it (Studies 3 and 4), the human span that contains it, carried with that span's sampling weight (Studies 1 and 2, whose labels are a weighted sample), or the majority of its frames in the frame-level truth, unsure frames excluded (enacted retail, synthetic mixtures). Words in spans labelled target or mostly target count as the wearer's; not target and mostly other as someone else's. Words in spans labelled mixed are left out: the label says both people spoke there without saying which words were whose, so any assignment would score label noise rather than the system; they are 32 % of the labelled word mass on Study 1, 10 % on Study 2 and under 5 % elsewhere. Words outside every labelled span are not scored. A word's score is the mean model probability over its frames, trimmed by 80 ms at each end (never more than a quarter of the word), the same guard the product applies. The default view, span-level, keeps the units above: product verification pieces on Studies 3 and 4, whole human spans weighted by their word count on Studies 1 and 2, and diarizer pieces weighted by frames on the two frame-labelled sets. The absent-speaker false-accept rate (Table 4, Figure 5) stays a fraction of frames: with nobody enrolled present there are no target words to count against.
Scoring geometry. Our generation-1 models are scored on non-overlapping 90 s blocks and the Isolate 2.0 models on 30 s blocks; Sortformer v1 runs on 90 s windows, Streaming Sortformer v2 on 15 s windows, and the whole-recording baselines on the full file. VibeVoice Streaming keeps its own native verifier. Every baseline is read on the same units as our models, at core frames, so precision–recall is comparable though unpaired, and wrong enrollments give the baselines' false-accept rate the same way. Thresholds and cosine cut-offs live on a 0.025 grid from 0.050 to 0.950; the transferability view reads exact per-span values at the starting thresholds and binned values, within about 0.001, while a slider is moved. Figure 6 clips values at 0.7.
Metrics
Each model emits a probability, so we sweep the threshold and report the whole curve, summarised by three numbers: average precision (AP), the area under the precision-recall curve; recall at precision ≥ 0.95, the operating point the product ships near; and best F1. Uncertainty is a 1,000-resample block bootstrap, paired across systems: labelled spans on Studies 1 and 2, utterances on Studies 3 and 4, conversations on the enacted-retail set and mixtures on the synthetic set are resampled, and every system is scored on the same draws.
Two conventional metrics are deliberately not headlined. ROC-AUC is insensitive to class balance and saturates near 0.99 for every model here, hiding the differences that matter; it is used here only as a threshold-free sanity check. Equal error rate presumes symmetric costs, and ours are not (a leaked sentence is not a missed word). Word error rate proper needs the wearer's true transcript, which the real sets lack; on these word units, recall at precision ≥ 0.95 is the fraction of the wearer's speech retained at a 5 % mis-attribution budget, in each set's labelled units (recognised words on the word-level tab), and is labelled as word attribution, not WER.
Leakage is measured as the absent-speaker false-accept rate: run each conversation again with an enrollment of a person who is not present (the enacted-retail set's own pool, plus a 14-person cross-set pool for the others) and count the fraction of frames emitted. This is a target-absent test on real conversational audio, not on silence.
12. References
- T. J. Park, I. Medennikov, K. Dhawan, W. Wang, H. Huang, N. R. Koluguri, K. C. Puvvada, J. Balam, B. Ginsburg. Sortformer: Seamless Integration of Speaker Diarization and ASR by Bridging Timestamps and Tokens. arXiv:2409.06656, 2024.
- I. Medennikov, T. J. Park, W. Wang, H. Huang, K. Dhawan, J. Balam, B. Ginsburg. Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering. arXiv:2507.18446, 2025.
- NVIDIA. Nemotron-3 Diarization, model card. huggingface.co/nvidia/Nemotron-3-Diarization, 2026.
- Q. Wang, Y. Huang, G. Zhao, E. Clark, W. Xia, H. Liao. DiarizationLM: Speaker Diarization Post-Processing with Large Language Models. arXiv:2401.03506, 2024.
- L. El Shafey, H. Soltau, I. Shafran. Joint Speech Recognition and Speaker Diarization via Sequence Transduction. Interspeech 2019; arXiv:1907.05337.
- Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. V. Le, R. Salakhutdinov. Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context. ACL 2019; arXiv:1901.02860.
- E. Ekstedt, G. Skantze. Voice Activity Projection: Self-supervised Learning of Turn-taking Events. Interspeech 2022; arXiv:2205.09812.
- J. Yu, C.-C. Chiu, B. Li, S. Chang, T. N. Sainath, Y. He, A. Narayanan, W. Han, A. Gulati, Y. Wu, R. Pang. FastEmit: Low-latency Streaming ASR with Sequence-level Emission Regularization. ICASSP 2021; arXiv:2010.11148.