At Ethosphere, we're building an end-to-end platform for making real-world conversations measurable.
A small wearable microphone captures a conversation. A target-speaker extraction system determines which speech belongs to the person wearing it. And our backend turns the resulting transcript into structured data that can be analyzed for things like coaching, customer experience, and operational insight.
The middle step exists for an important reason:
A microphone captures everybody.
The employee wearing the device may have explicitly opted in to being recorded. The customer standing across from them may not have. The same may be true of the coworkers talking nearby, the person walking past, or the people at the next table.
The public debate around increasingly capable wearable cameras and microphones is a useful reminder of something fairly intuitive: people care deeply about what these devices capture, particularly when they are not the person using them.
For us, that makes privacy a first-order technical problem. It is not enough to record everything and decide much later what should have been retained.
Our approach is to make the identity of the consenting speaker part of the capture pipeline itself.
We take a short enrollment recording of the wearer, with the goal of capturing it once and reusing it across sessions and days—even as their voice and recording conditions change. During a conversation, our Isolate model uses that enrollment to determine, moment by moment, whether the person speaking is the enrolled speaker. Speech attributed to the wearer continues through the transcription and analysis pipeline; speech belonging to customers, coworkers, and other people is discarded.
That gives us a useful middle ground: make real-world conversations measurable without treating every voice within microphone range as data we are entitled to collect.
But that simple idea hides most of the engineering challenge.
A useful system cannot respond to uncertainty by rejecting everything. That would reduce privacy risk, but it would also discard too much of the wearer’s own speech to be useful.
The opposite failure is more serious: a system that preserves nearly everything the wearer says but occasionally admits complete phrases or sentences from someone nearby is not doing its job.
So the problem is not simply to recognize the wearer. It is to preserve as much of their speech as possible while keeping speech from everyone else out.
So throughout this post we care about three related questions.
- When the wearer speaks, how much do we keep?
- When somebody else speaks, how much do we accidentally attribute to the wearer?
And an important but easy to overlook and surprisingly hard part:
- When the enrolled speaker is not in the conversation at all, does the system know that?
Those requirements have to hold outside a lab. Real conversations contain interruptions, overlapping speech, one-word acknowledgements, similar-sounding voices, music, reverberation, far-field pickup, and enrollment audio recorded under different conditions from the conversation itself.
The system also needs bounded latency and has to be cheap enough to run continuously for many simultaneous conversations.
The natural way to solve this is with speaker diarization.
We built that system too.
It turns out to be a strong baseline—and understanding where it breaks is the easiest way to understand why Isolate looks the way it does.
A natural solution: diarize, then verify
Suppose we have a recording that contains three people.
A speaker diarizer answers who spoke when, but it does not know their identities. It groups the recording into anonymous speaker tracks— all the time segments it believes belong to the same person. With three people, it might produce Speaker 0, Speaker 1, and Speaker 2.
We can then collect the audio belonging to each track, produce a speaker embedding, compare that embedding with the wearer's enrollment, using cosine similarity.
If Speaker 1 is sufficiently similar to the enrollment, we keep Speaker 1 and reject other tracks.
This is the diarize-then-verify baseline we compare against throughout the post. For the diarization stage, we evaluated several state-of-the-art commercial and open-source systems. For speaker verification, we tested three external models alongside our in-house speaker identity model. Unless noted otherwise, the results shown here use the best-performing identity model for the verification stage.
Figure 1 · Diarize, then verify. The diarizer turns the recording into anonymous speaker tracks. Audio from each track is pooled into a speaker embedding and compared with the wearer’s enrollment. The key consequence is that identity is decided once for an entire track: a speaker is either kept or rejected.
This is not a straw-man baseline. With a strong diarizer and a strong speaker-identity model, it works surprisingly well.
But the structure of the diarize-then-verify pipeline creates a few problems for our use case.
A speaker track is a very coarse unit of decision
Each speaker track gets one identity score.
If the score clears the cut-off, every frame assigned to the track is admitted. If it does not, every frame is rejected.
That works well when the diarization is exactly right. The problem is that any diarization error is inherited by the identity decision.
If two voices are merged into the same track, accepting the wearer means accepting the other voice too. Once the diarizer has made that decision there is nothing downstream that can split the track again.
Concurrent speech creates the same problem. Both people can be speaking during the same instant, while the identity decision exists at the level of the track.
The operating point moves in speaker-track-sized jumps
For us, the operating point matters.
We might decide, for example, that a small amount of fragmented leakage is tolerable— an isolated 80 ms frame or a clipped part of a word is very different from admitting a coherent phrase, sentence, or longer span from another speaker. Subject to that constraint, we want to preserve as much of the wearer’s speech as possible. The same is true in the other direction: if the model misses some of the wearer’s speech, losing a few scattered frames or part of a word may still preserve most important parts whereas dropping an entire utterance or speaker track removes that information altogether.
A model that assigns a score to every short interval gives us fine-grained control over that trade-off. Raising the threshold can reject a few uncertain moments while leaving the rest of the conversation unchanged.
Diarize-then-verify behaves differently. Each speaker track has one identity score, so changing the cut-off usually does nothing until it crosses the score of one of those tracks. When that happens, all the speech assigned to the track enters or leaves the output at once. Instead of giving up—or admitting—a few uncertain moments, the system may have to give up or admit an entire speaker track.
Which operating points exist, and where they fall, depends on who was in the room and how the diarizer split them.
A verification cosine cut-off does not necessarily travel
Speaker-verification scores depend on more than identity.
They can also shift with how much speech is available, the recording and acoustic conditions, and what the diarizer has grouped into each track. A clean ten-second track and a short or contaminated track from the same person need not receive the same score. And finally it also depends on how robust is the speaker verification model.
A cut-off that cleanly separates the wearer from everyone else in one conversation may sit much closer to the wrong speakers in another.
For a privacy-sensitive system, we cannot assume access to labeled data from every environment in which it will eventually run. The operating point therefore has to be chosen in advance and remain useful as the speakers, acoustics, and recording conditions change.
Later, we test this directly by choosing the threshold in one set of environments and measuring how well it transfers to others (see Figure 6, and the Threshold transferability and Precision–recall tabs of Figure 5).
Streaming changes the problem again
Many of the strongest diarizers can use context from an entire recording.
Our product cannot wait for conversation to end.
When we take a strong whole-recording diarizer and instead run it on bounded windows, some of their advantages disappears. On our hardest real conversation, for example, moving Nemotron 3 from the whole recording to 90-second windows reduces its fixed-cut-off F1 from 0.956 to 0.939, and recall at 95% precision from 0.967 to 0.926.
This highlights a fundamental difference in operating conditions: offline diarization can use long-range context, while our system has to make decisions continuously with bounded latency.
What we wanted instead was a model with four properties: Fine-grained decisions in time, rather than one decision for an entire speaker track. A score that lets us adjust the privacy–utility trade-off gradually. An enrollment representation that can be computed once and reused across conversations. And enough speaker-level context to say none of these people is the enrolled speaker.
That became Isolate.
Isolate: identity as a probability over time
Isolate is a frame-level classification model at the core of target-speaker extraction system. Every 80 ms, it produces a model-estimated probability that the enrolled person is speaking.
Around it, the target-speaker extraction system runs an on-device speech recognizer on the original conversation to generate the transcript and uses estimated probabilities to decide which recognized words to keep or discard.
That distinction is crucial. We are not trying to reconstruct a clean waveform containing only the wearer. The product ultimately consumes a transcript, so classification gives us exactly the control we need: a frame-level probability that can be thresholded at different operating points, and an error we can inspect directly—a word belonging to the wearer was missed, or a word belonging to somebody else was admitted.
The basic question is therefore simple:
For this 80 ms frame, what is the probability that the enrolled person is speaking?
Producing that probability reliably in a real conversation is the harder part.
Figure 2 · Isolate at a glance. The enrollment is encoded once and cached. Each block produces anonymous speaker activity and a frame-level target score; a speaker-level binding decision adds longer-range identity evidence before the two target scores are blended into one model-estimated probability per 80 ms frame.
There are two timescales in the model.
The first is local. Before the session begins, the enrollment recording is encoded once and cached. As conversation audio arrives, the model tracks anonymous speaker activity while simultaneously producing the frame representations used for target classification.
For identity, we also encode short windows of the conversation and compare them with the enrollment. These identity-evidence tokens tell the frame model not just that someone is speaking, but how much the nearby speech sounds like the enrolled person.
Figure 3 · How identity evidence enters the frame model. Each frame representation attends to identity-evidence tokens near it in time. The result is injected through a learned gated residual connection whose gate starts at zero, allowing the model to learn how much identity evidence to use.
This gives us a fine-grained target score, but it creates another problem.
Identity is much easier to establish from a ten-second explanation than from a 300 ms “yeah.” Short turns, overlapping speech, and similar voices can all be ambiguous when considered frame by frame.
So Isolate also reasons at a second, longer timescale.
Figure 4 · The Speaker Binding Module. Instead of deciding identity independently on every frame, the binder accumulates evidence for each tracked speaker across the block and asks which speaker, if any, is the enrolled person. That speaker-level decision is then placed back onto the frames and blended with the frame-level score.
For each anonymous speaker tracked through the block, the model pools that speaker's own frames into an identity representation and compares it with the enrollment. It then makes one decision over:
{speaker 1, speaker 2, … speaker K, none of them}.
We call this the Speaker Binding Module, or SBM.
The two decisions complement each other. The frame path determines where in time the target is active, with fine resolution around short turns, overlaps, and speaker boundaries. The binder contributes the longer-range evidence needed to determine which tracked voice actually belongs to the enrolled person—or whether the enrolled person is absent altogether.
The two scores are blended, passed through a sigmoid, resulting in a model-estimated probability comsumed by the rest of the system:
p(target) for every 80 ms frame.
What this looks like on a real conversation
The architecture above eventually reduces to something much simpler: a score moving through time.
The player below shows one real conversation as Isolate sees it. The top rows are the people actually speaking. The blue trace is the model's target-speaker score, and the horizontal line is the operating threshold. Frames above that threshold are treated as the wearer; frames below it are rejected.
Interactive example · Isolate running on a real conversation. The speaker tracks at the top are reference labels; the model does not receive those labels. The score trace is the model's output. Blue regions are wearer speech that survives, red regions are speech from somebody else that crossed the threshold, and missed wearer speech is shown separately.
A few things are easier to see here than in a benchmark table.
Most of the decision is not difficult: when the wearer speaks for several seconds, the score is comfortably on the correct side of the threshold. The interesting cases are the boundaries—short acknowledgements, interruptions, overlap, and moments where another voice briefly resembles the enrollment.
Changing the threshold exposes the actual product trade-off. Move it upward and fewer other-speaker fragments survive, but some wearer speech starts disappearing with them. Move it downward and wearer recall improves, at the cost of admitting more speech from the people around them.
That trade-off is what the rest of the evaluation measures across many conversations rather than one.
Training uses real and synthetic conversations extracted from more than 10,000 hours of publicly available audio and podcasts. The synthetic data includes adversarial combinations such as similar-sounding or otherwise confusable speakers. The model versions, scoring protocol and metric definitions are given in the Isolate Model Technical Report.
How we test it
There are many ways to make a target-speaker system look good accidentally.
Three evaluation rules matter more to us than any individual benchmark.
First, the enrollment is never taken from the test conversation. In every evaluation the model identifies the target from a separately recorded sample.
Second, we explicitly test the case where the enrollment belongs to somebody who is not present. That is not silence; other people are still talking. The correct output is simply nothing.
Third, we score the thing the product consumes. Wherever possible, that means recognized words rather than raw frames.
If a word belongs to the wearer, did we keep it?
If it belongs to somebody else, did we admit it?
Spans labelled mixed, where the ground truth says both people spoke but cannot identify which words belong to whom, are excluded rather than guessed.
Evaluation sets
| set | audio | ground truth |
|---|---|---|
| Study 1 | 178 min | human-labelled real conversations |
| Study 2 | 120 min | human-labelled real conversations |
| Study 3 | 2.4 min | human-labelled real conversation |
| Study 4 | 15.9 min | human-labelled real conversation |
| Enacted retail | 60.5 min | 16 consented retail and restaurant scenes, frame-level truth |
| Synthetic mixtures | 600 min | 2,400 two-speaker mixtures with exact speaker timelines |
Study 4 is particularly useful because it contains a second talker whose voice is similar enough to the target to expose weaknesses that barely appear on the easier conversations.
The baselines cover nine diarize/transcribe configurations: Sortformer v1, Streaming Sortformer v2, Nemotron 3 Diarization, DiariZen WavLM-Large s80 v2, VibeVoice-ASR 7B, ElevenLabs Scribe v1 with the speaker count fixed at two, Scribe v1 with the speaker count unconstrained, Scribe v2, and Meta Muse Voice Transcribe.
For the diarizer baselines, anonymous speakers are mapped to the enrollment using the same basic procedure from Figure 1: pool a track, embed it, compare it with the enrollment, and apply a cosine cut-off.
Now we can ask the question we actually care about.
Does the same operating point work when the wearer, the place and the recording conditions change?
One threshold, from synthetic mixtures to the shop floor
A benchmark is easy to overfit without training a single parameter.
Choose a different threshold for every evaluation set and almost any reasonably ranked classifier looks better.
A deployed product does not get to do that.
We want one operating point that remains useful when the conversation changes.
At one fixed threshold across all six sets, Isolate 2.0 + SBM-1 reaches a mean F1 of 0.958, compared with 0.925 for the strongest diarize-then-verify baseline.
Isolate 2.0 without the binder is marginally higher on this particular aggregate, at 0.961. That is useful context: SBM is not there to win a mean-F1 benchmark. Its main contribution is the speaker-level identity behavior we will test in the next section.
At the high-precision end of the curve—the region that matters most to the product—the advantage becomes easier to interpret.
At a 5% mis-attribution budget, Isolate 2.0 retains 96.8% of the wearer's speech, compared with 91.9% for the strongest baseline (span-level units; on recognized words, 95.8% against 80.6%).
The difference comes from the geometry of the decision.
With a frame score, the threshold can give up a small amount of uncertain speech while keeping the rest of the speaker.
With a speaker-track score, crossing one threshold can add or remove an entire voice.
But the more important result is not the best point on the curve.
It is what happens when the threshold is chosen somewhere else.
Figure 6 · Does a threshold picked elsewhere hold? In one view, each set is scored at a threshold selected using the other five. In the other, the threshold is selected on one set and then applied to the remaining five. The experiment asks whether calibration survives a change of set: other wearers, places and recording conditions.
For Isolate, the useful region of the curve is broad.
The F1 optimum is effectively a plateau rather than a spike. For Isolate 2.0 it spans roughly 0.60–0.85.
That means an imperfectly chosen threshold can still land in a good part of the curve.
Even the worst case—choosing the threshold from a single evaluation environment and applying it to the others—still gives 0.932 mean F1 elsewhere.
The strongest diarize-then-verify systems can fall to roughly 0.65 and 0.56 under the same kind of bad calibration.
That, more than the peak benchmark number, is the result we care about.
A production model needs an operating point that travels.
This does not mean that threshold 0.725, or any other number in this report, is universally optimal. In production the threshold is chosen from the privacy–utility budget appropriate to the deployment.
The important property is that the model does not sit on a narrow cliff where a modest shift in threshold or acoustic conditions suddenly changes its behavior.
What happens when the wearer isn't there?
Precision and recall still assume that the target speaker appears somewhere in the conversation.
That leaves a much harder privacy question.
What if they do not?
We test that directly.
Figure 7 · The wrong-enrollment protocol. Every recording is run once with the correct enrollment and again with an enrollment belonging to somebody who is not in the recording. In the second run, every emitted frame is a false accept.
This experiment matters because the input is not silence.
People are talking. They simply are not the enrolled person.
We measure the absent-speaker false-accept rate (FAR): the fraction of frames emitted under a wrong enrollment.
At threshold 0.5:
| system | Study 1 | Study 2 | Study 3 | Study 4 | Enacted retail |
|---|---|---|---|---|---|
| Isolate 1.1 | 5.89% | 7.77% | 8.25% | 5.33% | 4.14% |
| Isolate 1.1 + SBM-1 | 3.68% | 4.29% | 3.51% | 1.74% | 3.01% |
| Isolate 2.0 | 3.96% | 4.03% | 4.82% | 3.63% | 3.11% |
| Isolate 2.0 + SBM-1 | 2.85% | 2.75% | 3.35% | 1.56% | 2.64% |
At each model's operating point, the lowest threshold that holds precision at or above 95% on every set (the rule of the Threshold transferability tab in Figure 5), so the threshold differs per model:
| system | threshold | Study 1 | Study 2 | Study 3 | Study 4 | Enacted retail |
|---|---|---|---|---|---|---|
| Isolate 1.1 | 0.800 | 1.95% | 2.63% | 1.35% | 1.01% | 1.20% |
| Isolate 1.1 + SBM-1 | 0.775 | 1.28% | 1.28% | 0.95% | 0.35% | 0.76% |
| Isolate 2.0 | 0.875 | 1.05% | 1.06% | 0.88% | 0.45% | 0.87% |
| Isolate 2.0 + SBM-1 | 0.825 | 0.97% | 0.89% | 0.60% | 0.26% | 0.87% |
At its operating point every model emits under 3% of frames under a wrong enrollment, and Isolate 2.0 + SBM-1 under 1% on every set. The binder's advantage looks different here: Isolate 2.0 + SBM-1 reaches its 95% floor at 0.825 instead of 0.875 and keeps the same 94% of the wearer's speech on average, with fewer false accepts on four of the five sets.
Read together with the table above, this also says something about the generations: the Speaker Binding Module is not as crucial for Isolate 2.0, whose stronger identity evidence and trunk already keep false accepts low at the operating point, as it was for generation 1, where it cut the rate by up to 3×.
The generation-1 numbers look good in this table only because their operating points are conservative. At those thresholds Isolate 1.1 keeps about 87% of the wearer's speech on average and Isolate 1.1 + SBM-1 about 90%, against 94% for both generation-2 models. Figure 8 lets you move the threshold and see the whole recall–false-accept trade-off.
The important comparison is within each generation.
Adding the speaker-level binding decision reduces false accepts on every evaluation set.
On generation 1, the reduction ranges from roughly 1.4× to 3.1×. Generation 2 shows the same direction.
That is the behavior SBM was designed to add: not simply “does this 80 ms fragment resemble the enrollment?” but “among the voices I have been tracking, is any one of them actually this person?”
It helps substantially especially for less conservative threshold (0.5). We can see the privacy–utility trade-off directly by sweeping Isolate's threshold.
On the enacted-retail set at threshold 0.5, Isolate 1.0 sits at roughly 8.0% false accept / 96.8% recall.
Isolate 2.0 + SBM-1 is at roughly 2.6% false accept / 97.8% recall.
So the improvement did not come from simply becoming more conservative.
The curve itself moved. At more conservative thresholds (precision at or above 95% and FAR of about 1% or less), Isolate 2.0 and Isolate 2.0 + SBM-1 both offer about 94% recall on average.
That is the deployment decision we actually want: choose the amount of third-party leakage the system is allowed to tolerate, then read the corresponding wearer recall from the curve.
The remaining false accepts could be resolved with a single conversation-level decision rather than a frame-by-frame one: is the enrolled person in this conversation at all? That is a decision the binder can make over a whole turn or conversation, and it is the first of the design changes proposed in the next steps below.
Where the remaining errors are
A single precision or F1 number hides what the transcript actually looks like.
Our downstream product consumes words and utterances, so we also measure word attribution error:
(missed target words + admitted non-target words) / target words
The two components matter separately.
A miss removes something the wearer said.
A leak adds something another person said.
For the current model at threshold 0.5:
| set | attribution error | missed target | leaked non-target |
|---|---|---|---|
| Study 1 | 0.077 | 0.018 | 0.059 |
| Study 2 | 0.065 | 0.029 | 0.036 |
| Study 3 | 0.078 | 0.004 | 0.074 |
| Study 4 | 0.138 | 0.037 | 0.102 |
| Enacted retail | 0.029 | 0.015 | 0.014 |
| Synthetic mixtures | 0.057 | 0.026 | 0.031 |
And at its operating point, 0.825, the threshold from the false-accept table in the previous section that holds 95% precision on every set:
| set | attribution error | missed target | leaked non-target |
|---|---|---|---|
| Study 1 | 0.108 | 0.064 | 0.045 |
| Study 2 | 0.098 | 0.079 | 0.019 |
| Study 3 | 0.048 | 0.015 | 0.033 |
| Study 4 | 0.160 | 0.102 | 0.058 |
| Enacted retail | 0.031 | 0.025 | 0.006 |
| Synthetic mixtures | 0.071 | 0.053 | 0.018 |
Moving from 0.5 to the operating point flips the composition of the error: leaks fall on every set, misses rise, and on most sets the total goes up. That is the price of the precision floor, and it is what the product sees.
The hardest real conversation, Study 4, shows the trade-off.
At threshold 0.5 the model is not primarily failing by throwing away the wearer's speech: it misses about 3.7% of target words but admits non-target words equal to about 10.2% of them. At the operating point the balance reverses, about 10.2% missed against 5.8% leaked.
The remaining error is an attribution-boundary problem either way: which side of a turn boundary a word lands on, not a broad failure to hear the wearer.
Short utterances remain the weak end.
At the operating point, Isolate 2.0 + SBM-1 preserves about 57% of two-word target utterances, compared with roughly 95% of target utterances containing seven words or more, while admitting under 4% of long non-target utterances.
That is not surprising acoustically. Short turns provide less identity evidence, and their boundaries occupy a much larger fraction of the utterance.
It also explains why a span-level metric is more informative than a frame-level metric and complementary to word-level metric. A second long “yeah sure” and a ten-second recommendation are both speech, but they are not equally informative to the application and they do not fail for the same reasons.
Post-hoc filtering strategies
Once the model is working reasonably well, we also use a set of additional post-hoc filter to clean up some of the fractional leakage. We present one such filter using two off-the shelf open source diarizers:
Take the audio Isolate already decided to emit. Diarize that smaller stream. Find the dominant speaker inside it. Remove frames belonging only to the other speakers.
The intuition is straightforward: if Isolate is mostly correct, then the target should dominate its output and the mistakes should appear as smaller fragments belonging to other voices.
We ran this experiment on Isolate 1.1 + SBM-1, using Sortformer v1 and Nemotron 3 as the cleanup diarizer.
A reduced view of the result:
| set | system | recall @ 0.5 | attribution error @ 0.5 |
|---|---|---|---|
| Study 1 | unfiltered | 0.946 | 0.184 |
| + Sortformer v1 | 0.946 | 0.133 | |
| + Nemotron 3 | 0.935 | 0.140 | |
| Study 2 | unfiltered | 0.988 | 0.084 |
| + Sortformer v1 | 0.961 | 0.111 | |
| + Nemotron 3 | 0.925 | 0.116 | |
| Study 4 | unfiltered | 0.939 | 0.135 |
| + Sortformer v1 | 0.917 | 0.153 | |
| + Nemotron 3 | 0.931 | 0.124 | |
| Enacted retail | unfiltered | 0.980 | 0.064 |
| + Sortformer v1 | 0.970 | 0.070 | |
| + Nemotron 3 | 0.966 | 0.069 |
The filter does exactly what its mechanism suggests.
When fragments from several voices leak into the accepted audio, it can remove some of them. On Study 1, attribution error falls from 0.184 to 0.133 with Sortformer and 0.140 with Nemotron. On Study 4, Nemotron reduces it from 0.135 to 0.124.
Which diarizer helps depends on the conversation: Sortformer does better on Study 1, Nemotron on Study 4, and on Study 2 both make things worse, because the wearer is not the dominant voice in what was emitted and the filter removes the wearer instead. The filter also does not touch the privacy problem: under a wrong enrollment Isolate usually emits one impostor voice, the cleanup diarizer declares that voice the dominant speaker and keeps it, so the absent-speaker false-accept rate barely moves (Study 4: 1.74% to 1.31% with Nemotron; elsewhere within 0.4 points). Fragmented leaks carry little identity information, so a diarizer that needs long stretches of audio to build confidence cannot be relied on to filter them.
In other words, post-hoc filters can be useful to reduce leakage but a better solution requires stronger identity evidence to be incorporated before or during model's decision-making.
What we learned
The main lesson from this work is that the two natural timescales of the problem are both useful.
Frame-level decisions give us the resolution to handle overlap, avoid admitting an entire incorrectly merged speaker track, and move smoothly along a precision–recall curve.
They are also what make the operating point transferable. The useful region of Isolate's curve is broad enough that a threshold chosen in one environment generally remains useful in another.
But speaker-level identity solves a different problem.
A short acoustic fragment can look like the enrolled speaker by accident. When the enrolled person is absent entirely, enough of those fragments occasionally cross the frame threshold to produce the false accepts we still see.
The Speaker Binding Module substantially reduces that behavior by asking a longer-range question: which tracked voice is actually the wearer—or is the answer none of them?
Next steps
Two experiments are cheap because they use outputs we already have. One pools the frame probabilities inside each recognised word and decides once per word, so a threshold crossing can no longer split a word at its boundary. The other measures whether short isolated runs in the output can be rejected without losing genuine short turns; most absent-speaker errors are bursts of a few seconds. Both trim errors at the edges.
The bigger next steps are as follows:
Make one global decision: is the wearer in this conversation at all? Today the model answers that question frame by frame, so when the wearer is absent every frame has to be rejected on its own and a few slip through. The binder already makes a per-speaker call for every block, including "none of these speakers is the wearer". Pooling that call over a whole turn or conversation gives a single yes-or-no answer with far more evidence behind it. If the answer is no, nothing is emitted; if yes, the frame-level score takes over and decides where within the conversation the wearer speaks.
Decide the timing now and the identity a little later. Today the frame score and the binder's decision are fixed inside each block. Instead, the model could emit frames provisionally, let the binder re-label the whole turn once it has heard enough of that speaker, and remember each speaker's identity and verdict from one block to the next. A speaker identified correctly would stay identified, which is the fix for the block-level flips we still see on the hardest conversation.
Train for the threshold we actually ship. The model is trained on a per-frame loss, which does not know about the operating threshold or the absent-speaker rate. Adding both to the training objective would keep the best threshold from drifting between model versions and would make false accepts something the model is trained to avoid, not just measured on.
Close the gap that synthetic data cannot cover. Our training mixtures simulate overlap, pauses, noise, loudness, on a different microphone, through a different codec, and similar voices, but they cannot simulate the same person a month later, with a cold, and other sources of intra-person variability. The likely fixes are incremental enrollment, where the cached enrollment is refreshed with high-confidence wearer speech from each session.
Furthermore, the evaluation points at three more directions. The transcript itself carries evidence of who is speaking that the audio path never sees, such as a question followed by an answer or a one-word backchannel, and an LLM-based re-scorer over the words can use it. The 80 ms frame is coarse at word boundaries, a small upsampling network before the classification head would place boundaries at 10 ms at almost no added cost, and the same kind of change can extend the number of simultaneous talkers the model tracks. And the streaming client still works in fixed blocks, where the block edge is the place identity is lost; a rolling state with a predicted end of turn removes that edge and gives the presence decision a principled moment to fire.
The goal remains the same: keep the temporal resolution that makes overlapping real-world conversations tractable, while giving identity enough context to know when the right answer is nobody.
That is the balance we need if real-world conversations are going to be useful data without turning every person within microphone range into part of the dataset.
Detailed results
The full generation-by-generation tables, span-level evaluation, bootstrap intervals, metric definitions, scoring protocol, and baseline details are available in the accompanying Isolate Model Technical Report, along with the references for the baseline systems technical report.