Open-weights speaker diarization for pyannote.audio 4.x: WavLM-Base+ / Conformer segmentation + SimAM-ResNet34 embeddings, by Reway.
Python
23
0 commits
updated Sep 10, 2026

SUP as in "what's up": who is speaking, and when.
The most accurate open-weights speaker diarization we know of, on both public suites:
15.86 macro DER on pyannote's 8-corpus benchmark — ahead of pyannoteAI's commercial
precision-2 (16.06) — and 20.94 across all 12 corpora. From the research team at
Re:WayAI, packaged for
pyannote.audio 4.x.
Two models on one recipe: SUPlime on a WavLM-Base+ backbone (114 M params) and SUPlime-L on WavLM-Large (349 M). SUPlime-L is the better model on meeting and conversational audio; the base model wins on far-field and dinner-party recordings (CHiME-6, DiPCo, NOTSOFAR-1). Per-corpus numbers for all 12 corpora are in Results below.
pip install suplime # pulls pyannote.audio>=4.0.7,<5
pip install does not bring FFmpeg, and pyannote.audio reads audio files through
torchcodec, which loads FFmpeg's shared libraries at runtime — without them the first call
raises RuntimeError: Could not load libtorchcodec. Install it into the same environment:
micromamba install -c conda-forge ffmpeg # or conda/mamba; on Linux also apt install ffmpeg
Passing a waveform instead of a path needs no decoder at all:
import soundfile as sf, torch
x, sr = sf.read("meeting.wav", dtype="float32", always_2d=True)
output = pipeline({"waveform": torch.from_numpy(x.T), "sample_rate": sr}) # (channel, time)
import torch
from pyannote.audio import Pipeline
pipeline = Pipeline.from_pretrained("rewayai/suplime").to(torch.device("cuda"))
# the WavLM-Large variant is a drop-in replacement, same API and same threshold:
# pipeline = Pipeline.from_pretrained("rewayai/suplime-large").to(torch.device("cuda"))
output = pipeline("meeting.wav")
for turn, _, speaker in output.speaker_diarization.itertracks(yield_label=True):
print(f"{turn.start:.2f} {turn.end:.2f} {speaker}")
No Hugging Face token is needed: the repository is not gated.
Or build the pipeline in Python, without the repo's config.yaml — useful for pointing
at local checkpoints or setting the parameters yourself:
from suplime import SuplimeDiarization
pipeline = SuplimeDiarization() # defaults = rewayai/suplime
pipeline.instantiate(pipeline.default_parameters()) # threshold 0.72, min_cluster_size 12
L = "rewayai/suplime-large" # the same, for SUPlime-L
pipeline = SuplimeDiarization(segmentation={"checkpoint": L, "subfolder": "segmentation"},
embedding={"checkpoint": L, "subfolder": "embedding"})
Individual models load through the usual API:
from pyannote.audio import Model
seg = Model.from_pretrained("rewayai/suplime", subfolder="segmentation")
emb = Model.from_pretrained("rewayai/suplime", subfolder="embedding")
Both repositories have the same layout, so every call above works with
rewayai/suplime-large; the embedding model is byte-identical between them.
Set SUPLIME_FP16=0 to run the WavLM backbone in fp32 (default: fp16 autocast
on CUDA during inference — the published numbers were produced this way).
Diarization error rate (%), lower is better:
| Corpus (test unless noted) | SUPlime | SUPlime-L | DiariZen-L-s80-v2 | pyannoteAI precision-2 | pyannote community-1 |
|---|---|---|---|---|---|
| AISHELL-4 | 11.55 | 11.21 | 10.1 | 11.4 | 11.7 |
| AliMeeting (far, ch1) | 14.45 | 14.55 | 10.8 | 15.2 | 20.3 |
| AMI (IHM, Mix-Headset) | 12.59 | 11.91 | 25.69 † | 12.9 | 17.0 |
| AMI (SDM) | 15.14 | 14.85 | 13.9 | 15.6 | 19.9 |
| AVA-AVD | 38.09 | 36.83 | 42.62 † | 37.1 | 44.6 |
| MSDWild (few.val) | 17.81 | 17.48 | 15.8 | 17.3 | 22.8 |
| RAMC | 10.86 | 11.10 | 11.0 | 10.5 | 20.8 |
| VoxConverse (v0.3) | 9.21 | 8.93 | 9.1 | 8.5 | 11.2 |
| macro average (8) | 16.21 | 15.86 | 17.38 | 16.06 | 21.04 |
| NOTSOFAR-1 (80-session split) | 20.04 | 22.70 | 18.86 † | — | 27.67 † |
| ICSI | 22.97 | 22.77 | 26.54 † | — | 30.84 † |
| CHiME-6 | 48.11 | 48.98 | 48.39 † | — | 51.98 † |
| DiPCo | 30.44 | 30.92 | 37.56 † | — | 34.38 † |
| macro average (12) | 20.94 | 21.02 | 22.53 | — | 26.10 |
SUPlime-L leads the 8-corpus benchmark at 15.86 macro DER — ahead of precision-2
(16.06), the base model (16.21) and DiariZen-L-s80-v2 (17.38) — while across all 12 corpora
the base model takes it back by 0.06, the four extra sets being far-field and dinner-party
audio where the larger backbone does not pay off. Both beat DiariZen by ~1.5 DER on the
12-corpus average, with the largest margins on close-talk AMI and on CHiME-6 / DiPCo, while
DiariZen stays clearly better on the Mandarin meeting corpora and MSDWild. Per-corpus
hypothesis RTTMs ship with each model; the cards cover threshold behaviour and known
limitations.
Scoring conditions. Collar 0 s, overlapped speech scored, no oracle speaker count, reference cropped to each corpus' UEM, one clustering threshold (0.72) everywhere. The other systems' columns are their authors' published numbers (DiariZen, pyannote), except where † marks our own re-run of their open-source pipeline, unchanged and on the same files — and for NOTSOFAR-1 the DiariZen authors report 16.7 on a different session split.
voxblink2_samresnet34_ft, 25 M params, 256-dim). Shared, byte for byte.min_cluster_size 12,
overlap excluded from embeddings. Shared.The full training pipeline is released as well — corpus preparation, the augmentation setup,
the training command and the checkpoint soup — so the models can be reproduced from scratch
or trained further on your own data. Recipe and commands for both variants:
training/README.md.
Code: MIT (src/suplime/models/samresnet.py: Apache-2.0, ported from
WeSpeaker). Model weights: CC BY-NC 4.0.
Non-commercial because several training corpora (RAMC, MSDWild, AVA-AVD) are
licensed for research use only and others (AISHELL-4, AliMeeting) are share-alike;
the most restrictive terms apply to the derived model. Details in the model card. Built on pyannote.audio (MIT), WavLM (MIT) and WeSpeaker (Apache-2.0).
SUPlime tells you who spoke and when. If you also want to know what they said, and who they actually are, that is our day job: Re:WayAI is the speech API that knows who's talking.
SPEAKER_03Pay-as-you-go, 50 free hours to start, no subscription. The lime approves.
46 followers · starred Sep 2026
Python
96.3%
Shell
3.7%
Open-weights speaker diarization for pyannote.audio 4.x: WavLM-Base+ / Conformer segmentation + SimAM-ResNet34 embeddings, by Reway.
Python
23
0 commits
updated Sep 10, 2026

SUP as in "what's up": who is speaking, and when.
The most accurate open-weights speaker diarization we know of, on both public suites:
15.86 macro DER on pyannote's 8-corpus benchmark — ahead of pyannoteAI's commercial
precision-2 (16.06) — and 20.94 across all 12 corpora. From the research team at
Re:WayAI, packaged for
pyannote.audio 4.x.
Two models on one recipe: SUPlime on a WavLM-Base+ backbone (114 M params) and SUPlime-L on WavLM-Large (349 M). SUPlime-L is the better model on meeting and conversational audio; the base model wins on far-field and dinner-party recordings (CHiME-6, DiPCo, NOTSOFAR-1). Per-corpus numbers for all 12 corpora are in Results below.
pip install suplime # pulls pyannote.audio>=4.0.7,<5
pip install does not bring FFmpeg, and pyannote.audio reads audio files through
torchcodec, which loads FFmpeg's shared libraries at runtime — without them the first call
raises RuntimeError: Could not load libtorchcodec. Install it into the same environment:
micromamba install -c conda-forge ffmpeg # or conda/mamba; on Linux also apt install ffmpeg
Passing a waveform instead of a path needs no decoder at all:
import soundfile as sf, torch
x, sr = sf.read("meeting.wav", dtype="float32", always_2d=True)
output = pipeline({"waveform": torch.from_numpy(x.T), "sample_rate": sr}) # (channel, time)
import torch
from pyannote.audio import Pipeline
pipeline = Pipeline.from_pretrained("rewayai/suplime").to(torch.device("cuda"))
# the WavLM-Large variant is a drop-in replacement, same API and same threshold:
# pipeline = Pipeline.from_pretrained("rewayai/suplime-large").to(torch.device("cuda"))
output = pipeline("meeting.wav")
for turn, _, speaker in output.speaker_diarization.itertracks(yield_label=True):
print(f"{turn.start:.2f} {turn.end:.2f} {speaker}")
No Hugging Face token is needed: the repository is not gated.
Or build the pipeline in Python, without the repo's config.yaml — useful for pointing
at local checkpoints or setting the parameters yourself:
from suplime import SuplimeDiarization
pipeline = SuplimeDiarization() # defaults = rewayai/suplime
pipeline.instantiate(pipeline.default_parameters()) # threshold 0.72, min_cluster_size 12
L = "rewayai/suplime-large" # the same, for SUPlime-L
pipeline = SuplimeDiarization(segmentation={"checkpoint": L, "subfolder": "segmentation"},
embedding={"checkpoint": L, "subfolder": "embedding"})
Individual models load through the usual API:
from pyannote.audio import Model
seg = Model.from_pretrained("rewayai/suplime", subfolder="segmentation")
emb = Model.from_pretrained("rewayai/suplime", subfolder="embedding")
Both repositories have the same layout, so every call above works with
rewayai/suplime-large; the embedding model is byte-identical between them.
Set SUPLIME_FP16=0 to run the WavLM backbone in fp32 (default: fp16 autocast
on CUDA during inference — the published numbers were produced this way).
Diarization error rate (%), lower is better:
| Corpus (test unless noted) | SUPlime | SUPlime-L | DiariZen-L-s80-v2 | pyannoteAI precision-2 | pyannote community-1 |
|---|---|---|---|---|---|
| AISHELL-4 | 11.55 | 11.21 | 10.1 | 11.4 | 11.7 |
| AliMeeting (far, ch1) | 14.45 | 14.55 | 10.8 | 15.2 | 20.3 |
| AMI (IHM, Mix-Headset) | 12.59 | 11.91 | 25.69 † | 12.9 | 17.0 |
| AMI (SDM) | 15.14 | 14.85 | 13.9 | 15.6 | 19.9 |
| AVA-AVD | 38.09 | 36.83 | 42.62 † | 37.1 | 44.6 |
| MSDWild (few.val) | 17.81 | 17.48 | 15.8 | 17.3 | 22.8 |
| RAMC | 10.86 | 11.10 | 11.0 | 10.5 | 20.8 |
| VoxConverse (v0.3) | 9.21 | 8.93 | 9.1 | 8.5 | 11.2 |
| macro average (8) | 16.21 | 15.86 | 17.38 | 16.06 | 21.04 |
| NOTSOFAR-1 (80-session split) | 20.04 | 22.70 | 18.86 † | — | 27.67 † |
| ICSI | 22.97 | 22.77 | 26.54 † | — | 30.84 † |
| CHiME-6 | 48.11 | 48.98 | 48.39 † | — | 51.98 † |
| DiPCo | 30.44 | 30.92 | 37.56 † | — | 34.38 † |
| macro average (12) | 20.94 | 21.02 | 22.53 | — | 26.10 |
SUPlime-L leads the 8-corpus benchmark at 15.86 macro DER — ahead of precision-2
(16.06), the base model (16.21) and DiariZen-L-s80-v2 (17.38) — while across all 12 corpora
the base model takes it back by 0.06, the four extra sets being far-field and dinner-party
audio where the larger backbone does not pay off. Both beat DiariZen by ~1.5 DER on the
12-corpus average, with the largest margins on close-talk AMI and on CHiME-6 / DiPCo, while
DiariZen stays clearly better on the Mandarin meeting corpora and MSDWild. Per-corpus
hypothesis RTTMs ship with each model; the cards cover threshold behaviour and known
limitations.
Scoring conditions. Collar 0 s, overlapped speech scored, no oracle speaker count, reference cropped to each corpus' UEM, one clustering threshold (0.72) everywhere. The other systems' columns are their authors' published numbers (DiariZen, pyannote), except where † marks our own re-run of their open-source pipeline, unchanged and on the same files — and for NOTSOFAR-1 the DiariZen authors report 16.7 on a different session split.
voxblink2_samresnet34_ft, 25 M params, 256-dim). Shared, byte for byte.min_cluster_size 12,
overlap excluded from embeddings. Shared.The full training pipeline is released as well — corpus preparation, the augmentation setup,
the training command and the checkpoint soup — so the models can be reproduced from scratch
or trained further on your own data. Recipe and commands for both variants:
training/README.md.
Code: MIT (src/suplime/models/samresnet.py: Apache-2.0, ported from
WeSpeaker). Model weights: CC BY-NC 4.0.
Non-commercial because several training corpora (RAMC, MSDWild, AVA-AVD) are
licensed for research use only and others (AISHELL-4, AliMeeting) are share-alike;
the most restrictive terms apply to the derived model. Details in the model card. Built on pyannote.audio (MIT), WavLM (MIT) and WeSpeaker (Apache-2.0).
SUPlime tells you who spoke and when. If you also want to know what they said, and who they actually are, that is our day job: Re:WayAI is the speech API that knows who's talking.
SPEAKER_03Pay-as-you-go, 50 free hours to start, no subscription. The lime approves.
46 followers · starred Sep 2026
Python
96.3%
Shell
3.7%