rewayai/suplime

Open-weights speaker diarization for pyannote.audio 4.x: WavLM-Base+ / Conformer segmentation + SimAM-ResNet34 embeddings, by Reway.

Python

23

0 commits

updated Sep 10, 2026

See the code

README

SUPlime

SUPlime speaker diarization

SUP as in "what's up": who is speaking, and when.

The most accurate open-weights speaker diarization we know of, on both public suites: 15.86 macro DER on pyannote's 8-corpus benchmark — ahead of pyannoteAI's commercial precision-2 (16.06) — and 20.94 across all 12 corpora. From the research team at Re:WayAI, packaged for pyannote.audio 4.x.

Macro-average DER of SUPlime, SUPlime-L, DiariZen-L-s80-v2, pyannoteAI precision-2 and pyannote community-1 on the 8-corpus benchmark and on all 12 corpora

Two models on one recipe: SUPlime on a WavLM-Base+ backbone (114 M params) and SUPlime-L on WavLM-Large (349 M). SUPlime-L is the better model on meeting and conversational audio; the base model wins on far-field and dinner-party recordings (CHiME-6, DiPCo, NOTSOFAR-1). Per-corpus numbers for all 12 corpora are in Results below.

Install

pip install suplime            # pulls pyannote.audio>=4.0.7,<5

pip install does not bring FFmpeg, and pyannote.audio reads audio files through torchcodec, which loads FFmpeg's shared libraries at runtime — without them the first call raises RuntimeError: Could not load libtorchcodec. Install it into the same environment:

micromamba install -c conda-forge ffmpeg     # or conda/mamba; on Linux also apt install ffmpeg

Passing a waveform instead of a path needs no decoder at all:

import soundfile as sf, torch
x, sr = sf.read("meeting.wav", dtype="float32", always_2d=True)
output = pipeline({"waveform": torch.from_numpy(x.T), "sample_rate": sr})   # (channel, time)

Use

import torch
from pyannote.audio import Pipeline

pipeline = Pipeline.from_pretrained("rewayai/suplime").to(torch.device("cuda"))
# the WavLM-Large variant is a drop-in replacement, same API and same threshold:
# pipeline = Pipeline.from_pretrained("rewayai/suplime-large").to(torch.device("cuda"))
output = pipeline("meeting.wav")
for turn, _, speaker in output.speaker_diarization.itertracks(yield_label=True):
    print(f"{turn.start:.2f} {turn.end:.2f} {speaker}")

No Hugging Face token is needed: the repository is not gated.

Or build the pipeline in Python, without the repo's config.yaml — useful for pointing at local checkpoints or setting the parameters yourself:

from suplime import SuplimeDiarization
pipeline = SuplimeDiarization()                       # defaults = rewayai/suplime
pipeline.instantiate(pipeline.default_parameters())   # threshold 0.72, min_cluster_size 12

L = "rewayai/suplime-large"                           # the same, for SUPlime-L
pipeline = SuplimeDiarization(segmentation={"checkpoint": L, "subfolder": "segmentation"},
                              embedding={"checkpoint": L, "subfolder": "embedding"})

Individual models load through the usual API:

from pyannote.audio import Model
seg = Model.from_pretrained("rewayai/suplime", subfolder="segmentation")
emb = Model.from_pretrained("rewayai/suplime", subfolder="embedding")

Both repositories have the same layout, so every call above works with rewayai/suplime-large; the embedding model is byte-identical between them.

Set SUPLIME_FP16=0 to run the WavLM backbone in fp32 (default: fp16 autocast on CUDA during inference — the published numbers were produced this way).

Results

Diarization error rate (%), lower is better:

Corpus (test unless noted)SUPlimeSUPlime-LDiariZen-L-s80-v2pyannoteAI precision-2pyannote community-1
AISHELL-411.5511.2110.111.411.7
AliMeeting (far, ch1)14.4514.5510.815.220.3
AMI (IHM, Mix-Headset)12.5911.9125.69 †12.917.0
AMI (SDM)15.1414.8513.915.619.9
AVA-AVD38.0936.8342.62 †37.144.6
MSDWild (few.val)17.8117.4815.817.322.8
RAMC10.8611.1011.010.520.8
VoxConverse (v0.3)9.218.939.18.511.2
macro average (8)16.2115.8617.3816.0621.04
NOTSOFAR-1 (80-session split)20.0422.7018.86 †—27.67 †
ICSI22.9722.7726.54 †—30.84 †
CHiME-648.1148.9848.39 †—51.98 †
DiPCo30.4430.9237.56 †—34.38 †
macro average (12)20.9421.0222.53—26.10

SUPlime-L leads the 8-corpus benchmark at 15.86 macro DER — ahead of precision-2 (16.06), the base model (16.21) and DiariZen-L-s80-v2 (17.38) — while across all 12 corpora the base model takes it back by 0.06, the four extra sets being far-field and dinner-party audio where the larger backbone does not pay off. Both beat DiariZen by ~1.5 DER on the 12-corpus average, with the largest margins on close-talk AMI and on CHiME-6 / DiPCo, while DiariZen stays clearly better on the Mandarin meeting corpora and MSDWild. Per-corpus hypothesis RTTMs ship with each model; the cards cover threshold behaviour and known limitations.

Scoring conditions. Collar 0 s, overlapped speech scored, no oracle speaker count, reference cropped to each corpus' UEM, one clustering threshold (0.72) everywhere. The other systems' columns are their authors' published numbers (DiariZen, pyannote), except where † marks our own re-run of their open-source pipeline, unchanged and on the same files — and for NOTSOFAR-1 the DiariZen authors report 16.7 on a different session split.

Architecture

  • Segmentation: WavLM (learned mixture of all layers) → 4-layer Conformer → powerset classifier (4 speakers / 10 s window, up to 2 simultaneous). WavLM-Base+, 114 M params for SUPlime; WavLM-Large, 349 M for SUPlime-L. This is the only difference between them.
  • Embedding: WeSpeaker SimAM-ResNet34 with attentive statistics pooling (voxblink2_samresnet34_ft, 25 M params, 256-dim). Shared, byte for byte.
  • Clustering: agglomerative (centroid linkage), threshold 0.72, min_cluster_size 12, overlap excluded from embeddings. Shared.

Training

The full training pipeline is released as well — corpus preparation, the augmentation setup, the training command and the checkpoint soup — so the models can be reproduced from scratch or trained further on your own data. Recipe and commands for both variants: training/README.md.

License

Code: MIT (src/suplime/models/samresnet.py: Apache-2.0, ported from WeSpeaker). Model weights: CC BY-NC 4.0. Non-commercial because several training corpora (RAMC, MSDWild, AVA-AVD) are licensed for research use only and others (AISHELL-4, AliMeeting) are share-alike; the most restrictive terms apply to the derived model. Details in the model card. Built on pyannote.audio (MIT), WavLM (MIT) and WeSpeaker (Apache-2.0).

Diarization is not enough?

SUPlime tells you who spoke and when. If you also want to know what they said, and who they actually are, that is our day job: Re:WayAI is the speech API that knows who's talking.

  • Transcription in 25 European languages (40 for real-time streaming), with overlap-robust diarization from the same team that built SUPlime
  • Voice enrollment: speakers get their real names, not SPEAKER_03
  • Real-time streaming over WebSocket, transcript Q&A, one-click DOCX / PDF exports
  • On-premise deployment, audio deleted right after processing, ISO 27001

Pay-as-you-go, 50 free hours to start, no subscription. The lime approves.

Significant stargazers

Michael Pedersen

46 followers · starred Sep 2026

rewayai/suplime

Open-weights speaker diarization for pyannote.audio 4.x: WavLM-Base+ / Conformer segmentation + SimAM-ResNet34 embeddings, by Reway.

Python

23

0 commits

updated Sep 10, 2026

See the code

README

SUPlime

SUPlime speaker diarization

SUP as in "what's up": who is speaking, and when.

The most accurate open-weights speaker diarization we know of, on both public suites: 15.86 macro DER on pyannote's 8-corpus benchmark — ahead of pyannoteAI's commercial precision-2 (16.06) — and 20.94 across all 12 corpora. From the research team at Re:WayAI, packaged for pyannote.audio 4.x.

Macro-average DER of SUPlime, SUPlime-L, DiariZen-L-s80-v2, pyannoteAI precision-2 and pyannote community-1 on the 8-corpus benchmark and on all 12 corpora

Two models on one recipe: SUPlime on a WavLM-Base+ backbone (114 M params) and SUPlime-L on WavLM-Large (349 M). SUPlime-L is the better model on meeting and conversational audio; the base model wins on far-field and dinner-party recordings (CHiME-6, DiPCo, NOTSOFAR-1). Per-corpus numbers for all 12 corpora are in Results below.

Install

pip install suplime            # pulls pyannote.audio>=4.0.7,<5

pip install does not bring FFmpeg, and pyannote.audio reads audio files through torchcodec, which loads FFmpeg's shared libraries at runtime — without them the first call raises RuntimeError: Could not load libtorchcodec. Install it into the same environment:

micromamba install -c conda-forge ffmpeg     # or conda/mamba; on Linux also apt install ffmpeg

Passing a waveform instead of a path needs no decoder at all:

import soundfile as sf, torch
x, sr = sf.read("meeting.wav", dtype="float32", always_2d=True)
output = pipeline({"waveform": torch.from_numpy(x.T), "sample_rate": sr})   # (channel, time)

Use

import torch
from pyannote.audio import Pipeline

pipeline = Pipeline.from_pretrained("rewayai/suplime").to(torch.device("cuda"))
# the WavLM-Large variant is a drop-in replacement, same API and same threshold:
# pipeline = Pipeline.from_pretrained("rewayai/suplime-large").to(torch.device("cuda"))
output = pipeline("meeting.wav")
for turn, _, speaker in output.speaker_diarization.itertracks(yield_label=True):
    print(f"{turn.start:.2f} {turn.end:.2f} {speaker}")

No Hugging Face token is needed: the repository is not gated.

Or build the pipeline in Python, without the repo's config.yaml — useful for pointing at local checkpoints or setting the parameters yourself:

from suplime import SuplimeDiarization
pipeline = SuplimeDiarization()                       # defaults = rewayai/suplime
pipeline.instantiate(pipeline.default_parameters())   # threshold 0.72, min_cluster_size 12

L = "rewayai/suplime-large"                           # the same, for SUPlime-L
pipeline = SuplimeDiarization(segmentation={"checkpoint": L, "subfolder": "segmentation"},
                              embedding={"checkpoint": L, "subfolder": "embedding"})

Individual models load through the usual API:

from pyannote.audio import Model
seg = Model.from_pretrained("rewayai/suplime", subfolder="segmentation")
emb = Model.from_pretrained("rewayai/suplime", subfolder="embedding")

Both repositories have the same layout, so every call above works with rewayai/suplime-large; the embedding model is byte-identical between them.

Set SUPLIME_FP16=0 to run the WavLM backbone in fp32 (default: fp16 autocast on CUDA during inference — the published numbers were produced this way).

Results

Diarization error rate (%), lower is better:

Corpus (test unless noted)SUPlimeSUPlime-LDiariZen-L-s80-v2pyannoteAI precision-2pyannote community-1
AISHELL-411.5511.2110.111.411.7
AliMeeting (far, ch1)14.4514.5510.815.220.3
AMI (IHM, Mix-Headset)12.5911.9125.69 †12.917.0
AMI (SDM)15.1414.8513.915.619.9
AVA-AVD38.0936.8342.62 †37.144.6
MSDWild (few.val)17.8117.4815.817.322.8
RAMC10.8611.1011.010.520.8
VoxConverse (v0.3)9.218.939.18.511.2
macro average (8)16.2115.8617.3816.0621.04
NOTSOFAR-1 (80-session split)20.0422.7018.86 †—27.67 †
ICSI22.9722.7726.54 †—30.84 †
CHiME-648.1148.9848.39 †—51.98 †
DiPCo30.4430.9237.56 †—34.38 †
macro average (12)20.9421.0222.53—26.10

SUPlime-L leads the 8-corpus benchmark at 15.86 macro DER — ahead of precision-2 (16.06), the base model (16.21) and DiariZen-L-s80-v2 (17.38) — while across all 12 corpora the base model takes it back by 0.06, the four extra sets being far-field and dinner-party audio where the larger backbone does not pay off. Both beat DiariZen by ~1.5 DER on the 12-corpus average, with the largest margins on close-talk AMI and on CHiME-6 / DiPCo, while DiariZen stays clearly better on the Mandarin meeting corpora and MSDWild. Per-corpus hypothesis RTTMs ship with each model; the cards cover threshold behaviour and known limitations.

Scoring conditions. Collar 0 s, overlapped speech scored, no oracle speaker count, reference cropped to each corpus' UEM, one clustering threshold (0.72) everywhere. The other systems' columns are their authors' published numbers (DiariZen, pyannote), except where † marks our own re-run of their open-source pipeline, unchanged and on the same files — and for NOTSOFAR-1 the DiariZen authors report 16.7 on a different session split.

Architecture

  • Segmentation: WavLM (learned mixture of all layers) → 4-layer Conformer → powerset classifier (4 speakers / 10 s window, up to 2 simultaneous). WavLM-Base+, 114 M params for SUPlime; WavLM-Large, 349 M for SUPlime-L. This is the only difference between them.
  • Embedding: WeSpeaker SimAM-ResNet34 with attentive statistics pooling (voxblink2_samresnet34_ft, 25 M params, 256-dim). Shared, byte for byte.
  • Clustering: agglomerative (centroid linkage), threshold 0.72, min_cluster_size 12, overlap excluded from embeddings. Shared.

Training

The full training pipeline is released as well — corpus preparation, the augmentation setup, the training command and the checkpoint soup — so the models can be reproduced from scratch or trained further on your own data. Recipe and commands for both variants: training/README.md.

License

Code: MIT (src/suplime/models/samresnet.py: Apache-2.0, ported from WeSpeaker). Model weights: CC BY-NC 4.0. Non-commercial because several training corpora (RAMC, MSDWild, AVA-AVD) are licensed for research use only and others (AISHELL-4, AliMeeting) are share-alike; the most restrictive terms apply to the derived model. Details in the model card. Built on pyannote.audio (MIT), WavLM (MIT) and WeSpeaker (Apache-2.0).

Diarization is not enough?

SUPlime tells you who spoke and when. If you also want to know what they said, and who they actually are, that is our day job: Re:WayAI is the speech API that knows who's talking.

  • Transcription in 25 European languages (40 for real-time streaming), with overlap-robust diarization from the same team that built SUPlime
  • Voice enrollment: speakers get their real names, not SPEAKER_03
  • Real-time streaming over WebSocket, transcript Q&A, one-click DOCX / PDF exports
  • On-premise deployment, audio deleted right after processing, ISO 27001

Pay-as-you-go, 50 free hours to start, no subscription. The lime approves.

Significant stargazers

Michael Pedersen

46 followers · starred Sep 2026

Languages

Python

96.3%

Shell

3.7%