videosdk-live/Namo-Turn-Detector-v1-Multilingual

Model

๐ŸŽฏ Namo Turn Detector v1 - MultiLingual

27

14 commits

1 linked in READMEs

updated Oct 15, 2025

See the code

README

๐ŸŽฏ Namo Turn Detector v1 - MultiLingual

License ONNX Model Size Inference Speed

๐Ÿš€ Namo Turn Detection Model for Multiple Languages

๐Ÿ‡ธ๐Ÿ‡ฆ Arabic, ๐Ÿ‡ฎ๐Ÿ‡ณ Bengali, ๐Ÿ‡จ๐Ÿ‡ณ Chinese, ๐Ÿ‡ฉ๐Ÿ‡ฐ Danish, ๐Ÿ‡ณ๐Ÿ‡ฑ Dutch, ๐Ÿ‡ฉ๐Ÿ‡ช German, ๐Ÿ‡ฌ๐Ÿ‡ง๐Ÿ‡บ๐Ÿ‡ธ English, ๐Ÿ‡ซ๐Ÿ‡ฎ Finnish, ๐Ÿ‡ซ๐Ÿ‡ท French, ๐Ÿ‡ฎ๐Ÿ‡ณ Hindi, ๐Ÿ‡ฎ๐Ÿ‡ฉ Indonesian, ๐Ÿ‡ฎ๐Ÿ‡น Italian, ๐Ÿ‡ฏ๐Ÿ‡ต Japanese, ๐Ÿ‡ฐ๐Ÿ‡ท Korean, ๐Ÿ‡ฎ๐Ÿ‡ณ Marathi, ๐Ÿ‡ณ๐Ÿ‡ด Norwegian, ๐Ÿ‡ต๐Ÿ‡ฑ Polish, ๐Ÿ‡ต๐Ÿ‡น Portuguese, ๐Ÿ‡ท๐Ÿ‡บ Russian, ๐Ÿ‡ช๐Ÿ‡ธ Spanish, ๐Ÿ‡น๐Ÿ‡ท Turkish, ๐Ÿ‡บ๐Ÿ‡ฆ Ukrainian, and ๐Ÿ‡ป๐Ÿ‡ณ Vietnamese


๐Ÿ“‹ Overview

The Namo Turn Detector is a specialized AI model designed to solve one of the most challenging problems in conversational AI: knowing when a user has finished speaking.

This Multilingual model uses advanced natural language understanding to distinguish between:

  • โœ… Complete utterances (user is done speaking)
  • ๐Ÿ”„ Incomplete utterances (user will continue speaking)

Built on mmBERT architecture and optimized with quantized ONNX format, it delivers enterprise-grade performance with minimal latency.

๐Ÿ”‘ Key Features

  • Turn Detection Specialist: Detects end-of-turn vs. continuation in multilingual speech transcripts.
  • Low Latency: Optimized with quantized ONNX for <29ms inference.
  • Robust Performance: Average 90.25% accuracy on multilingual utterances.
  • Easy Integration: Compatible with Python, ONNX Runtime, and VideoSDK Agents SDK.
  • Enterprise Ready: Supports real-time conversational AI and voice assistants.

๐Ÿ“Š Performance Metrics

MetricScore
โšก Latency<29ms
๐Ÿ’พ Model Size~295MB
LanguageAccuracyPrecisionRecallF1 ScoreSamples
๐Ÿ‡น๐Ÿ‡ท Turkish0.97310.96110.98530.9730966
๐Ÿ‡ฐ๐Ÿ‡ท Korean0.96850.95410.98420.9690890
๐Ÿ‡ฉ๐Ÿ‡ช German0.94250.91350.97720.94431322
๐Ÿ‡ฏ๐Ÿ‡ต Japanese0.94360.90990.98570.9463834
๐Ÿ‡ฎ๐Ÿ‡ณ Hindi0.93980.92760.96030.94361295
๐Ÿ‡ณ๐Ÿ‡ฑ Dutch0.92790.89590.97380.93321401
๐Ÿ‡ณ๐Ÿ‡ด Norwegian0.91650.87170.98010.92271976
๐Ÿ‡จ๐Ÿ‡ณ Chinese0.91640.88590.96080.9219945
๐Ÿ‡ซ๐Ÿ‡ฎ Finnish0.91580.87460.97020.91991010
๐Ÿ‡ฌ๐Ÿ‡ง English0.90860.85070.98010.91082845
๐Ÿ‡ฎ๐Ÿ‡ฉ Indonesian0.90220.85140.97070.9071971
๐Ÿ‡ฎ๐Ÿ‡น Italian0.90150.85620.96400.9069782
๐Ÿ‡ต๐Ÿ‡ฑ Polish0.90680.86190.95680.9069976
๐Ÿ‡ต๐Ÿ‡น Portuguese0.89560.84100.96760.89991398
๐Ÿ‡ฉ๐Ÿ‡ฐ Danish0.89730.85170.96440.9045779
๐Ÿ‡ช๐Ÿ‡ธ Spanish0.88880.83040.96810.89401295
๐Ÿ‡ฎ๐Ÿ‡ณ Marathi0.88500.87620.90080.8883774
๐Ÿ‡ท๐Ÿ‡บ Russian0.87480.83180.95470.88901470
๐Ÿ‡บ๐Ÿ‡ฆ Ukrainian0.87940.81640.95870.8819929
๐Ÿ‡ป๐Ÿ‡ณ Vietnamese0.86450.81350.94390.87381004
๐Ÿ‡ธ๐Ÿ‡ฆ Arabic0.84900.79650.94390.8639947
๐Ÿ‡ฎ๐Ÿ‡ณ Bengali0.79400.78740.79390.79071000

๐Ÿ“Š Evaluated on 25,000+ Multilingual utterances from diverse conversational contexts

โšก๏ธ Speed Analysis

Alt text

๐Ÿ”ง Train & Test Scripts

Train Script Test Script

๐Ÿ› ๏ธ Installation

To use this model, you will need to install the following libraries.

pip install onnxruntime transformers huggingface_hub

๐Ÿš€ Quick Start

You can run inference directly from Hugging Face repository.

import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer
from huggingface_hub import hf_hub_download

class TurnDetector:
    def __init__(self, repo_id="videosdk-live/Namo-Turn-Detector-v1-Multilingual"):
        """
        Initializes the detector by downloading the model and tokenizer
        from the Hugging Face Hub.
        """
        print(f"Loading model from repo: {repo_id}")
        
        # Download the model and tokenizer from the Hub
        # Authentication is handled automatically if you are logged in
        model_path = hf_hub_download(repo_id=repo_id, filename="model_quant.onnx")
        self.tokenizer = AutoTokenizer.from_pretrained(repo_id)
        
        # Set up the ONNX Runtime inference session
        self.session = ort.InferenceSession(model_path)
        self.max_length = 8192
        print("โœ… Model and tokenizer loaded successfully.")

    def predict(self, text: str) -> tuple:
        """
        Predicts if a given text utterance is the end of a turn.
        Returns (predicted_label, confidence) where:
        - predicted_label: 0 for "Not End of Turn", 1 for "End of Turn"
        - confidence: confidence score between 0 and 1
        """
        # Tokenize the input text
        inputs = self.tokenizer(
            text,
            truncation=True,
            max_length=self.max_length,
            return_tensors="np"
        )
        
        # Prepare the feed dictionary for the ONNX model
        feed_dict = {
            "input_ids": inputs["input_ids"],
            "attention_mask": inputs["attention_mask"]
        }
        
        # Run inference
        outputs = self.session.run(None, feed_dict)
        logits = outputs[0]

        probabilities = self._softmax(logits[0])
        predicted_label = np.argmax(probabilities)
        confidence = float(np.max(probabilities))

        return predicted_label, confidence

    def _softmax(self, x, axis=None):
        if axis is None:
            axis = -1
        exp_x = np.exp(x - np.max(x, axis=axis, keepdims=True))
        return exp_x / np.sum(exp_x, axis=axis, keepdims=True)

# --- Example Usage ---
if __name__ == "__main__":
    detector = TurnDetector()
    
    sentences = [
        "They're often made with oil or sugar.",                         # Expected: End of Turn
        "I think the next logical step is to",                           # Expected: Not End of Turn
        "What are you doing tonight?",                                   # Expected: End of Turn
        "The Revenue Act of 1862 adopted rates that increased with",     # Expected: Not End of Turn
    ]
    
    for sentence in sentences:
        predicted_label, confidence = detector.predict(sentence)
        result = "End of Turn" if predicted_label == 1 else "Not End of Turn"
        print(f"'{sentence}' -> {result} (confidence: {confidence:.3f})")
        print("-" * 50)

๐Ÿค– VideoSDK Agents Integration

Integrate this turn detector directly with VideoSDK Agents for production-ready conversational AI applications.

from videosdk_agents import NamoTurnDetectorV1, pre_download_namo_turn_v1_model

#download model
pre_download_namo_turn_v1_model()

# Initialize Multilingual turn detector for VideoSDK Agents
turn_detector = NamoTurnDetectorV1()

๐Ÿ“š Complete Integration Guide - Learn how to use NamoTurnDetectorV1 with VideoSDK Agents

๐Ÿ“– Citation

@model{namo_turn_detector_en_2025,
  title={Namo Turn Detector v1: Multilingual},
  author={VideoSDK Team},
  year={2025},
  publisher={Hugging Face},
  url={https://huggingface.co/videosdk-live/Namo-Turn-Detector-v1-Multilingual},
  note={ONNX-optimized mmBERT for turn detection in 23 Languages}
}

๐Ÿ“„ License

This project is licensed under the Apache License 2.0 - see the LICENSE file for details.

Made with โค๏ธ by the VideoSDK Team

VideoSDK

conversational-ai
end-of-utterance
mmbert
model-index
modernbert
onnx
onnxruntime
quantized
real-time
turn-detection
voice-activity-detection
voice-assistant

Contributors

videosdk-sdk

6 commits

aarya-vsdk

5 commits

poojann

2 commits

videosdk-live/Namo-Turn-Detector-v1-Multilingual

Model

๐ŸŽฏ Namo Turn Detector v1 - MultiLingual

27

14 commits

1 linked in READMEs

updated Oct 15, 2025

See the code

README

๐ŸŽฏ Namo Turn Detector v1 - MultiLingual

License ONNX Model Size Inference Speed

๐Ÿš€ Namo Turn Detection Model for Multiple Languages

๐Ÿ‡ธ๐Ÿ‡ฆ Arabic, ๐Ÿ‡ฎ๐Ÿ‡ณ Bengali, ๐Ÿ‡จ๐Ÿ‡ณ Chinese, ๐Ÿ‡ฉ๐Ÿ‡ฐ Danish, ๐Ÿ‡ณ๐Ÿ‡ฑ Dutch, ๐Ÿ‡ฉ๐Ÿ‡ช German, ๐Ÿ‡ฌ๐Ÿ‡ง๐Ÿ‡บ๐Ÿ‡ธ English, ๐Ÿ‡ซ๐Ÿ‡ฎ Finnish, ๐Ÿ‡ซ๐Ÿ‡ท French, ๐Ÿ‡ฎ๐Ÿ‡ณ Hindi, ๐Ÿ‡ฎ๐Ÿ‡ฉ Indonesian, ๐Ÿ‡ฎ๐Ÿ‡น Italian, ๐Ÿ‡ฏ๐Ÿ‡ต Japanese, ๐Ÿ‡ฐ๐Ÿ‡ท Korean, ๐Ÿ‡ฎ๐Ÿ‡ณ Marathi, ๐Ÿ‡ณ๐Ÿ‡ด Norwegian, ๐Ÿ‡ต๐Ÿ‡ฑ Polish, ๐Ÿ‡ต๐Ÿ‡น Portuguese, ๐Ÿ‡ท๐Ÿ‡บ Russian, ๐Ÿ‡ช๐Ÿ‡ธ Spanish, ๐Ÿ‡น๐Ÿ‡ท Turkish, ๐Ÿ‡บ๐Ÿ‡ฆ Ukrainian, and ๐Ÿ‡ป๐Ÿ‡ณ Vietnamese


๐Ÿ“‹ Overview

The Namo Turn Detector is a specialized AI model designed to solve one of the most challenging problems in conversational AI: knowing when a user has finished speaking.

This Multilingual model uses advanced natural language understanding to distinguish between:

  • โœ… Complete utterances (user is done speaking)
  • ๐Ÿ”„ Incomplete utterances (user will continue speaking)

Built on mmBERT architecture and optimized with quantized ONNX format, it delivers enterprise-grade performance with minimal latency.

๐Ÿ”‘ Key Features

  • Turn Detection Specialist: Detects end-of-turn vs. continuation in multilingual speech transcripts.
  • Low Latency: Optimized with quantized ONNX for <29ms inference.
  • Robust Performance: Average 90.25% accuracy on multilingual utterances.
  • Easy Integration: Compatible with Python, ONNX Runtime, and VideoSDK Agents SDK.
  • Enterprise Ready: Supports real-time conversational AI and voice assistants.

๐Ÿ“Š Performance Metrics

MetricScore
โšก Latency<29ms
๐Ÿ’พ Model Size~295MB
LanguageAccuracyPrecisionRecallF1 ScoreSamples
๐Ÿ‡น๐Ÿ‡ท Turkish0.97310.96110.98530.9730966
๐Ÿ‡ฐ๐Ÿ‡ท Korean0.96850.95410.98420.9690890
๐Ÿ‡ฉ๐Ÿ‡ช German0.94250.91350.97720.94431322
๐Ÿ‡ฏ๐Ÿ‡ต Japanese0.94360.90990.98570.9463834
๐Ÿ‡ฎ๐Ÿ‡ณ Hindi0.93980.92760.96030.94361295
๐Ÿ‡ณ๐Ÿ‡ฑ Dutch0.92790.89590.97380.93321401
๐Ÿ‡ณ๐Ÿ‡ด Norwegian0.91650.87170.98010.92271976
๐Ÿ‡จ๐Ÿ‡ณ Chinese0.91640.88590.96080.9219945
๐Ÿ‡ซ๐Ÿ‡ฎ Finnish0.91580.87460.97020.91991010
๐Ÿ‡ฌ๐Ÿ‡ง English0.90860.85070.98010.91082845
๐Ÿ‡ฎ๐Ÿ‡ฉ Indonesian0.90220.85140.97070.9071971
๐Ÿ‡ฎ๐Ÿ‡น Italian0.90150.85620.96400.9069782
๐Ÿ‡ต๐Ÿ‡ฑ Polish0.90680.86190.95680.9069976
๐Ÿ‡ต๐Ÿ‡น Portuguese0.89560.84100.96760.89991398
๐Ÿ‡ฉ๐Ÿ‡ฐ Danish0.89730.85170.96440.9045779
๐Ÿ‡ช๐Ÿ‡ธ Spanish0.88880.83040.96810.89401295
๐Ÿ‡ฎ๐Ÿ‡ณ Marathi0.88500.87620.90080.8883774
๐Ÿ‡ท๐Ÿ‡บ Russian0.87480.83180.95470.88901470
๐Ÿ‡บ๐Ÿ‡ฆ Ukrainian0.87940.81640.95870.8819929
๐Ÿ‡ป๐Ÿ‡ณ Vietnamese0.86450.81350.94390.87381004
๐Ÿ‡ธ๐Ÿ‡ฆ Arabic0.84900.79650.94390.8639947
๐Ÿ‡ฎ๐Ÿ‡ณ Bengali0.79400.78740.79390.79071000

๐Ÿ“Š Evaluated on 25,000+ Multilingual utterances from diverse conversational contexts

โšก๏ธ Speed Analysis

Alt text

๐Ÿ”ง Train & Test Scripts

Train Script Test Script

๐Ÿ› ๏ธ Installation

To use this model, you will need to install the following libraries.

pip install onnxruntime transformers huggingface_hub

๐Ÿš€ Quick Start

You can run inference directly from Hugging Face repository.

import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer
from huggingface_hub import hf_hub_download

class TurnDetector:
    def __init__(self, repo_id="videosdk-live/Namo-Turn-Detector-v1-Multilingual"):
        """
        Initializes the detector by downloading the model and tokenizer
        from the Hugging Face Hub.
        """
        print(f"Loading model from repo: {repo_id}")
        
        # Download the model and tokenizer from the Hub
        # Authentication is handled automatically if you are logged in
        model_path = hf_hub_download(repo_id=repo_id, filename="model_quant.onnx")
        self.tokenizer = AutoTokenizer.from_pretrained(repo_id)
        
        # Set up the ONNX Runtime inference session
        self.session = ort.InferenceSession(model_path)
        self.max_length = 8192
        print("โœ… Model and tokenizer loaded successfully.")

    def predict(self, text: str) -> tuple:
        """
        Predicts if a given text utterance is the end of a turn.
        Returns (predicted_label, confidence) where:
        - predicted_label: 0 for "Not End of Turn", 1 for "End of Turn"
        - confidence: confidence score between 0 and 1
        """
        # Tokenize the input text
        inputs = self.tokenizer(
            text,
            truncation=True,
            max_length=self.max_length,
            return_tensors="np"
        )
        
        # Prepare the feed dictionary for the ONNX model
        feed_dict = {
            "input_ids": inputs["input_ids"],
            "attention_mask": inputs["attention_mask"]
        }
        
        # Run inference
        outputs = self.session.run(None, feed_dict)
        logits = outputs[0]

        probabilities = self._softmax(logits[0])
        predicted_label = np.argmax(probabilities)
        confidence = float(np.max(probabilities))

        return predicted_label, confidence

    def _softmax(self, x, axis=None):
        if axis is None:
            axis = -1
        exp_x = np.exp(x - np.max(x, axis=axis, keepdims=True))
        return exp_x / np.sum(exp_x, axis=axis, keepdims=True)

# --- Example Usage ---
if __name__ == "__main__":
    detector = TurnDetector()
    
    sentences = [
        "They're often made with oil or sugar.",                         # Expected: End of Turn
        "I think the next logical step is to",                           # Expected: Not End of Turn
        "What are you doing tonight?",                                   # Expected: End of Turn
        "The Revenue Act of 1862 adopted rates that increased with",     # Expected: Not End of Turn
    ]
    
    for sentence in sentences:
        predicted_label, confidence = detector.predict(sentence)
        result = "End of Turn" if predicted_label == 1 else "Not End of Turn"
        print(f"'{sentence}' -> {result} (confidence: {confidence:.3f})")
        print("-" * 50)

๐Ÿค– VideoSDK Agents Integration

Integrate this turn detector directly with VideoSDK Agents for production-ready conversational AI applications.

from videosdk_agents import NamoTurnDetectorV1, pre_download_namo_turn_v1_model

#download model
pre_download_namo_turn_v1_model()

# Initialize Multilingual turn detector for VideoSDK Agents
turn_detector = NamoTurnDetectorV1()

๐Ÿ“š Complete Integration Guide - Learn how to use NamoTurnDetectorV1 with VideoSDK Agents

๐Ÿ“– Citation

@model{namo_turn_detector_en_2025,
  title={Namo Turn Detector v1: Multilingual},
  author={VideoSDK Team},
  year={2025},
  publisher={Hugging Face},
  url={https://huggingface.co/videosdk-live/Namo-Turn-Detector-v1-Multilingual},
  note={ONNX-optimized mmBERT for turn detection in 23 Languages}
}

๐Ÿ“„ License

This project is licensed under the Apache License 2.0 - see the LICENSE file for details.

Made with โค๏ธ by the VideoSDK Team

VideoSDK

conversational-ai
end-of-utterance
mmbert
model-index
modernbert
onnx
onnxruntime
quantized
real-time
turn-detection
voice-activity-detection
voice-assistant

Contributors

videosdk-sdk

6 commits

aarya-vsdk

5 commits

poojann

2 commits