LibreTranslate/MiniSBD

Free and open source library for fast sentence boundary detection

8

stars

17

commits

Python

primary language

Mar 23, 2026

updated

nlp
sbd
sentence-boundary

README

MiniSBD

Free and open source Python library for fast sentence boundary detection. It uses 8bit quantized ONNX models for inference, thus making it fast and lightweight.

The only dependency is onnxruntime / onnxruntime-gpu.

Installation

pip install -U minisbd

Usage

from minisbd import SBDetect

text = """
La Révolution française (1789-1799) est une période de bouleversements politiques et sociaux en France et dans ses colonies, ainsi qu'en Europe à la fin du XVIIIe siècle. Traditionnellement, on la fait commencer à l'ouverture des États généraux le 5 mai 1789 et finir au coup d'État de Napoléon Bonaparte le 9 novembre 1799 (18 brumaire de l'an VIII). En ce qui concerne l'histoire de France, elle met fin à l'Ancien Régime, notamment à la monarchie absolue, remplacée par la monarchie constitutionnelle (1789-1792) puis par la Première République.
"""

detector = SBDetect("fr", use_gpu=True)
for sent in detector.sentences(text):
    print(f"--> {sent}")

# --> La Révolution française (1789-1799) est une période de bouleversements politiques et sociaux en France et dans ses colonies, ainsi qu'en Europe à la fin du XVIIIe siècle.
# --> Traditionnellement, on la fait commencer à l'ouverture des États généraux le 5 mai 1789 et finir au coup d'État de Napoléon Bonaparte le 9 novembre 1799 (18 brumaire de l'an VIII).
# --> En ce qui concerne l'histoire de France, elle met fin à l'Ancien Régime, notamment à la monarchie absolue, remplacée par la monarchie constitutionnelle (1789-1792) puis par la Première République.

By default models are downloaded from GitHub and stored in the user's ~/.cache/minisbd folder. You can change this at runtime via:

from minisbd import models
models.cache_dir = '/path/to/cache'

You can optionally specify a path to a ONNX model instead of having MiniSBD download the model for you:

from minisbd import SBDetect
detector = SBDetect("/path/to/model.onnx")
# ...

Language Support

from minisbd.models import list_models
print(list_models())
LanguageCode
Afrikaansaf
Ancient Greekgrc
Ancient Hebrewhbo
Arabicar
Armenianhy
Basqueeu
Belarusianbe
Bulgarianbg
Buryatbxr
Catalanca
Chinese (Simplified)zh-hans
Chinese (Traditional)zh-hant
Classical Chineselzh
Copticcop
Croatianhr
Czechcs
Danishda
Dutchnl
Englishen
Erzyamyv
Estonianet
Faroesefo
Finnishfi
Frenchfr
Galiciangl
Germande
Gothicgot
Greekel
Hebrewhe
Hindihi
Hungarianhu
Icelandicis
Indonesianid
Irishga
Italianit
Japaneseja
Kazakhkk
Koreanko
Kurmanjikmr
Kyrgyzky
Latinla
Latvianlv
Ligurianlij
Lithuanianlt
Maghrebi Arabic Frenchqaf
Maltesemt
Manxgv
Marathimr
Naijapcm
North Samisme
Norwegiannb
Norwegian Nynorsknn
Old Church Slavoniccu
Old East Slavicorv
Old Frenchfro
Persianfa
Polishpl
Pomakqpm
Portuguesept
Romanianro
Russianru
Sanskritsa
Scottish Gaelicgd
Serbiansr
Slovaksk
Sloveniansl
Spanishes
Swedishsv
Tamilta
Telugute
Thaith
Turkishtr
Turkish Germanqtd
Ukrainianuk
Upper Sorbianhsb
Urduur
Uyghurug
Vietnamesevi
Welshcy
Western Armenianhyw
Wolofwo

Converting Stanza Models

The extract.py script can be used to extract existing Stanza models and convert them to ONNX. See the source code.

Credits

MiniSBD is a port of Stanza's tokenizer models to ONNX. The models are the same as those from Stanza, but have been converted to ONNX and quantized for faster inference and smaller size.

License

AGPLv3

Some code has been originally modified from Stanza.

Contributors

pierotofy

16 commits

cornmail

1 commits

LibreTranslate/MiniSBD

Free and open source library for fast sentence boundary detection

8

stars

17

commits

Python

primary language

Mar 23, 2026

updated

nlp
sbd
sentence-boundary

README

MiniSBD

Free and open source Python library for fast sentence boundary detection. It uses 8bit quantized ONNX models for inference, thus making it fast and lightweight.

The only dependency is onnxruntime / onnxruntime-gpu.

Installation

pip install -U minisbd

Usage

from minisbd import SBDetect

text = """
La Révolution française (1789-1799) est une période de bouleversements politiques et sociaux en France et dans ses colonies, ainsi qu'en Europe à la fin du XVIIIe siècle. Traditionnellement, on la fait commencer à l'ouverture des États généraux le 5 mai 1789 et finir au coup d'État de Napoléon Bonaparte le 9 novembre 1799 (18 brumaire de l'an VIII). En ce qui concerne l'histoire de France, elle met fin à l'Ancien Régime, notamment à la monarchie absolue, remplacée par la monarchie constitutionnelle (1789-1792) puis par la Première République.
"""

detector = SBDetect("fr", use_gpu=True)
for sent in detector.sentences(text):
    print(f"--> {sent}")

# --> La Révolution française (1789-1799) est une période de bouleversements politiques et sociaux en France et dans ses colonies, ainsi qu'en Europe à la fin du XVIIIe siècle.
# --> Traditionnellement, on la fait commencer à l'ouverture des États généraux le 5 mai 1789 et finir au coup d'État de Napoléon Bonaparte le 9 novembre 1799 (18 brumaire de l'an VIII).
# --> En ce qui concerne l'histoire de France, elle met fin à l'Ancien Régime, notamment à la monarchie absolue, remplacée par la monarchie constitutionnelle (1789-1792) puis par la Première République.

By default models are downloaded from GitHub and stored in the user's ~/.cache/minisbd folder. You can change this at runtime via:

from minisbd import models
models.cache_dir = '/path/to/cache'

You can optionally specify a path to a ONNX model instead of having MiniSBD download the model for you:

from minisbd import SBDetect
detector = SBDetect("/path/to/model.onnx")
# ...

Language Support

from minisbd.models import list_models
print(list_models())
LanguageCode
Afrikaansaf
Ancient Greekgrc
Ancient Hebrewhbo
Arabicar
Armenianhy
Basqueeu
Belarusianbe
Bulgarianbg
Buryatbxr
Catalanca
Chinese (Simplified)zh-hans
Chinese (Traditional)zh-hant
Classical Chineselzh
Copticcop
Croatianhr
Czechcs
Danishda
Dutchnl
Englishen
Erzyamyv
Estonianet
Faroesefo
Finnishfi
Frenchfr
Galiciangl
Germande
Gothicgot
Greekel
Hebrewhe
Hindihi
Hungarianhu
Icelandicis
Indonesianid
Irishga
Italianit
Japaneseja
Kazakhkk
Koreanko
Kurmanjikmr
Kyrgyzky
Latinla
Latvianlv
Ligurianlij
Lithuanianlt
Maghrebi Arabic Frenchqaf
Maltesemt
Manxgv
Marathimr
Naijapcm
North Samisme
Norwegiannb
Norwegian Nynorsknn
Old Church Slavoniccu
Old East Slavicorv
Old Frenchfro
Persianfa
Polishpl
Pomakqpm
Portuguesept
Romanianro
Russianru
Sanskritsa
Scottish Gaelicgd
Serbiansr
Slovaksk
Sloveniansl
Spanishes
Swedishsv
Tamilta
Telugute
Thaith
Turkishtr
Turkish Germanqtd
Ukrainianuk
Upper Sorbianhsb
Urduur
Uyghurug
Vietnamesevi
Welshcy
Western Armenianhyw
Wolofwo

Converting Stanza Models

The extract.py script can be used to extract existing Stanza models and convert them to ONNX. See the source code.

Credits

MiniSBD is a port of Stanza's tokenizer models to ONNX. The models are the same as those from Stanza, but have been converted to ONNX and quantized for faster inference and smaller size.

License

AGPLv3

Some code has been originally modified from Stanza.

Contributors

pierotofy

16 commits

cornmail

1 commits

Languages

Python

100.0%