napari-voxtell integrates VoxTell, a 3D vision-language segmentation model, into the napari ecosystem. This plugin enables text-based prompting for volumetric medical image segmentation, offering an alternative to traditional interaction methods such as bounding boxes, point clicks, or manual brush strokes, used e.g. in our nnInteractive plugin.
VoxTell accepts free-form text descriptions (e.g., "liver", "aortic arch", "brain tumor") to generate 3D segmentation masks. As an experimental research tool, napari-voxtell is designed to facilitate exploration and prototyping in medical image analysis workflows rather than production clinical use.
Note: VoxTell is an ongoing research project and may produce variable results depending on anatomical region, imaging modality, and prompt specificity. Users should validate outputs carefully and not rely on this tool for clinical decision-making without expert review.
⚠️ Image Orientation (Critical): For correct anatomical localization (e.g., distinguishing left from right), images must be in RAS orientation. VoxTell was trained on data reoriented using this specific reader. To make this robust, this plugin ships its own VoxTell reader that opens .nii.gz files through exactly that reader (see Getting Started), and it reorients results back to the image's original orientation on save. A quick way to spot a mismatch is if a simple prompt like "liver" fails and segments e.g. parts of the spleen instead.
Image Spacing: The model does not resample images to a standardized spacing for faster inference. Performance may degrade on images with very uncommon voxel spacings (e.g., super high-resolution brain MRI). In such cases, consider resampling the image to a more typical clinical spacing (e.g., 1.5×1.5×1.5 mm³) before segmentation.
VoxTell supports Python 3.10+ and works with Conda, pip, or any other virtual environment. Here's an example using Conda:
conda create -n voxtell python=3.12
conda activate voxtell
Install PyTorch compatible with your CUDA version. For example, for Ubuntu with a modern Nvidia GPU:
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu126
Any PyTorch >= 2.1.2 works except the 2.9.x series, which has a 3D-convolution memory regression (pytorch#166122; also excluded by VoxTell and nnU-Net).
For other configurations (Mac, CPU, different CUDA versions), please refer to the PyTorch Get Started page.
Install the latest release from PyPI (this also installs VoxTell 0.1.2 or newer for local inference):
pip install napari-voxtell
For the development version (the main branch, which may contain unreleased changes), install
directly from GitHub:
pip install "git+https://github.com/MIC-DKFZ/napari-voxtell.git"
For development, clone and install in editable mode (you can also use uv):
git clone https://github.com/MIC-DKFZ/napari-voxtell.git
cd napari-voxtell
pip install -e .
Note: Model weights are automatically downloaded from Hugging Face on first use. This may take a few minutes depending on your internet connection.
You can launch the plugin in three ways.
[!IMPORTANT] When opening a
.nii.gzfile, choose the VoxTell reader (notnapari-nifti). The VoxTell reader reorients the volume to RAS using the sameNibabelIOWithReorientreader the model was trained with, which is required for correct left/right and organ localization. Volumes opened with other readers may be mis-oriented and produce wrong-side / wrong-organ masks.
Option A: Start napari and activate manually
napari
Then go to Plugins > napari-voxtell.
Option B: Start napari with the widget open
napari -w napari-voxtell
Option C: Open an image directly with the widget
napari path/to/your/image.nii.gz -w napari-voxtell
Please carefully review all segmentation outputs. Model performance varies with anatomical complexity, imaging quality, spacing, and prompt clarity. This tool is intended for research exploration, not validated clinical workflows.
The plugin can run the napari GUI locally while inference runs on a remote GPU machine (workstation or cluster). Loading, display and saving stay identical to local mode — only the inference is offloaded over HTTP.
pip install "voxtell[server]"
voxtell-server --host 127.0.0.1 --port 1527
ssh -N -L 1527:127.0.0.1:1527 your-workstation
http://127.0.0.1:1527,
click Connect, then open an image and Submit as usual. The image is uploaded to the
server on the first Submit; the progress bar and Cancel work exactly as in local mode.For a trusted LAN you can instead bind the server to 0.0.0.0 and set an API key
(voxtell-server --host 0.0.0.0 --api-key ..., then paste the key into the widget).
If you use napari-voxtell in your research, please cite our paper:
@inproceedings{rokuss2026voxtell,
title={Voxtell: Free-text promptable universal 3d medical image segmentation},
author={Rokuss, Maximilian and Langenberg, Moritz and Kirchhoff, Yannick and Isensee, Fabian and Hamm, Benjamin and Ulrich, Constantin and Regnery, Sebastian and Bauer, Lukas and Katsigiannopulos, Efthimios and Norajitra, Tobias and Maier-Hein, Klaus},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
pages={37538--37557},
year={2026}
}
This repository is licensed under the Apache-2.0 License.
Important: The default model checkpoints downloaded by this plugin are licensed under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 (CC-BY-NC-SA 4.0). Please review the Hugging Face Model Card for details regarding model usage and limitations.
Contributions are welcome! Please feel free to submit a Pull Request or open an issue for bugs and feature requests.
Special shoutout to Benjamin Hamm who created the first version of this plugin. For questions, issues, or collaborations, please contact:
📧 maximilian.rokuss@dkfz-heidelberg.de / benjamin.hamm@dkfz-heidelberg.de
The remote (client/server) inference mode is inspired by MIC-DKFZ's nnInteractive and napari-nninteractive (Apache-2.0).
Python
100.0%
napari-voxtell integrates VoxTell, a 3D vision-language segmentation model, into the napari ecosystem. This plugin enables text-based prompting for volumetric medical image segmentation, offering an alternative to traditional interaction methods such as bounding boxes, point clicks, or manual brush strokes, used e.g. in our nnInteractive plugin.
VoxTell accepts free-form text descriptions (e.g., "liver", "aortic arch", "brain tumor") to generate 3D segmentation masks. As an experimental research tool, napari-voxtell is designed to facilitate exploration and prototyping in medical image analysis workflows rather than production clinical use.
Note: VoxTell is an ongoing research project and may produce variable results depending on anatomical region, imaging modality, and prompt specificity. Users should validate outputs carefully and not rely on this tool for clinical decision-making without expert review.
⚠️ Image Orientation (Critical): For correct anatomical localization (e.g., distinguishing left from right), images must be in RAS orientation. VoxTell was trained on data reoriented using this specific reader. To make this robust, this plugin ships its own VoxTell reader that opens .nii.gz files through exactly that reader (see Getting Started), and it reorients results back to the image's original orientation on save. A quick way to spot a mismatch is if a simple prompt like "liver" fails and segments e.g. parts of the spleen instead.
Image Spacing: The model does not resample images to a standardized spacing for faster inference. Performance may degrade on images with very uncommon voxel spacings (e.g., super high-resolution brain MRI). In such cases, consider resampling the image to a more typical clinical spacing (e.g., 1.5×1.5×1.5 mm³) before segmentation.
VoxTell supports Python 3.10+ and works with Conda, pip, or any other virtual environment. Here's an example using Conda:
conda create -n voxtell python=3.12
conda activate voxtell
Install PyTorch compatible with your CUDA version. For example, for Ubuntu with a modern Nvidia GPU:
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu126
Any PyTorch >= 2.1.2 works except the 2.9.x series, which has a 3D-convolution memory regression (pytorch#166122; also excluded by VoxTell and nnU-Net).
For other configurations (Mac, CPU, different CUDA versions), please refer to the PyTorch Get Started page.
Install the latest release from PyPI (this also installs VoxTell 0.1.2 or newer for local inference):
pip install napari-voxtell
For the development version (the main branch, which may contain unreleased changes), install
directly from GitHub:
pip install "git+https://github.com/MIC-DKFZ/napari-voxtell.git"
For development, clone and install in editable mode (you can also use uv):
git clone https://github.com/MIC-DKFZ/napari-voxtell.git
cd napari-voxtell
pip install -e .
Note: Model weights are automatically downloaded from Hugging Face on first use. This may take a few minutes depending on your internet connection.
You can launch the plugin in three ways.
[!IMPORTANT] When opening a
.nii.gzfile, choose the VoxTell reader (notnapari-nifti). The VoxTell reader reorients the volume to RAS using the sameNibabelIOWithReorientreader the model was trained with, which is required for correct left/right and organ localization. Volumes opened with other readers may be mis-oriented and produce wrong-side / wrong-organ masks.
Option A: Start napari and activate manually
napari
Then go to Plugins > napari-voxtell.
Option B: Start napari with the widget open
napari -w napari-voxtell
Option C: Open an image directly with the widget
napari path/to/your/image.nii.gz -w napari-voxtell
Please carefully review all segmentation outputs. Model performance varies with anatomical complexity, imaging quality, spacing, and prompt clarity. This tool is intended for research exploration, not validated clinical workflows.
The plugin can run the napari GUI locally while inference runs on a remote GPU machine (workstation or cluster). Loading, display and saving stay identical to local mode — only the inference is offloaded over HTTP.
pip install "voxtell[server]"
voxtell-server --host 127.0.0.1 --port 1527
ssh -N -L 1527:127.0.0.1:1527 your-workstation
http://127.0.0.1:1527,
click Connect, then open an image and Submit as usual. The image is uploaded to the
server on the first Submit; the progress bar and Cancel work exactly as in local mode.For a trusted LAN you can instead bind the server to 0.0.0.0 and set an API key
(voxtell-server --host 0.0.0.0 --api-key ..., then paste the key into the widget).
If you use napari-voxtell in your research, please cite our paper:
@inproceedings{rokuss2026voxtell,
title={Voxtell: Free-text promptable universal 3d medical image segmentation},
author={Rokuss, Maximilian and Langenberg, Moritz and Kirchhoff, Yannick and Isensee, Fabian and Hamm, Benjamin and Ulrich, Constantin and Regnery, Sebastian and Bauer, Lukas and Katsigiannopulos, Efthimios and Norajitra, Tobias and Maier-Hein, Klaus},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
pages={37538--37557},
year={2026}
}
This repository is licensed under the Apache-2.0 License.
Important: The default model checkpoints downloaded by this plugin are licensed under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 (CC-BY-NC-SA 4.0). Please review the Hugging Face Model Card for details regarding model usage and limitations.
Contributions are welcome! Please feel free to submit a Pull Request or open an issue for bugs and feature requests.
Special shoutout to Benjamin Hamm who created the first version of this plugin. For questions, issues, or collaborations, please contact:
📧 maximilian.rokuss@dkfz-heidelberg.de / benjamin.hamm@dkfz-heidelberg.de
The remote (client/server) inference mode is inspired by MIC-DKFZ's nnInteractive and napari-nninteractive (Apache-2.0).
Python
100.0%