This project provides a web-based interface for VoxTell, a state-of-the-art 3D vision-language model for medical image segmentation. VoxTell enables natural language-driven anatomical segmentation across CT, PET, and MRI modalities—simply describe what you want to segment in plain English.
VoxTell is a deep learning model that combines 3D image understanding with natural language processing to segment anatomical structures from text prompts. Instead of traditional segmentation tools that require manual annotation or predefined labels, VoxTell accepts prompts like:
"liver", "brain", "left kidney""right lung upper lobe", "L5 vertebra""prostate tumor", "pancreatic head"The model was trained on 158 public datasets with over 62,000 volumetric images, covering brain, thorax, abdomen, pelvis, musculoskeletal structures, and pathological findings.
This implementation includes critical optimizations to run VoxTell on consumer GPUs with limited VRAM (e.g., RTX 3060 12GB, RTX 4060 Ti 16GB):
| Component | Optimization | VRAM Savings |
|---|---|---|
| Text Encoder | Load Qwen3-Embedding-4B in float16 precision | ~15GB → ~7.5GB |
| Memory Allocator | PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True" | Reduces GPU memory fragmentation |
| Sliding Window | perform_everything_on_device = False | Offloads CPU-compatible ops to reduce peak VRAM |
[!IMPORTANT] Hardware Requirements
- Minimum: 12GB VRAM (tested on RTX 3080 12GB version)
- Recommended: 16GB+ VRAM for larger volumes
- CPU: Multi-core recommended for sliding window fallback
.nii.gz) or RTStruct files.nii, .nii.gz, and DICOM upload/processinggit clone https://github.com/gomesgustavoo/voxtell-web-plugin.git
cd voxtell-web-plugin
Create and activate a Conda environment:
conda create -n voxtell python=3.12
conda activate voxtell
Install PyTorch (adjust CUDA version as needed):
[!WARNING] PyTorch 2.9.0 Compatibility Issue
There is a known OOM bug in PyTorch 2.9.0 affecting 3D convolutions. Use PyTorch 2.8.0 or earlier until resolved (PyTorch Issue #166122).
pip install torch==2.8.0 torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu126
Install VoxTell dependencies:
pip install -e .
Download the VoxTell model (see download_model.py):
# huggingface_hub is required
python download_model.py
Navigate to the frontend directory and install dependencies:
cd frontend
npm install
From the project root:
chmod +x run.sh # only needed once
./run.sh
The script starts the backend (:8000) and the frontend (:5173) together. Open http://localhost:5173 in your browser. Press Ctrl+C to stop both servers.
[!NOTE] The backend loads the VoxTell model at startup, which takes 30–60 seconds. The frontend will be immediately available, but inference requests will fail until the backend prints
Model loaded successfully.
.nii, .nii.gz file or a DICOM series"liver", "prostate tumor", "left kidney").nii.gz or RTStruct files for further analysis| Prompt | Target Structure |
|---|---|
"lungs" | Both lungs |
"right kidney" | Right kidney only |
"prostate tumor" | Clinical target volume |
"L4 vertebra" | L4 vertebral body |
"thoracic aorta" | Descending thoracic aorta |
[!TIP] Image Orientation
VoxTell requires images in RAS orientation for correct left/right anatomical localization. If segmentations appear mirrored or incorrect (e.g., liver segments spleen instead), verify your NIfTI metadata.
┌─────────────────┐
│ React Frontend │ ← User uploads .nii.gz + text prompt
└────────┬────────┘
│ HTTP POST /predict
│
┌────────▼────────┐
│ FastAPI Server │ ← Receives file + prompt
└────────┬────────┘
│
┌────────▼────────┐
│ VoxTell Model │ ← 3D vision-language inference
│ - Image Encoder│ (optimized for low VRAM)
│ - Text Encoder │
│ - Fusion Decoder│
└────────┬────────┘
│
└─────────► Returns .nii.gz segmentation mask
| File | Description |
|---|---|
backend/server.py | FastAPI inference server with /predict endpoint |
voxtell/inference/predictor.py | VoxTell inference engine with FP16 text encoder |
frontend/src/App.tsx | Main React application with file upload and API logic |
frontend/src/components/Viewer.tsx | NiiVue-based 3D medical image viewer |
The following modifications were made to the original VoxTell implementation:
1. Text Encoder Precision Reduction
voxtell/inference/predictor.py
# Load Qwen3-Embedding-4B in float16 instead of float32
self.text_encoder = AutoModel.from_pretrained(
"Qwen/Qwen3-Embedding-4B",
torch_dtype=torch.float16 # ← Halves VRAM usage
)
2. Memory Fragmentation Mitigation
backend/server.py
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
3. Sliding Window CPU Offload
backend/server.py
predictor.perform_everything_on_device = False
# Allows nnUNet to use CPU for preprocessing/postprocessing
python download_model.py first — the model files must exist under models/voxtell_v1.1/.VoxTell requires images in RAS orientation for correct left/right anatomical localization. If results seem flipped (e.g., liver appears on the wrong side), verify your NIfTI file metadata and orientation.
This project inherits the Apache 2.0 License from the original VoxTell repository. See LICENSE for details.
This project builds upon VoxTell by the Division of Medical Image Computing (MIC), German Cancer Research Center (DKFZ):
Rokuss et al. (2025). VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentation. arXiv:2511.11450.
@misc{rokuss2025voxtell,
title={VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentation},
author={Maximilian Rokuss and Moritz Langenberg and Yannick Kirchhoff and Fabian Isensee and Benjamin Hamm and Constantin Ulrich and Sebastian Regnery and Lukas Bauer and Efthimios Katsigiannopulos and Tobias Norajitra and Klaus Maier-Hein},
year={2025},
eprint={2511.11450},
archivePrefix={arXiv}
}
Special thanks to the authors for open-sourcing this amazing work.
For questions about this web interface implementation, contact:
📧 https://www.linkedin.com/in/gustavoogomesss/
For questions about the original VoxTell model, contact:
📧 maximilian.rokuss@dkfz-heidelberg.de / moritz.langenberg@dkfz-heidelberg.de
Python
57.8%
TypeScript
37.9%
Shell
2.8%
This project provides a web-based interface for VoxTell, a state-of-the-art 3D vision-language model for medical image segmentation. VoxTell enables natural language-driven anatomical segmentation across CT, PET, and MRI modalities—simply describe what you want to segment in plain English.
VoxTell is a deep learning model that combines 3D image understanding with natural language processing to segment anatomical structures from text prompts. Instead of traditional segmentation tools that require manual annotation or predefined labels, VoxTell accepts prompts like:
"liver", "brain", "left kidney""right lung upper lobe", "L5 vertebra""prostate tumor", "pancreatic head"The model was trained on 158 public datasets with over 62,000 volumetric images, covering brain, thorax, abdomen, pelvis, musculoskeletal structures, and pathological findings.
This implementation includes critical optimizations to run VoxTell on consumer GPUs with limited VRAM (e.g., RTX 3060 12GB, RTX 4060 Ti 16GB):
| Component | Optimization | VRAM Savings |
|---|---|---|
| Text Encoder | Load Qwen3-Embedding-4B in float16 precision | ~15GB → ~7.5GB |
| Memory Allocator | PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True" | Reduces GPU memory fragmentation |
| Sliding Window | perform_everything_on_device = False | Offloads CPU-compatible ops to reduce peak VRAM |
[!IMPORTANT] Hardware Requirements
- Minimum: 12GB VRAM (tested on RTX 3080 12GB version)
- Recommended: 16GB+ VRAM for larger volumes
- CPU: Multi-core recommended for sliding window fallback
.nii.gz) or RTStruct files.nii, .nii.gz, and DICOM upload/processinggit clone https://github.com/gomesgustavoo/voxtell-web-plugin.git
cd voxtell-web-plugin
Create and activate a Conda environment:
conda create -n voxtell python=3.12
conda activate voxtell
Install PyTorch (adjust CUDA version as needed):
[!WARNING] PyTorch 2.9.0 Compatibility Issue
There is a known OOM bug in PyTorch 2.9.0 affecting 3D convolutions. Use PyTorch 2.8.0 or earlier until resolved (PyTorch Issue #166122).
pip install torch==2.8.0 torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu126
Install VoxTell dependencies:
pip install -e .
Download the VoxTell model (see download_model.py):
# huggingface_hub is required
python download_model.py
Navigate to the frontend directory and install dependencies:
cd frontend
npm install
From the project root:
chmod +x run.sh # only needed once
./run.sh
The script starts the backend (:8000) and the frontend (:5173) together. Open http://localhost:5173 in your browser. Press Ctrl+C to stop both servers.
[!NOTE] The backend loads the VoxTell model at startup, which takes 30–60 seconds. The frontend will be immediately available, but inference requests will fail until the backend prints
Model loaded successfully.
.nii, .nii.gz file or a DICOM series"liver", "prostate tumor", "left kidney").nii.gz or RTStruct files for further analysis| Prompt | Target Structure |
|---|---|
"lungs" | Both lungs |
"right kidney" | Right kidney only |
"prostate tumor" | Clinical target volume |
"L4 vertebra" | L4 vertebral body |
"thoracic aorta" | Descending thoracic aorta |
[!TIP] Image Orientation
VoxTell requires images in RAS orientation for correct left/right anatomical localization. If segmentations appear mirrored or incorrect (e.g., liver segments spleen instead), verify your NIfTI metadata.
┌─────────────────┐
│ React Frontend │ ← User uploads .nii.gz + text prompt
└────────┬────────┘
│ HTTP POST /predict
│
┌────────▼────────┐
│ FastAPI Server │ ← Receives file + prompt
└────────┬────────┘
│
┌────────▼────────┐
│ VoxTell Model │ ← 3D vision-language inference
│ - Image Encoder│ (optimized for low VRAM)
│ - Text Encoder │
│ - Fusion Decoder│
└────────┬────────┘
│
└─────────► Returns .nii.gz segmentation mask
| File | Description |
|---|---|
backend/server.py | FastAPI inference server with /predict endpoint |
voxtell/inference/predictor.py | VoxTell inference engine with FP16 text encoder |
frontend/src/App.tsx | Main React application with file upload and API logic |
frontend/src/components/Viewer.tsx | NiiVue-based 3D medical image viewer |
The following modifications were made to the original VoxTell implementation:
1. Text Encoder Precision Reduction
voxtell/inference/predictor.py
# Load Qwen3-Embedding-4B in float16 instead of float32
self.text_encoder = AutoModel.from_pretrained(
"Qwen/Qwen3-Embedding-4B",
torch_dtype=torch.float16 # ← Halves VRAM usage
)
2. Memory Fragmentation Mitigation
backend/server.py
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"
3. Sliding Window CPU Offload
backend/server.py
predictor.perform_everything_on_device = False
# Allows nnUNet to use CPU for preprocessing/postprocessing
python download_model.py first — the model files must exist under models/voxtell_v1.1/.VoxTell requires images in RAS orientation for correct left/right anatomical localization. If results seem flipped (e.g., liver appears on the wrong side), verify your NIfTI file metadata and orientation.
This project inherits the Apache 2.0 License from the original VoxTell repository. See LICENSE for details.
This project builds upon VoxTell by the Division of Medical Image Computing (MIC), German Cancer Research Center (DKFZ):
Rokuss et al. (2025). VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentation. arXiv:2511.11450.
@misc{rokuss2025voxtell,
title={VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentation},
author={Maximilian Rokuss and Moritz Langenberg and Yannick Kirchhoff and Fabian Isensee and Benjamin Hamm and Constantin Ulrich and Sebastian Regnery and Lukas Bauer and Efthimios Katsigiannopulos and Tobias Norajitra and Klaus Maier-Hein},
year={2025},
eprint={2511.11450},
archivePrefix={arXiv}
}
Special thanks to the authors for open-sourcing this amazing work.
For questions about this web interface implementation, contact:
📧 https://www.linkedin.com/in/gustavoogomesss/
For questions about the original VoxTell model, contact:
📧 maximilian.rokuss@dkfz-heidelberg.de / moritz.langenberg@dkfz-heidelberg.de
Python
57.8%
TypeScript
37.9%
Shell
2.8%