ZhaoxuLi123/SpecDETR

SpecDETR: A Transformer-based Hyperspectral Point Object Detection Network

Python

54

8 commits

updated Jun 29, 2025

See the code

README

SpecDETR

This is the official repository for "SpecDETR: A transformer-based hyperspectral point object detection network", and it is also part of the open-source hyperspectral object detection toolbox HODToolbox.

Paper link: ISPRS P&RS or arXiv.


Contributions

  1. We verified that existing visual objec detection networks possess the capability for spatial-spectral integrated semantic representation of tiny objects. These networks can effectively detect subpixel-level objects in hyperspectral images, outperforming pixel-wise hyperspectral target detection methods.

  2. We developed the first multi-class hyperspectral tiny object detectio benchmark dataset, SPOD. Based on this dataset, a comprehensive evaluation was conducted on the performance of mainstream visual object detection networks and hyperspectral target detection methods on the hyperspectral tiny object detection task.

  3. We proposed SpecDETR, an innovate tiny object detection network that introduces a self-excited mechanism to enhance feature extraction and designs a streamlined yet efficient novel DETR decoder architecture tailored for tiny object characteristics. Experiments demonstrate that SpecDETR significantly outperforms existing approaches in the hyperspectral tiny object detection task.

  4. We developed an open-source hyperspectral object detection toolbox, HODToolbox, facilitating the paradigm shift from traditional pixel-level hyperspectral target detection to hyperspectral object detection. The toolbox integrates the following core functionalities:

    • Convert traditional hyperspectral target detection datasets into object detection datasets, and use single-target prior spectra to generate large-scale training image sets for object detection networks training.
    • Train and test mainstream visual object detection networks on hyperspectral object detection datasets.
    • Quantitative evaluation and visual analysis of detection results.

News & Updates

  • June 29, 2025: We make the following updates:

    1. Open-source the simulated training sets for three public HTD datasets: Avon, SanDiego, and MUUFLGulfport.
    2. Add support for the infrared video satellite flying airplane detection dataset IRAir.
    3. Provide pre-trained SpecDETR models for hyperspectral tiny object detection on three public datasets (Avon, SanDiego, MUUFLGulfport) and single-frame infrared tiny object detection on IRAir dataset.
    4. Released the companion toolbox HODToolbox
  • May 08, 2025: We are pleased to announce that our work SpecDETR has been accepted by ISPRS Journal of Photogrammetry and Remote Sensing!


Abstract

Hyperspectral target detection (HTD) aims to identify specific materials based on spectral information in hyperspectral imagery and can detect extremely small-sized objects, some of which occupy a smaller than one-pixel area. However, existing HTD methods are developed based on per-pixel binary classification, neglecting the three-dimensional cube structure of hyperspectral images (HSIs) that integrates both spatial and spectral dimensions. The synergistic existence of spatial and spectral features in HSIs enable objects to simultaneously exhibit both, yet the per-pixel HTD framework limits the joint expression of these features. In this paper, we rethink HTD from the perspective of spatial–spectral synergistic representation and propose hyperspectral point object detection as an innovative task framework. We introduce SpecDETR, the first specialized network for hyperspectral multi-class point object detection, which eliminates dependence on pre-trained backbone networks commonly required by vision-based object detectors. SpecDETR uses a multi-layer Transformer encoder with self-excited subpixel-scale attention modules to directly extract deep spatial–spectral joint features from hyperspectral cubes. During feature extraction, we introduce a self-excited mechanism to enhance object features through self-excited amplification, thereby accelerating network convergence. Additionally, SpecDETR regards point object detection as a one-to-many set prediction problem, thereby achieving a concise and efficient DETR decoder that surpasses the state-of-the-art (SOTA) DETR decoder. We develop a simulated hyperSpectral Point Object Detection benchmark termed SPOD, and for the first time, evaluate and compare the performance of visual object detection networks and HTD methods on hyperspectral point object detection. Extensive experiments demonstrate that our proposed SpecDETR outperforms SOTA visual object detection networks and HTD methods.


Installation

The project is built upon mmdetection 3.0.0 and runs on the Ubuntu system.

  1. Create a new conda environment and activate the environment. Requires Python>=3.7.

  2. Install Pytorch. Requires torch>=1.8.

  3. Install mmengine and mmcv.

    pip install mmengine==0.7.3
    pip install mmcv==2.0.0
    
  4. Clone this repository:

    SpecDETR_ROOT=/path/to/clone/SpecDETR
    git clone https://github.com/ZhaoxuLi123/SpecDETR $SpecDETR_ROOT
    
  5. Compile and install mmdet. If mmdet>=3.0.0 is already installed in the environment, skip this step.

    cd $SpecDETR_ROOT
    pip install -v -e .
    

Dataset Preparation

SPOD Dataset

  1. Download SPOD_30b_8c.zip from Baidu Drive (key: 2789) or OneDrive.

  2. Unzip SPOD_30b_8c.zip locally.

  3. Update the dataset configuration files configs/_base_/datasetshsi_detection.py and configs/VisualObjectDetectionNetwork/_base_/datasets/hsi_detection4x.py:

      data_root = 'Dataset_Root/SPOD_30b_8c/'  # ← Update to the local path
    

Avon Dataset, SanDiego Dataset, and MUUFLGulfport Dataset

  1. Download dataset zip files:

  2. Unzip dataset zip files locally.

  3. Update the dataset configuration files configs/_base_/datasetshsi_detection.py and configs/VisualObjectDetectionNetwork/_base_/datasets/hsi_detection4x.py:

      data_root = 'Dataset_path/'  # ← Update to the local path of the dataset.
    

Other HTD Dataset

We provide conversion scripts from HTD Dataset to HOD Dataset in HODToolbox.

For details on adding new HOD datasets, please refer to the Training and Inference of the HOD Task section in HODToolbox.

IRAir Dataset

We will release the IRAir dataset soon. Please stay tuned on the project website: IRAir

The IRAir dataset configuration file is located at:
configs/_base_/datasets/irair_real_label.py


Model Zoo

We provide pre-trained SpecDETR models for SPOD Dataset, Avon Dataset, SanDiego Dataset, MUUFLGulfport Dataset, and IRAir Dataset.

MethodDatasetConfigurationBaidu DriveOneDrive
SpecDETRSPOD DatasetLinkLink (key: 2789)Link
SpecDETRAvon DatasetLinkLink (key: 2789)Link
SpecDETRSanDiego DatasetLinkLink (key: 2789)Link
SpecDETRMUUFLGulfport DatasetLinkLink (key: 2789)Link
SpecDETRIRAir DatasetLinkLink (key: 2789)Link

Model Training

SpecDETR

To train SpecDETR on the SPOD dataset, execute either of the following commands:

python train.py --dataset SPOD

or

python train.py --config ./configs/specdetr/SpecDETR_SPOD_100e.py --work-dir ./work_dirs/SpecDETR/SPOD/

Note: Although we have fixed all random seeds, there may still be slight differences in AP performance each time you run. This difference originates from the underlying mechanism of CUDA.

Other Model

We provide configuration files for existing visual object detection networks on the SPOD dataset in the configs/VisualObjectDetectionNetwork directory. Execute the following command to train these models:

python train.py --config ./configs/VisualObjectDetectionNetwork/dino-5scale_swin-l-100e_hsi4x.py --work-dir ./work_dirs/dino-5scale_swin-l/SPOD/

Model Evaluation

Use SPOD dataset as example.

  1. Model Inference
    Run the following command to obtain SpecDETR's inference results on the SPOD test set:

    python test.py --dataset SPOD  
    

    or

    python test.py --config ./configs/specdetr/SpecDETR_SPOD_100e.py \  
                  --work-dir ./work_dirs/SpecDETR/SPOD/ \  
                  --checkpoint ./work_dirs/SpecDETR/SpecDETR_SPOD_100e.pth \  
                  --out ./work_dirs/SpecDETR/SPOD/SpecDETR.pkl  
    
  2. Detection Accuracy Evaluation
    Please refer to the Quantitative Evaluation of Results section in HODToolbox.

  3. Inference Speed Evaluation
    Execute the following command to evaluate SpecDETR's inference speed on the SPOD test set:

    python benchmark.py --config ./configs/specdetr/SpecDETR_SPOD_100e.py \  
                       --checkpoint ./work_dirs/SpecDETR/SpecDETR_SPOD_100e.pth  
    
  4. FLOPs Calculation
    Run the following command to compute FLOPs:

    python get_flops.py --config ./configs/specdetr/SpecDETR_SPOD_100e.py  
    

Benchmark

Partial Quantitative Results of Visual Object Detection Networks on SPOD Dataset

MethodBackboneImage SizemAP50:95mAP25mAP50mAP75FLOPsParams
Faster R-CNNResNet50x40.1970.3770.3740.17968.8G41.5M
Faster R-CNNRegNetXx40.2270.3790.3780.24257.7G31.6M
Faster R-CNNResNeSt50x40.2460.3160.3160.277185.1G44.6M
Faster R-CNNResNeXt101x40.2200.3680.3660.231128.4G99.4M
Faster R-CNNHRNetx40.3200.4040.4020.345104.4G63.2M
TOODResNeXt101x40.3040.4640.4400.303114.3G97.7M
CentripetalNetHourglassNet104x40.6950.8290.8050.673501.3G205.9M
CornerNetHourglassNet104x40.6260.7360.7120.609462.6G201.1M
RepPointsResNet50x40.2070.6910.5720.07454.1G36.9M
RepPointsResNeXt101x40.4850.8060.7900.54075.0G58.1M
RetinaNetEfficientNetx40.4620.8360.8110.46636.1G18.5M
RetinaNetPVTv2-B3x40.4260.7570.7340.44271.3G52.4M
DeformableDETRResNet50x40.2310.6920.5600.14758.7G41.2M
DINOResNet50x40.1680.4910.4180.09786.3G47.6M
DINOSwin-Lx40.7570.8520.8420.764203.9G218.3M
SpecDETR--x10.8560.9380.9300.863139.7G16.1M

Partial Quantitative Results of HTD Methods on SPOD Dataset

MethodmAP50:95mAP25mAP50mAP75
ASD0.1820.2860.2600.182
CEM0.0400.1220.0750.035
CRBBH0.0360.1290.0830.028
CSRBBH0.0340.1160.0760.028
HSS0.0730.3030.1790.058
IRN0.0000.0020.0010.000
KMSD0.1080.2850.2070.095
KOSP0.0170.0830.0440.014
KSMF0.0030.0150.0090.002
KTCIMF0.0010.0080.0020.000
LSSA0.0410.0930.0710.037
MSD0.2480.5210.4020.228
OSP0.0310.1080.0630.027
SMF0.0030.0160.0080.002
SRBBH0.0190.0920.0510.013
SRBBH_PBD0.0130.0880.0380.007
TCIMF0.0090.0610.0250.007
TSTTD0.0440.0570.0550.043
SpecDETR0.8560.9380.9300.863

Citation

If the work or the code is helpful, please cite the paper:

@article{li2025specdetr,
  title={SpecDETR: A transformer-based hyperspectral point object detection network},
  author={Li, Zhaoxu and An, Wei and Guo, Gaowei and Wang, Longguang and Wang, Yingqian and Lin, Zaiping},
  journal={ISPRS Journal of Photogrammetry and Remote Sensing},
  volume={226},
  pages={221--246},
  year={2025},
  publisher={Elsevier}
}

Contact

For further questions or details, please directly reach out to lizhaoxu@nudt.edu.cn.

ZhaoxuLi123/SpecDETR

SpecDETR: A Transformer-based Hyperspectral Point Object Detection Network

Python

54

8 commits

updated Jun 29, 2025

See the code

README

SpecDETR

This is the official repository for "SpecDETR: A transformer-based hyperspectral point object detection network", and it is also part of the open-source hyperspectral object detection toolbox HODToolbox.

Paper link: ISPRS P&RS or arXiv.


Contributions

  1. We verified that existing visual objec detection networks possess the capability for spatial-spectral integrated semantic representation of tiny objects. These networks can effectively detect subpixel-level objects in hyperspectral images, outperforming pixel-wise hyperspectral target detection methods.

  2. We developed the first multi-class hyperspectral tiny object detectio benchmark dataset, SPOD. Based on this dataset, a comprehensive evaluation was conducted on the performance of mainstream visual object detection networks and hyperspectral target detection methods on the hyperspectral tiny object detection task.

  3. We proposed SpecDETR, an innovate tiny object detection network that introduces a self-excited mechanism to enhance feature extraction and designs a streamlined yet efficient novel DETR decoder architecture tailored for tiny object characteristics. Experiments demonstrate that SpecDETR significantly outperforms existing approaches in the hyperspectral tiny object detection task.

  4. We developed an open-source hyperspectral object detection toolbox, HODToolbox, facilitating the paradigm shift from traditional pixel-level hyperspectral target detection to hyperspectral object detection. The toolbox integrates the following core functionalities:

    • Convert traditional hyperspectral target detection datasets into object detection datasets, and use single-target prior spectra to generate large-scale training image sets for object detection networks training.
    • Train and test mainstream visual object detection networks on hyperspectral object detection datasets.
    • Quantitative evaluation and visual analysis of detection results.

News & Updates

  • June 29, 2025: We make the following updates:

    1. Open-source the simulated training sets for three public HTD datasets: Avon, SanDiego, and MUUFLGulfport.
    2. Add support for the infrared video satellite flying airplane detection dataset IRAir.
    3. Provide pre-trained SpecDETR models for hyperspectral tiny object detection on three public datasets (Avon, SanDiego, MUUFLGulfport) and single-frame infrared tiny object detection on IRAir dataset.
    4. Released the companion toolbox HODToolbox
  • May 08, 2025: We are pleased to announce that our work SpecDETR has been accepted by ISPRS Journal of Photogrammetry and Remote Sensing!


Abstract

Hyperspectral target detection (HTD) aims to identify specific materials based on spectral information in hyperspectral imagery and can detect extremely small-sized objects, some of which occupy a smaller than one-pixel area. However, existing HTD methods are developed based on per-pixel binary classification, neglecting the three-dimensional cube structure of hyperspectral images (HSIs) that integrates both spatial and spectral dimensions. The synergistic existence of spatial and spectral features in HSIs enable objects to simultaneously exhibit both, yet the per-pixel HTD framework limits the joint expression of these features. In this paper, we rethink HTD from the perspective of spatial–spectral synergistic representation and propose hyperspectral point object detection as an innovative task framework. We introduce SpecDETR, the first specialized network for hyperspectral multi-class point object detection, which eliminates dependence on pre-trained backbone networks commonly required by vision-based object detectors. SpecDETR uses a multi-layer Transformer encoder with self-excited subpixel-scale attention modules to directly extract deep spatial–spectral joint features from hyperspectral cubes. During feature extraction, we introduce a self-excited mechanism to enhance object features through self-excited amplification, thereby accelerating network convergence. Additionally, SpecDETR regards point object detection as a one-to-many set prediction problem, thereby achieving a concise and efficient DETR decoder that surpasses the state-of-the-art (SOTA) DETR decoder. We develop a simulated hyperSpectral Point Object Detection benchmark termed SPOD, and for the first time, evaluate and compare the performance of visual object detection networks and HTD methods on hyperspectral point object detection. Extensive experiments demonstrate that our proposed SpecDETR outperforms SOTA visual object detection networks and HTD methods.


Installation

The project is built upon mmdetection 3.0.0 and runs on the Ubuntu system.

  1. Create a new conda environment and activate the environment. Requires Python>=3.7.

  2. Install Pytorch. Requires torch>=1.8.

  3. Install mmengine and mmcv.

    pip install mmengine==0.7.3
    pip install mmcv==2.0.0
    
  4. Clone this repository:

    SpecDETR_ROOT=/path/to/clone/SpecDETR
    git clone https://github.com/ZhaoxuLi123/SpecDETR $SpecDETR_ROOT
    
  5. Compile and install mmdet. If mmdet>=3.0.0 is already installed in the environment, skip this step.

    cd $SpecDETR_ROOT
    pip install -v -e .
    

Dataset Preparation

SPOD Dataset

  1. Download SPOD_30b_8c.zip from Baidu Drive (key: 2789) or OneDrive.

  2. Unzip SPOD_30b_8c.zip locally.

  3. Update the dataset configuration files configs/_base_/datasetshsi_detection.py and configs/VisualObjectDetectionNetwork/_base_/datasets/hsi_detection4x.py:

      data_root = 'Dataset_Root/SPOD_30b_8c/'  # ← Update to the local path
    

Avon Dataset, SanDiego Dataset, and MUUFLGulfport Dataset

  1. Download dataset zip files:

  2. Unzip dataset zip files locally.

  3. Update the dataset configuration files configs/_base_/datasetshsi_detection.py and configs/VisualObjectDetectionNetwork/_base_/datasets/hsi_detection4x.py:

      data_root = 'Dataset_path/'  # ← Update to the local path of the dataset.
    

Other HTD Dataset

We provide conversion scripts from HTD Dataset to HOD Dataset in HODToolbox.

For details on adding new HOD datasets, please refer to the Training and Inference of the HOD Task section in HODToolbox.

IRAir Dataset

We will release the IRAir dataset soon. Please stay tuned on the project website: IRAir

The IRAir dataset configuration file is located at:
configs/_base_/datasets/irair_real_label.py


Model Zoo

We provide pre-trained SpecDETR models for SPOD Dataset, Avon Dataset, SanDiego Dataset, MUUFLGulfport Dataset, and IRAir Dataset.

MethodDatasetConfigurationBaidu DriveOneDrive
SpecDETRSPOD DatasetLinkLink (key: 2789)Link
SpecDETRAvon DatasetLinkLink (key: 2789)Link
SpecDETRSanDiego DatasetLinkLink (key: 2789)Link
SpecDETRMUUFLGulfport DatasetLinkLink (key: 2789)Link
SpecDETRIRAir DatasetLinkLink (key: 2789)Link

Model Training

SpecDETR

To train SpecDETR on the SPOD dataset, execute either of the following commands:

python train.py --dataset SPOD

or

python train.py --config ./configs/specdetr/SpecDETR_SPOD_100e.py --work-dir ./work_dirs/SpecDETR/SPOD/

Note: Although we have fixed all random seeds, there may still be slight differences in AP performance each time you run. This difference originates from the underlying mechanism of CUDA.

Other Model

We provide configuration files for existing visual object detection networks on the SPOD dataset in the configs/VisualObjectDetectionNetwork directory. Execute the following command to train these models:

python train.py --config ./configs/VisualObjectDetectionNetwork/dino-5scale_swin-l-100e_hsi4x.py --work-dir ./work_dirs/dino-5scale_swin-l/SPOD/

Model Evaluation

Use SPOD dataset as example.

  1. Model Inference
    Run the following command to obtain SpecDETR's inference results on the SPOD test set:

    python test.py --dataset SPOD  
    

    or

    python test.py --config ./configs/specdetr/SpecDETR_SPOD_100e.py \  
                  --work-dir ./work_dirs/SpecDETR/SPOD/ \  
                  --checkpoint ./work_dirs/SpecDETR/SpecDETR_SPOD_100e.pth \  
                  --out ./work_dirs/SpecDETR/SPOD/SpecDETR.pkl  
    
  2. Detection Accuracy Evaluation
    Please refer to the Quantitative Evaluation of Results section in HODToolbox.

  3. Inference Speed Evaluation
    Execute the following command to evaluate SpecDETR's inference speed on the SPOD test set:

    python benchmark.py --config ./configs/specdetr/SpecDETR_SPOD_100e.py \  
                       --checkpoint ./work_dirs/SpecDETR/SpecDETR_SPOD_100e.pth  
    
  4. FLOPs Calculation
    Run the following command to compute FLOPs:

    python get_flops.py --config ./configs/specdetr/SpecDETR_SPOD_100e.py  
    

Benchmark

Partial Quantitative Results of Visual Object Detection Networks on SPOD Dataset

MethodBackboneImage SizemAP50:95mAP25mAP50mAP75FLOPsParams
Faster R-CNNResNet50x40.1970.3770.3740.17968.8G41.5M
Faster R-CNNRegNetXx40.2270.3790.3780.24257.7G31.6M
Faster R-CNNResNeSt50x40.2460.3160.3160.277185.1G44.6M
Faster R-CNNResNeXt101x40.2200.3680.3660.231128.4G99.4M
Faster R-CNNHRNetx40.3200.4040.4020.345104.4G63.2M
TOODResNeXt101x40.3040.4640.4400.303114.3G97.7M
CentripetalNetHourglassNet104x40.6950.8290.8050.673501.3G205.9M
CornerNetHourglassNet104x40.6260.7360.7120.609462.6G201.1M
RepPointsResNet50x40.2070.6910.5720.07454.1G36.9M
RepPointsResNeXt101x40.4850.8060.7900.54075.0G58.1M
RetinaNetEfficientNetx40.4620.8360.8110.46636.1G18.5M
RetinaNetPVTv2-B3x40.4260.7570.7340.44271.3G52.4M
DeformableDETRResNet50x40.2310.6920.5600.14758.7G41.2M
DINOResNet50x40.1680.4910.4180.09786.3G47.6M
DINOSwin-Lx40.7570.8520.8420.764203.9G218.3M
SpecDETR--x10.8560.9380.9300.863139.7G16.1M

Partial Quantitative Results of HTD Methods on SPOD Dataset

MethodmAP50:95mAP25mAP50mAP75
ASD0.1820.2860.2600.182
CEM0.0400.1220.0750.035
CRBBH0.0360.1290.0830.028
CSRBBH0.0340.1160.0760.028
HSS0.0730.3030.1790.058
IRN0.0000.0020.0010.000
KMSD0.1080.2850.2070.095
KOSP0.0170.0830.0440.014
KSMF0.0030.0150.0090.002
KTCIMF0.0010.0080.0020.000
LSSA0.0410.0930.0710.037
MSD0.2480.5210.4020.228
OSP0.0310.1080.0630.027
SMF0.0030.0160.0080.002
SRBBH0.0190.0920.0510.013
SRBBH_PBD0.0130.0880.0380.007
TCIMF0.0090.0610.0250.007
TSTTD0.0440.0570.0550.043
SpecDETR0.8560.9380.9300.863

Citation

If the work or the code is helpful, please cite the paper:

@article{li2025specdetr,
  title={SpecDETR: A transformer-based hyperspectral point object detection network},
  author={Li, Zhaoxu and An, Wei and Guo, Gaowei and Wang, Longguang and Wang, Yingqian and Lin, Zaiping},
  journal={ISPRS Journal of Photogrammetry and Remote Sensing},
  volume={226},
  pages={221--246},
  year={2025},
  publisher={Elsevier}
}

Contact

For further questions or details, please directly reach out to lizhaoxu@nudt.edu.cn.

Languages

Python

100.0%