PSRben/VisionHOPE

Official PyTorch implementation of VisionHOPE: Visual Backbones as Self-Modifying Learning Systems.

Python

58

9 commits

updated Sep 29, 2026

See the code

README

VisionHOPE

Visual Backbones as Self-Modifying Learning Systems

Siran Peng* · Tianshuo Zhang* · Tianyu Fu · Weisong Zhao · Haoyuan Zhang
Jiankuo Zhao · Minghui Wu · Ping Jiang · Xiangyu Zhu · Chenxu Zhao† · Zhen Lei†

* Equal contribution.   †Corresponding authors.

arXiv: 2609.33325 PyTorch 2.1 Hugging Face Models License: MIT

📄 Paper · 🤗 Hugging Face Paper · Updates · Overview · Results & checkpoints · Citation

Installation · Models · SRNL & operator · Fast inference · Datasets · Training · Evaluation · Efficiency

PyTorch implementation of VisionHOPE, a visual backbone based on self-referential nested learning (SRNL). This repository provides ImageNet classification, COCO detection and instance segmentation, and ADE20K semantic segmentation, along with standalone SRNL and VisionHOPE operator modules.

Pretrained weights: GitHub Releases · Hugging Face · Baidu Netdisk (visionhope_weights, access code: 8866).

Progression from CNNs, ViTs, SSMs, and TTT to nested learning: VisionHOPE adapts both its memory and the rule that updates it within an image.

From adaptive visual computation to self-modifying learning within an image.

Updates

Overview

VisionHOPE builds on the self-referential construction in Nested Learning: The Illusion of Deep Learning Architectures. It lets what the model remembers and how it learns evolve together while processing an image. The VisionHOPE operator combines:

  1. Coupled memories. Five memories store content, generate key and value representations, and govern learning rate and retention. They co-evolve as visual context accumulates.
  2. Stable updates. A soft injection cap and a spectral clamp provide non-expansion guarantees for both token-wise and chunk-wise memory recurrences.
  3. Spatially aligned scans. Four directional scans use row- and column-aligned chunks, followed by spatial restoration and learned channel-wise fusion.

VisionHOPE operator and residual block: four directional SRNL scans with five coupled memories, stability-matched step-size control, and channel-wise output fusion.

The VisionHOPE operator (left) and residual block (right). Click the figure to inspect the full resolution.

The hierarchical Tiny / Small / Base backbones support classification and dense prediction. For use in other architectures, see the standalone SRNL, operator, and block interfaces.

Installation

Use Linux with an NVIDIA GPU and the CUDA toolkit, including nvcc and a C++ compiler. Python 3.10 is recommended for all tasks; classification also supports Python 3.11. Run the following from the repository root, with CUDA 12.1 installed:

python -m pip install torch==2.1.0 torchvision==0.16.0 \
    --index-url https://download.pytorch.org/whl/cu121
python -m pip install -r requirements/classification.txt
python -m pip install --no-deps -e .

CUDA extensions compile on first use. Keep model parameters in FP32 and use torch.autocast for mixed precision.

Additional dependencies for COCO and ADE20K

Use a separate Python 3.10 environment for each task, with the PyTorch and CUDA versions above. Install MMCV, then the task dependencies:

python -m pip install --only-binary=mmcv mmcv==2.1.0 \
    -f https://download.openmmlab.com/mmcv/dist/cu121/torch2.1/index.html

For COCO:

python -m pip install -r requirements/detection.txt
python -m pip install --no-deps -e .

For ADE20K:

python -m pip install -r requirements/segmentation.txt
python -m pip install --no-deps -e .

Models

Use these model names and checkpoint filenames for ImageNet-1K classification at 224 × 224. Accuracy and complexity are listed in Main results.

ModelNameCheckpoint filename
VisionHOPE-Tvisionhope_tinyvisionhope_tiny.pth
VisionHOPE-Svisionhope_smallvisionhope_small.pth
VisionHOPE-Bvisionhope_basevisionhope_base.pth

The train/test scripts look for checkpoints in weights/; set WEIGHTS_ROOT to use another directory. COCO checkpoints are named visionhope_<size>_coco_<schedule>.pth (e.g. visionhope_small_coco_3x.pth); ADE20K checkpoints are named visionhope_<size>_ade20k.pth.

import torch
from visionhope.models import create_model

model = create_model(
    "visionhope_small", checkpoint_path="weights/visionhope_small.pth"
).cuda().eval()

with torch.inference_mode():
    logits = model(torch.randn(1, 3, 224, 224, device="cuda"))  # [1, 1000]

Omit checkpoint_path to create a model for training.

SRNL and VisionHOPE operator

SRNL accepts sequences of shape [batch, tokens, dim] and returns the same shape:

import torch
from visionhope.models import SRNL

srnl = SRNL(dim=64, head_dim=16, chunk_size=64).cuda()
x = torch.randn(2, 257, 64, device="cuda", requires_grad=True)
y = srnl(x)
y.square().mean().backward()

Pass srnl(x, queries=q) to supply queries of the same shape as x; otherwise, x is also used as the query. Each forward call starts a new sequence. dim must be divisible by head_dim, which supports 4, 8, 16, 32, and 64.

VisionHOPEOperator and VisionHOPEBlock accept image features in [batch, channels, height, width] format and preserve their shape:

import torch
from visionhope.models import VisionHOPEOperator, VisionHOPEBlock

x = torch.randn(2, 64, 14, 20, device="cuda")
operator = VisionHOPEOperator(dim=64, head_dim=16).cuda()
block = VisionHOPEBlock(dim=64, mixer_dim=32, head_dim=16).cuda()

y = operator(x)
z = block(x)

mixer_dim sets the block's internal width. All three components require CUDA and support training with autograd.

chunk_size=None uses 64-token chunks for standalone SRNL and row/column chunks for image operators and blocks. You can also specify a positive chunk length; rectangular feature maps and sequences with an incomplete final chunk are supported.

Fast inference

For standalone SRNL and operators, set fast_inference=True, call .eval(), and disable gradients:

import torch
from visionhope.models import SRNL, VisionHOPEOperator

srnl = SRNL(dim=64, fast_inference=True).cuda().eval()
operator = VisionHOPEOperator(dim=64, fast_inference=True).cuda().eval()

with torch.inference_mode():
    y = srnl(torch.randn(2, 257, 64, device="cuda"))
    z = operator(torch.randn(2, 64, 14, 20, device="cuda"))

The same flag can be changed after construction, for example srnl.fast_inference = True. It leaves weights unchanged and does not affect training. If you supply separate SRNL queries, their dtype must match the input for fast inference.

For blocks and complete models, use prepare_for_inference after loading weights:

import torch
from visionhope.models import create_model
from visionhope.inference.prepare import prepare_for_inference

model = create_model(
    "visionhope_small", checkpoint_path="weights/visionhope_small.pth"
).cuda().eval()
prepare_for_inference(model)

with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
    logits = model(torch.randn(1, 3, 224, 224, device="cuda"))

prepare_for_inference fuses model parameters in place for fast inference. Keep an unprepared copy if you also need to continue training. The test scripts below enable fast inference by default; pass --mode ordinary to disable it.

Datasets

Download ImageNet-1K (ILSVRC 2012), COCO 2017, and ADE20K from their official websites.

Arrange the datasets as follows, or point the scripts to equivalent locations with --data-root:

data/
├── imagenet/
│   ├── train/<class_name>/*.JPEG
│   └── val/<class_name>/*.JPEG
├── coco/
│   ├── train2017/
│   ├── val2017/
│   └── annotations/instances_{train,val}2017.json
└── ade20k/
    ├── images/{training,validation}/
    └── annotations/{training,validation}/

You can also set IMAGENET_ROOT, COCO_ROOT, and ADE20K_ROOT to your dataset directories. For ADE20K, use the extracted ADEChallengeData2016 directory as the root.

Training

The scripts in scripts/train contain the training recipes. Matching evaluation scripts are in scripts/test.

TaskSizesScript name
ImageNet-1Ktiny, small, basevisionhope_<size>.sh
COCO, Mask R-CNN 1×tiny, small, basevisionhope_<size>_coco_1x.sh
COCO, Mask R-CNN 3×tiny, smallvisionhope_<size>_coco_3x.sh
ADE20K, UPerNettiny, small, basevisionhope_<size>_ade20k.sh

For example, to train VisionHOPE-S on eight GPUs:

# ImageNet-1K
NPROC_PER_NODE=8 bash scripts/train/visionhope_small.sh \
    --data-root data/imagenet --output outputs/visionhope_small

# COCO
NPROC_PER_NODE=8 bash scripts/train/visionhope_small_coco_1x.sh \
    --data-root data/coco --pretrained weights/visionhope_small.pth

# ADE20K
NPROC_PER_NODE=8 bash scripts/train/visionhope_small_ade20k.sh \
    --data-root data/ade20k --pretrained weights/visionhope_small.pth

Training defaults to eight GPUs. Set NPROC_PER_NODE=1 for a single GPU. --batch-size is per GPU; the learning rate scales with the effective global batch unless --lr is given. Arguments appended to a script override its defaults. COCO and ADE20K training use ImageNet-pretrained backbones.

Outputs go to outputs/train/<script-name>/ unless --output is supplied. Use best.pth for evaluation and training_state.pth to resume training:

bash scripts/train/visionhope_small.sh \
    --resume outputs/visionhope_small/training_state.pth \
    --output outputs/visionhope_small

For argument descriptions, defaults, and choices:

python -m visionhope.tasks.cli train classification --help

Replace classification with detection or segmentation for the other tasks.

Evaluation

Run the matching test script with a checkpoint and dataset:

CUDA_VISIBLE_DEVICES=0 bash scripts/test/visionhope_small.sh \
    --data-root data/imagenet --checkpoint weights/visionhope_small.pth

CUDA_VISIBLE_DEVICES=0 bash scripts/test/visionhope_small_coco_1x.sh \
    --data-root data/coco --checkpoint weights/visionhope_small_coco_1x.pth

CUDA_VISIBLE_DEVICES=0 bash scripts/test/visionhope_small_ade20k.sh \
    --data-root data/ade20k --checkpoint weights/visionhope_small_ade20k.pth

The scripts set the evaluation precision and preprocessing for each task. ImageNet evaluation uses FP32, 224 × 224 inputs, and crop ratio 1.0. Pass --precision fp16 or --precision bf16 to a test script for mixed-precision evaluation. Classification accepts --num-classes to match the checkpoint's output classes and --no-cudnn-benchmark to disable cuDNN algorithm benchmarking. For COCO and ADE20K, use --test-scale W H to change the resize scale while keeping the aspect ratio; defaults are 1333 800 and 2048 512, respectively. ADE20K evaluation defaults to batch size 1; larger batches require images of the same size after preprocessing. Results are saved under outputs/test/<script-name>/; pass --output with a different directory for another evaluation of the same model. Use python -m visionhope.tasks.cli inference classification --help for evaluation options; the other task names work here too.

Main results

The tables below summarize ImageNet-1K, COCO, and ADE20K results for the hierarchical VisionHOPE-T / S / B models. See the paper for experimental details. Pretrained checkpoints are available from GitHub Releases, Hugging Face, and Baidu Netdisk (access code: 8866).

ImageNet-1K classification

Trained on ImageNet-1K; single-crop evaluation at 224 × 224. FLOPs include the classification head.

ModelParams (M)FLOPs (G)Top-1 (%) ↑RecipeCKPT
VisionHOPE-T26.64.984.1Train / EvalGitHub / Hugging Face
VisionHOPE-S52.99.885.2Train / EvalGitHub / Hugging Face
VisionHOPE-B91.217.385.6Train / EvalGitHub / Hugging Face

COCO detection and instance segmentation

Mask R-CNN, initialized with ImageNet-1K-pretrained backbones and evaluated on val2017. FLOPs are for the full detector at 1280 × 800. Box AP and mask AP are reported in percent.

1× schedule

BackboneFLOPs (G)Box AP ↑Mask AP ↑RecipeCKPT
VisionHOPE-T26647.943.1Train / EvalGitHub / Hugging Face
VisionHOPE-S36549.544.2Train / EvalGitHub / Hugging Face
VisionHOPE-B51650.545.0Train / EvalGitHub / Hugging Face

3× schedule

BackboneFLOPs (G)Box AP ↑Mask AP ↑RecipeCKPT
VisionHOPE-T26649.444.1Train / EvalGitHub / Hugging Face
VisionHOPE-S36550.545.0Train / EvalGitHub / Hugging Face

ADE20K semantic segmentation

UPerNet, initialized with ImageNet-1K-pretrained backbones; single-scale validation. Parameters are for the complete model; inference FLOPs are measured at 512 × 2048.

BackboneParams (M)FLOPs (G)mIoU (%) ↑RecipeCKPT
VisionHOPE-T55.494249.4Train / EvalGitHub / Hugging Face
VisionHOPE-S81.8104350.3Train / EvalGitHub / Hugging Face
VisionHOPE-B121.2120051.8Train / EvalGitHub / Hugging Face

See Evaluation for precision and preprocessing, and Efficiency for FLOP-counting and timing conventions.

Efficiency

Measure parameters, inference FLOPs, throughput, and GPU memory without a dataset:

bash run_inference.sh complexity --model visionhope_small
bash run_inference.sh efficiency --model visionhope_small

bash run_inference.sh complexity --task detection --model visionhope_small \
    --checkpoint weights/visionhope_small_coco_1x.pth --shape 1280 800
bash run_inference.sh efficiency --task segmentation --model visionhope_small \
    --checkpoint weights/visionhope_small_ade20k.pth --shape 512 2048

--shape takes height and width. Use --help for batch size, precision, and timing options, and --output result.json to save a report. Throughput and memory measurements use fast inference. They measure the model forward pass on synthetic inputs, excluding data loading, preprocessing, and prediction postprocessing. FLOPs count one multiply-add as one operation. COCO counts depend on the proposals generated for the input; the report lists operations outside the FLOP count. memory.eager reports peak CUDA memory, including model weights and inputs; CUDA Graph memory is reported separately.

Citation

If VisionHOPE is useful for your research, please cite the arXiv preprint:

@article{peng2026visionhope,
  title   = {VisionHOPE: Visual Backbones as Self-Modifying Learning Systems},
  author  = {Peng, Siran and Zhang, Tianshuo and Fu, Tianyu and Zhao, Weisong
             and Zhang, Haoyuan and Zhao, Jiankuo and Wu, Minghui and Jiang, Ping
             and Zhu, Xiangyu and Zhao, Chenxu and Lei, Zhen},
  journal = {arXiv preprint arXiv:2609.33325},
  year    = {2026},
  doi     = {10.48550/arXiv.2609.33325},
  url     = {https://arxiv.org/abs/2609.33325}
}

License

VisionHOPE is released under the MIT License. See Third-party notices for third-party attribution and licenses.

PSRben/VisionHOPE

Official PyTorch implementation of VisionHOPE: Visual Backbones as Self-Modifying Learning Systems.

Python

58

9 commits

updated Sep 29, 2026

See the code

README

VisionHOPE

Visual Backbones as Self-Modifying Learning Systems

Siran Peng* · Tianshuo Zhang* · Tianyu Fu · Weisong Zhao · Haoyuan Zhang
Jiankuo Zhao · Minghui Wu · Ping Jiang · Xiangyu Zhu · Chenxu Zhao† · Zhen Lei†

* Equal contribution.   †Corresponding authors.

arXiv: 2609.33325 PyTorch 2.1 Hugging Face Models License: MIT

📄 Paper · 🤗 Hugging Face Paper · Updates · Overview · Results & checkpoints · Citation

Installation · Models · SRNL & operator · Fast inference · Datasets · Training · Evaluation · Efficiency

PyTorch implementation of VisionHOPE, a visual backbone based on self-referential nested learning (SRNL). This repository provides ImageNet classification, COCO detection and instance segmentation, and ADE20K semantic segmentation, along with standalone SRNL and VisionHOPE operator modules.

Pretrained weights: GitHub Releases · Hugging Face · Baidu Netdisk (visionhope_weights, access code: 8866).

Progression from CNNs, ViTs, SSMs, and TTT to nested learning: VisionHOPE adapts both its memory and the rule that updates it within an image.

From adaptive visual computation to self-modifying learning within an image.

Updates

Overview

VisionHOPE builds on the self-referential construction in Nested Learning: The Illusion of Deep Learning Architectures. It lets what the model remembers and how it learns evolve together while processing an image. The VisionHOPE operator combines:

  1. Coupled memories. Five memories store content, generate key and value representations, and govern learning rate and retention. They co-evolve as visual context accumulates.
  2. Stable updates. A soft injection cap and a spectral clamp provide non-expansion guarantees for both token-wise and chunk-wise memory recurrences.
  3. Spatially aligned scans. Four directional scans use row- and column-aligned chunks, followed by spatial restoration and learned channel-wise fusion.

VisionHOPE operator and residual block: four directional SRNL scans with five coupled memories, stability-matched step-size control, and channel-wise output fusion.

The VisionHOPE operator (left) and residual block (right). Click the figure to inspect the full resolution.

The hierarchical Tiny / Small / Base backbones support classification and dense prediction. For use in other architectures, see the standalone SRNL, operator, and block interfaces.

Installation

Use Linux with an NVIDIA GPU and the CUDA toolkit, including nvcc and a C++ compiler. Python 3.10 is recommended for all tasks; classification also supports Python 3.11. Run the following from the repository root, with CUDA 12.1 installed:

python -m pip install torch==2.1.0 torchvision==0.16.0 \
    --index-url https://download.pytorch.org/whl/cu121
python -m pip install -r requirements/classification.txt
python -m pip install --no-deps -e .

CUDA extensions compile on first use. Keep model parameters in FP32 and use torch.autocast for mixed precision.

Additional dependencies for COCO and ADE20K

Use a separate Python 3.10 environment for each task, with the PyTorch and CUDA versions above. Install MMCV, then the task dependencies:

python -m pip install --only-binary=mmcv mmcv==2.1.0 \
    -f https://download.openmmlab.com/mmcv/dist/cu121/torch2.1/index.html

For COCO:

python -m pip install -r requirements/detection.txt
python -m pip install --no-deps -e .

For ADE20K:

python -m pip install -r requirements/segmentation.txt
python -m pip install --no-deps -e .

Models

Use these model names and checkpoint filenames for ImageNet-1K classification at 224 × 224. Accuracy and complexity are listed in Main results.

ModelNameCheckpoint filename
VisionHOPE-Tvisionhope_tinyvisionhope_tiny.pth
VisionHOPE-Svisionhope_smallvisionhope_small.pth
VisionHOPE-Bvisionhope_basevisionhope_base.pth

The train/test scripts look for checkpoints in weights/; set WEIGHTS_ROOT to use another directory. COCO checkpoints are named visionhope_<size>_coco_<schedule>.pth (e.g. visionhope_small_coco_3x.pth); ADE20K checkpoints are named visionhope_<size>_ade20k.pth.

import torch
from visionhope.models import create_model

model = create_model(
    "visionhope_small", checkpoint_path="weights/visionhope_small.pth"
).cuda().eval()

with torch.inference_mode():
    logits = model(torch.randn(1, 3, 224, 224, device="cuda"))  # [1, 1000]

Omit checkpoint_path to create a model for training.

SRNL and VisionHOPE operator

SRNL accepts sequences of shape [batch, tokens, dim] and returns the same shape:

import torch
from visionhope.models import SRNL

srnl = SRNL(dim=64, head_dim=16, chunk_size=64).cuda()
x = torch.randn(2, 257, 64, device="cuda", requires_grad=True)
y = srnl(x)
y.square().mean().backward()

Pass srnl(x, queries=q) to supply queries of the same shape as x; otherwise, x is also used as the query. Each forward call starts a new sequence. dim must be divisible by head_dim, which supports 4, 8, 16, 32, and 64.

VisionHOPEOperator and VisionHOPEBlock accept image features in [batch, channels, height, width] format and preserve their shape:

import torch
from visionhope.models import VisionHOPEOperator, VisionHOPEBlock

x = torch.randn(2, 64, 14, 20, device="cuda")
operator = VisionHOPEOperator(dim=64, head_dim=16).cuda()
block = VisionHOPEBlock(dim=64, mixer_dim=32, head_dim=16).cuda()

y = operator(x)
z = block(x)

mixer_dim sets the block's internal width. All three components require CUDA and support training with autograd.

chunk_size=None uses 64-token chunks for standalone SRNL and row/column chunks for image operators and blocks. You can also specify a positive chunk length; rectangular feature maps and sequences with an incomplete final chunk are supported.

Fast inference

For standalone SRNL and operators, set fast_inference=True, call .eval(), and disable gradients:

import torch
from visionhope.models import SRNL, VisionHOPEOperator

srnl = SRNL(dim=64, fast_inference=True).cuda().eval()
operator = VisionHOPEOperator(dim=64, fast_inference=True).cuda().eval()

with torch.inference_mode():
    y = srnl(torch.randn(2, 257, 64, device="cuda"))
    z = operator(torch.randn(2, 64, 14, 20, device="cuda"))

The same flag can be changed after construction, for example srnl.fast_inference = True. It leaves weights unchanged and does not affect training. If you supply separate SRNL queries, their dtype must match the input for fast inference.

For blocks and complete models, use prepare_for_inference after loading weights:

import torch
from visionhope.models import create_model
from visionhope.inference.prepare import prepare_for_inference

model = create_model(
    "visionhope_small", checkpoint_path="weights/visionhope_small.pth"
).cuda().eval()
prepare_for_inference(model)

with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
    logits = model(torch.randn(1, 3, 224, 224, device="cuda"))

prepare_for_inference fuses model parameters in place for fast inference. Keep an unprepared copy if you also need to continue training. The test scripts below enable fast inference by default; pass --mode ordinary to disable it.

Datasets

Download ImageNet-1K (ILSVRC 2012), COCO 2017, and ADE20K from their official websites.

Arrange the datasets as follows, or point the scripts to equivalent locations with --data-root:

data/
├── imagenet/
│   ├── train/<class_name>/*.JPEG
│   └── val/<class_name>/*.JPEG
├── coco/
│   ├── train2017/
│   ├── val2017/
│   └── annotations/instances_{train,val}2017.json
└── ade20k/
    ├── images/{training,validation}/
    └── annotations/{training,validation}/

You can also set IMAGENET_ROOT, COCO_ROOT, and ADE20K_ROOT to your dataset directories. For ADE20K, use the extracted ADEChallengeData2016 directory as the root.

Training

The scripts in scripts/train contain the training recipes. Matching evaluation scripts are in scripts/test.

TaskSizesScript name
ImageNet-1Ktiny, small, basevisionhope_<size>.sh
COCO, Mask R-CNN 1×tiny, small, basevisionhope_<size>_coco_1x.sh
COCO, Mask R-CNN 3×tiny, smallvisionhope_<size>_coco_3x.sh
ADE20K, UPerNettiny, small, basevisionhope_<size>_ade20k.sh

For example, to train VisionHOPE-S on eight GPUs:

# ImageNet-1K
NPROC_PER_NODE=8 bash scripts/train/visionhope_small.sh \
    --data-root data/imagenet --output outputs/visionhope_small

# COCO
NPROC_PER_NODE=8 bash scripts/train/visionhope_small_coco_1x.sh \
    --data-root data/coco --pretrained weights/visionhope_small.pth

# ADE20K
NPROC_PER_NODE=8 bash scripts/train/visionhope_small_ade20k.sh \
    --data-root data/ade20k --pretrained weights/visionhope_small.pth

Training defaults to eight GPUs. Set NPROC_PER_NODE=1 for a single GPU. --batch-size is per GPU; the learning rate scales with the effective global batch unless --lr is given. Arguments appended to a script override its defaults. COCO and ADE20K training use ImageNet-pretrained backbones.

Outputs go to outputs/train/<script-name>/ unless --output is supplied. Use best.pth for evaluation and training_state.pth to resume training:

bash scripts/train/visionhope_small.sh \
    --resume outputs/visionhope_small/training_state.pth \
    --output outputs/visionhope_small

For argument descriptions, defaults, and choices:

python -m visionhope.tasks.cli train classification --help

Replace classification with detection or segmentation for the other tasks.

Evaluation

Run the matching test script with a checkpoint and dataset:

CUDA_VISIBLE_DEVICES=0 bash scripts/test/visionhope_small.sh \
    --data-root data/imagenet --checkpoint weights/visionhope_small.pth

CUDA_VISIBLE_DEVICES=0 bash scripts/test/visionhope_small_coco_1x.sh \
    --data-root data/coco --checkpoint weights/visionhope_small_coco_1x.pth

CUDA_VISIBLE_DEVICES=0 bash scripts/test/visionhope_small_ade20k.sh \
    --data-root data/ade20k --checkpoint weights/visionhope_small_ade20k.pth

The scripts set the evaluation precision and preprocessing for each task. ImageNet evaluation uses FP32, 224 × 224 inputs, and crop ratio 1.0. Pass --precision fp16 or --precision bf16 to a test script for mixed-precision evaluation. Classification accepts --num-classes to match the checkpoint's output classes and --no-cudnn-benchmark to disable cuDNN algorithm benchmarking. For COCO and ADE20K, use --test-scale W H to change the resize scale while keeping the aspect ratio; defaults are 1333 800 and 2048 512, respectively. ADE20K evaluation defaults to batch size 1; larger batches require images of the same size after preprocessing. Results are saved under outputs/test/<script-name>/; pass --output with a different directory for another evaluation of the same model. Use python -m visionhope.tasks.cli inference classification --help for evaluation options; the other task names work here too.

Main results

The tables below summarize ImageNet-1K, COCO, and ADE20K results for the hierarchical VisionHOPE-T / S / B models. See the paper for experimental details. Pretrained checkpoints are available from GitHub Releases, Hugging Face, and Baidu Netdisk (access code: 8866).

ImageNet-1K classification

Trained on ImageNet-1K; single-crop evaluation at 224 × 224. FLOPs include the classification head.

ModelParams (M)FLOPs (G)Top-1 (%) ↑RecipeCKPT
VisionHOPE-T26.64.984.1Train / EvalGitHub / Hugging Face
VisionHOPE-S52.99.885.2Train / EvalGitHub / Hugging Face
VisionHOPE-B91.217.385.6Train / EvalGitHub / Hugging Face

COCO detection and instance segmentation

Mask R-CNN, initialized with ImageNet-1K-pretrained backbones and evaluated on val2017. FLOPs are for the full detector at 1280 × 800. Box AP and mask AP are reported in percent.

1× schedule

BackboneFLOPs (G)Box AP ↑Mask AP ↑RecipeCKPT
VisionHOPE-T26647.943.1Train / EvalGitHub / Hugging Face
VisionHOPE-S36549.544.2Train / EvalGitHub / Hugging Face
VisionHOPE-B51650.545.0Train / EvalGitHub / Hugging Face

3× schedule

BackboneFLOPs (G)Box AP ↑Mask AP ↑RecipeCKPT
VisionHOPE-T26649.444.1Train / EvalGitHub / Hugging Face
VisionHOPE-S36550.545.0Train / EvalGitHub / Hugging Face

ADE20K semantic segmentation

UPerNet, initialized with ImageNet-1K-pretrained backbones; single-scale validation. Parameters are for the complete model; inference FLOPs are measured at 512 × 2048.

BackboneParams (M)FLOPs (G)mIoU (%) ↑RecipeCKPT
VisionHOPE-T55.494249.4Train / EvalGitHub / Hugging Face
VisionHOPE-S81.8104350.3Train / EvalGitHub / Hugging Face
VisionHOPE-B121.2120051.8Train / EvalGitHub / Hugging Face

See Evaluation for precision and preprocessing, and Efficiency for FLOP-counting and timing conventions.

Efficiency

Measure parameters, inference FLOPs, throughput, and GPU memory without a dataset:

bash run_inference.sh complexity --model visionhope_small
bash run_inference.sh efficiency --model visionhope_small

bash run_inference.sh complexity --task detection --model visionhope_small \
    --checkpoint weights/visionhope_small_coco_1x.pth --shape 1280 800
bash run_inference.sh efficiency --task segmentation --model visionhope_small \
    --checkpoint weights/visionhope_small_ade20k.pth --shape 512 2048

--shape takes height and width. Use --help for batch size, precision, and timing options, and --output result.json to save a report. Throughput and memory measurements use fast inference. They measure the model forward pass on synthetic inputs, excluding data loading, preprocessing, and prediction postprocessing. FLOPs count one multiply-add as one operation. COCO counts depend on the proposals generated for the input; the report lists operations outside the FLOP count. memory.eager reports peak CUDA memory, including model weights and inputs; CUDA Graph memory is reported separately.

Citation

If VisionHOPE is useful for your research, please cite the arXiv preprint:

@article{peng2026visionhope,
  title   = {VisionHOPE: Visual Backbones as Self-Modifying Learning Systems},
  author  = {Peng, Siran and Zhang, Tianshuo and Fu, Tianyu and Zhao, Weisong
             and Zhang, Haoyuan and Zhao, Jiankuo and Wu, Minghui and Jiang, Ping
             and Zhu, Xiangyu and Zhao, Chenxu and Lei, Zhen},
  journal = {arXiv preprint arXiv:2609.33325},
  year    = {2026},
  doi     = {10.48550/arXiv.2609.33325},
  url     = {https://arxiv.org/abs/2609.33325}
}

License

VisionHOPE is released under the MIT License. See Third-party notices for third-party attribution and licenses.

Languages

Python

60.5%

Cuda

35.5%

Shell

2.6%

C++

1.4%