mlboydaisuke/ssdlite320-mobilenetv3-litert

Model

SSDLite320-MobileNetV3-Large — LiteRT (CompiledModel GPU)

0

6 commits

2 linked in READMEs

updated Aug 15, 2026

See the code

README

SSDLite320-MobileNetV3-Large — LiteRT (CompiledModel GPU)

SSDLite320-MobileNetV3 object detection on-device (LiteRT GPU, Pixel 8a)

torchvision ssdlite320_mobilenet_v3_large (COCO, BSD-3) converted patch-free via litert-torch for the LiteRT CompiledModel GPU (ML Drift) path. FP16, 7.2 MB — the fast clean detector (0.59 GMACs).

Device-verified on Pixel 8a (Tensor G3): CompiledModel GPU delegates all 286 graph nodes to OpenCL (Replacing 286 out of 286 node(s) with delegate (LITERT_CL), 1 partition, no CPU fallback) and runs the live camera at ~30 FPS.

I/O

  • Input [1, 3, 320, 320] NCHW, RGB, normalized pixel/127.5 - 1 ∈ [-1, 1] (mean = std = 0.5, not ImageNet); bilinear stretch-resize to 320×320.
  • Output 12 raw head tensors — (cls, box) per 6 feature levels (H = W = 20,10,5,3,2,1):
    • cls[i] = [1, 546, H, W] = 6 anchors × 91 classes
    • box[i] = [1, 24, H, W] = 6 anchors × 4 box deltas
    • 3234 anchors total, 91 classes (COCO 90 + background at index 0).

Decode (runs in app code)

The graph returns raw head outputs so it stays GPU-clean — SSD's built-in DefaultBoxGenerator + NMS postprocess lowers to GPU-rejected GATHER_ND/TOPK/>4D. Decode mirrors torchvision SSD.postprocess_detections + BoxCoder(10,10,5,5): rebuild the 3234 default boxes (scales 0.2–0.95, aspect ratios {1, 2, 3, ½, ⅓}), softmax over 91 → best non-background class → score threshold → decode (dx,dy,dw,dh) against the anchor → per-class NMS. End-to-end this matches stock torchvision 298/300 boxes @ IoU 0.99 on the FP16 tflite.

GPU compatibility

BANNED NONE, Flex/Custom NONE, max tensor ndim 4, 0 dynamic dims. Ops include SUM×8 (MobileNetV3 SqueezeExcitation global pools) and TRANSPOSE×11 — all accepted by Mali ML Drift on device. FP16 is bit-faithful to PyTorch (per-output corr ≥ 0.99999).

Why raw heads: SSD's postprocess lowers to GATHER_ND/TOPK/>4D (GPU-rejected); tapping the 4D head convs and keeping NCHW I/O (no to_channel_last_io, which would blow the MobileNetV3 SqueezeExcitation pools up to GATHER_ND) converts stock-clean with no model-internal patch.

Sample app & conversion

Android sample (live camera) + conversion / validation scripts: https://github.com/john-rocky/LiteRT-Models/tree/main/ssdlite

License

BSD-3-Clause (torchvision code + weights). COCO dataset terms apply to the training data.


Want a different model on-device? Open a request — free, open weights only; the export and its measured numbers get published publicly.

coco
litert
mobilenetv3
object-detection
on-device
ssdlite
tflite

mlboydaisuke/ssdlite320-mobilenetv3-litert

Model

SSDLite320-MobileNetV3-Large — LiteRT (CompiledModel GPU)

0

6 commits

2 linked in READMEs

updated Aug 15, 2026

See the code

README

SSDLite320-MobileNetV3-Large — LiteRT (CompiledModel GPU)

SSDLite320-MobileNetV3 object detection on-device (LiteRT GPU, Pixel 8a)

torchvision ssdlite320_mobilenet_v3_large (COCO, BSD-3) converted patch-free via litert-torch for the LiteRT CompiledModel GPU (ML Drift) path. FP16, 7.2 MB — the fast clean detector (0.59 GMACs).

Device-verified on Pixel 8a (Tensor G3): CompiledModel GPU delegates all 286 graph nodes to OpenCL (Replacing 286 out of 286 node(s) with delegate (LITERT_CL), 1 partition, no CPU fallback) and runs the live camera at ~30 FPS.

I/O

  • Input [1, 3, 320, 320] NCHW, RGB, normalized pixel/127.5 - 1 ∈ [-1, 1] (mean = std = 0.5, not ImageNet); bilinear stretch-resize to 320×320.
  • Output 12 raw head tensors — (cls, box) per 6 feature levels (H = W = 20,10,5,3,2,1):
    • cls[i] = [1, 546, H, W] = 6 anchors × 91 classes
    • box[i] = [1, 24, H, W] = 6 anchors × 4 box deltas
    • 3234 anchors total, 91 classes (COCO 90 + background at index 0).

Decode (runs in app code)

The graph returns raw head outputs so it stays GPU-clean — SSD's built-in DefaultBoxGenerator + NMS postprocess lowers to GPU-rejected GATHER_ND/TOPK/>4D. Decode mirrors torchvision SSD.postprocess_detections + BoxCoder(10,10,5,5): rebuild the 3234 default boxes (scales 0.2–0.95, aspect ratios {1, 2, 3, ½, ⅓}), softmax over 91 → best non-background class → score threshold → decode (dx,dy,dw,dh) against the anchor → per-class NMS. End-to-end this matches stock torchvision 298/300 boxes @ IoU 0.99 on the FP16 tflite.

GPU compatibility

BANNED NONE, Flex/Custom NONE, max tensor ndim 4, 0 dynamic dims. Ops include SUM×8 (MobileNetV3 SqueezeExcitation global pools) and TRANSPOSE×11 — all accepted by Mali ML Drift on device. FP16 is bit-faithful to PyTorch (per-output corr ≥ 0.99999).

Why raw heads: SSD's postprocess lowers to GATHER_ND/TOPK/>4D (GPU-rejected); tapping the 4D head convs and keeping NCHW I/O (no to_channel_last_io, which would blow the MobileNetV3 SqueezeExcitation pools up to GATHER_ND) converts stock-clean with no model-internal patch.

Sample app & conversion

Android sample (live camera) + conversion / validation scripts: https://github.com/john-rocky/LiteRT-Models/tree/main/ssdlite

License

BSD-3-Clause (torchvision code + weights). COCO dataset terms apply to the training data.


Want a different model on-device? Open a request — free, open weights only; the export and its measured numbers get published publicly.

coco
litert
mobilenetv3
object-detection
on-device
ssdlite
tflite