internlm/InternLumina-U2

Model

Intern Lumina U2

15

7 commits

3 linked in READMEs

updated Sep 24, 2026

See the code

README

Intern Lumina U2

A Multi-Codebook Diffusion Large Language Model for Omni-Visual Understanding, Image Generation and Editing.

[πŸ“‘ Technical Report (Coming Soon)] Β· 🌐 Project Page Β· πŸ’» Code

Overview

Intern Lumina U2 is a unified multimodal model that brings language, image, video and 3D into a single framework, covering text QA, text-to-image generation, image understanding, image editing, video understanding and 3D understanding with one model. It is a 16B-parameter MoE with 1B active parameters (16B-A1B), pairing an efficient sparse backbone with an 8-codebook fully-discrete visual representation built on AToken.

Checkpoints in this repo

This repo hosts checkpoints of the same architecture trained on different hardware stacks:

  • ascend/ β€” trained on Huawei Ascend NPUs.
  • nvidia/ β€” trained on NVIDIA GPUs.

Usage

Inference code, setup instructions and examples live in the GitHub repo: https://github.com/InternLM/InternLumina-U2. Point CHECKPOINT at one of the subfolders above when running infer_1024_sft.sh.

Benchmarks

Preliminary, partial results. Full comparison tables will appear in the upcoming technical report. See the GitHub repo for the current table.

License

Apache 2.0.

Citation

@misc{internluminau2,
  title  = {Intern Lumina U2},
  author = {{Intern Lumina U2 Team, Shanghai AI Laboratory}},
  year   = {2026},
  note   = {Tech report coming soon}
}
diffusion
endpoints_compatible
image-editing
image-generation
multimodal
safetensors
transformers
unified-model
vision-language

Contributors

qianyu1217

6 commits

qzhangFDU

1 commits

internlm/InternLumina-U2

Model

Intern Lumina U2

15

7 commits

3 linked in READMEs

updated Sep 24, 2026

See the code

README

Intern Lumina U2

A Multi-Codebook Diffusion Large Language Model for Omni-Visual Understanding, Image Generation and Editing.

[πŸ“‘ Technical Report (Coming Soon)] Β· 🌐 Project Page Β· πŸ’» Code

Overview

Intern Lumina U2 is a unified multimodal model that brings language, image, video and 3D into a single framework, covering text QA, text-to-image generation, image understanding, image editing, video understanding and 3D understanding with one model. It is a 16B-parameter MoE with 1B active parameters (16B-A1B), pairing an efficient sparse backbone with an 8-codebook fully-discrete visual representation built on AToken.

Checkpoints in this repo

This repo hosts checkpoints of the same architecture trained on different hardware stacks:

  • ascend/ β€” trained on Huawei Ascend NPUs.
  • nvidia/ β€” trained on NVIDIA GPUs.

Usage

Inference code, setup instructions and examples live in the GitHub repo: https://github.com/InternLM/InternLumina-U2. Point CHECKPOINT at one of the subfolders above when running infer_1024_sft.sh.

Benchmarks

Preliminary, partial results. Full comparison tables will appear in the upcoming technical report. See the GitHub repo for the current table.

License

Apache 2.0.

Citation

@misc{internluminau2,
  title  = {Intern Lumina U2},
  author = {{Intern Lumina U2 Team, Shanghai AI Laboratory}},
  year   = {2026},
  note   = {Tech report coming soon}
}
diffusion
endpoints_compatible
image-editing
image-generation
multimodal
safetensors
transformers
unified-model
vision-language

Contributors

qianyu1217

6 commits

qzhangFDU

1 commits