A Multi-Codebook Diffusion Large Language Model for Omni-Visual Understanding, Image Generation and Editing.
[π Technical Report (Coming Soon)] Β· π Project Page Β· π» Code
Intern Lumina U2 is a unified multimodal model that brings language, image, video and 3D into a single framework, covering text QA, text-to-image generation, image understanding, image editing, video understanding and 3D understanding with one model. It is a 16B-parameter MoE with 1B active parameters (16B-A1B), pairing an efficient sparse backbone with an 8-codebook fully-discrete visual representation built on AToken.
This repo hosts checkpoints of the same architecture trained on different hardware stacks:
ascend/ β trained on Huawei Ascend NPUs.nvidia/ β trained on NVIDIA GPUs.Inference code, setup instructions and examples live in the GitHub repo:
https://github.com/InternLM/InternLumina-U2. Point CHECKPOINT at one of the subfolders above
when running infer_1024_sft.sh.
Preliminary, partial results. Full comparison tables will appear in the upcoming technical report. See the GitHub repo for the current table.
Apache 2.0.
@misc{internluminau2,
title = {Intern Lumina U2},
author = {{Intern Lumina U2 Team, Shanghai AI Laboratory}},
year = {2026},
note = {Tech report coming soon}
}
6 commits
1 commits
A Multi-Codebook Diffusion Large Language Model for Omni-Visual Understanding, Image Generation and Editing.
[π Technical Report (Coming Soon)] Β· π Project Page Β· π» Code
Intern Lumina U2 is a unified multimodal model that brings language, image, video and 3D into a single framework, covering text QA, text-to-image generation, image understanding, image editing, video understanding and 3D understanding with one model. It is a 16B-parameter MoE with 1B active parameters (16B-A1B), pairing an efficient sparse backbone with an 8-codebook fully-discrete visual representation built on AToken.
This repo hosts checkpoints of the same architecture trained on different hardware stacks:
ascend/ β trained on Huawei Ascend NPUs.nvidia/ β trained on NVIDIA GPUs.Inference code, setup instructions and examples live in the GitHub repo:
https://github.com/InternLM/InternLumina-U2. Point CHECKPOINT at one of the subfolders above
when running infer_1024_sft.sh.
Preliminary, partial results. Full comparison tables will appear in the upcoming technical report. See the GitHub repo for the current table.
Apache 2.0.
@misc{internluminau2,
title = {Intern Lumina U2},
author = {{Intern Lumina U2 Team, Shanghai AI Laboratory}},
year = {2026},
note = {Tech report coming soon}
}
6 commits
1 commits